benchgap
Calibration

CharXiv → OmniDocBench 1.5

OmniDocBench 1.5 is estimated from CharXiv with a Michaelis–Menten curve fitted on 5 models measured on both: y = 2.0000·x / (1.08154 + x), R² = 0.38, cross-validated error 7.4 pp. It is used for 8 estimates.

Estimated modelCharXivOmniDocBench 1.5Source
Claude Mythos 593.5%92.7%estimated ± 7.4 pp, low confidence
Claude Opus 4.889.9%90.8%estimated ± 7.4 pp, low confidence
Gemini 3.1 Flash-Lite73.2%80.7%estimated ± 7.4 pp, low confidence
GLM-5.3-Flash89.4%90.5%estimated ± 7.4 pp, low confidence
Muse Spark 1.188.4%89.9%estimated ± 7.4 pp, low confidence
Qwen3.8-Omni-Flash91.4%91.6%estimated ± 7.4 pp, low confidence
Sakana Fugu85.1%88.1%estimated ± 7.4 pp, low confidence
Sakana Fugu-Ultra86.6%88.9%estimated ± 7.4 pp, low confidence