benchgap
Calibration

MMMU-Pro → OmniDocBench 1.5

OmniDocBench 1.5 is estimated from MMMU-Pro with a offset logistic curve fitted on 5 models measured on both: y = 0.0000 + (0.9171 − 0.0000) / (1 + exp(−180.33·(x − 0.7313))), R² = 1.00, cross-validated error 5.9 pp. It is used for 37 estimates.

Estimated modelMMMU-ProOmniDocBench 1.5Source
Claude Opus 4.570.6%0.9%estimated ± 5.9 pp, low confidence
Claude Opus 4.677.3%91.7%estimated ± 5.9 pp, low confidence
Command A+63.0%0.0%estimated ± 5.9 pp, low confidence
dots3-note Preview79.1%91.7%estimated ± 5.9 pp, low confidence
Gemini 3.1 Pro83.9%91.7%estimated ± 5.9 pp, low confidence
Gemini 3.5 Flash83.6%91.7%estimated ± 5.9 pp, low confidence
Gemini 3 Pro81.0%91.7%estimated ± 5.9 pp, low confidence
Gemma 4 12B69.1%0.1%estimated ± 5.9 pp, low confidence
Gemma 4 26B A4B73.8%70.5%estimated ± 5.9 pp, low confidence
Gemma 4 31B76.9%91.6%estimated ± 5.9 pp, low confidence
GPT-5.279.5%91.7%estimated ± 5.9 pp, low confidence
GPT-5.481.2%91.7%estimated ± 5.9 pp, low confidence
GPT-5.4 mini76.6%91.5%estimated ± 5.9 pp, low confidence
GPT-5.4 nano66.1%0.0%estimated ± 5.9 pp, low confidence
GPT-5.581.2%91.7%estimated ± 5.9 pp, low confidence
GPT-5.6 Luna78.4%91.7%estimated ± 5.9 pp, low confidence
GPT-5.6 Sol83.0%91.7%estimated ± 5.9 pp, low confidence
GPT-5.6 Terra80.7%91.7%estimated ± 5.9 pp, low confidence
Grok 4.2075.2%89.5%estimated ± 5.9 pp, low confidence
Grok 4.378.1%91.7%estimated ± 5.9 pp, low confidence
Inkling73.5%60.5%estimated ± 5.9 pp, low confidence
Inkling-Small74.0%75.8%estimated ± 5.9 pp, low confidence
Interfaze Beta71.1%2.3%estimated ± 5.9 pp, low confidence
Kimi K2.679.4%91.7%estimated ± 5.9 pp, low confidence
Kimi K2.578.5%91.7%estimated ± 5.9 pp, low confidence
Kimi K2.5 (Reasoning)78.5%91.7%estimated ± 5.9 pp, low confidence
Kimi K381.6%91.7%estimated ± 5.9 pp, low confidence
LFM2.5-VL-3B30.5%0.0%estimated ± 5.9 pp, low confidence
MiMo-V2.577.9%91.7%estimated ± 5.9 pp, low confidence
Muse Spark80.4%91.7%estimated ± 5.9 pp, low confidence
Pareto 26.978.0%91.7%estimated ± 5.9 pp, low confidence
Qwen3.5 397B79.0%91.7%estimated ± 5.9 pp, low confidence
Qwen3.6-27B75.8%91.0%estimated ± 5.9 pp, low confidence
Qwen3.6 Plus78.8%91.7%estimated ± 5.9 pp, low confidence
Seed 2.1 Pro81.6%91.7%estimated ± 5.9 pp, low confidence
Seed 2.1 Turbo80.1%91.7%estimated ± 5.9 pp, low confidence
Step 5 Preview76.0%91.2%estimated ± 5.9 pp, low confidence