Calibration
MMMU-Pro → OmniDocBench 1.5
OmniDocBench 1.5 is estimated from MMMU-Pro with a offset logistic curve fitted on 5 models measured on both: y = 0.0000 + (0.9171 − 0.0000) / (1 + exp(−180.33·(x − 0.7313))), R² = 1.00, cross-validated error 5.9 pp. It is used for 37 estimates.
| Estimated model | MMMU-Pro | OmniDocBench 1.5 | Source |
|---|---|---|---|
| Claude Opus 4.5 | 70.6% | 0.9% | estimated ± 5.9 pp, low confidence |
| Claude Opus 4.6 | 77.3% | 91.7% | estimated ± 5.9 pp, low confidence |
| Command A+ | 63.0% | 0.0% | estimated ± 5.9 pp, low confidence |
| dots3-note Preview | 79.1% | 91.7% | estimated ± 5.9 pp, low confidence |
| Gemini 3.1 Pro | 83.9% | 91.7% | estimated ± 5.9 pp, low confidence |
| Gemini 3.5 Flash | 83.6% | 91.7% | estimated ± 5.9 pp, low confidence |
| Gemini 3 Pro | 81.0% | 91.7% | estimated ± 5.9 pp, low confidence |
| Gemma 4 12B | 69.1% | 0.1% | estimated ± 5.9 pp, low confidence |
| Gemma 4 26B A4B | 73.8% | 70.5% | estimated ± 5.9 pp, low confidence |
| Gemma 4 31B | 76.9% | 91.6% | estimated ± 5.9 pp, low confidence |
| GPT-5.2 | 79.5% | 91.7% | estimated ± 5.9 pp, low confidence |
| GPT-5.4 | 81.2% | 91.7% | estimated ± 5.9 pp, low confidence |
| GPT-5.4 mini | 76.6% | 91.5% | estimated ± 5.9 pp, low confidence |
| GPT-5.4 nano | 66.1% | 0.0% | estimated ± 5.9 pp, low confidence |
| GPT-5.5 | 81.2% | 91.7% | estimated ± 5.9 pp, low confidence |
| GPT-5.6 Luna | 78.4% | 91.7% | estimated ± 5.9 pp, low confidence |
| GPT-5.6 Sol | 83.0% | 91.7% | estimated ± 5.9 pp, low confidence |
| GPT-5.6 Terra | 80.7% | 91.7% | estimated ± 5.9 pp, low confidence |
| Grok 4.20 | 75.2% | 89.5% | estimated ± 5.9 pp, low confidence |
| Grok 4.3 | 78.1% | 91.7% | estimated ± 5.9 pp, low confidence |
| Inkling | 73.5% | 60.5% | estimated ± 5.9 pp, low confidence |
| Inkling-Small | 74.0% | 75.8% | estimated ± 5.9 pp, low confidence |
| Interfaze Beta | 71.1% | 2.3% | estimated ± 5.9 pp, low confidence |
| Kimi K2.6 | 79.4% | 91.7% | estimated ± 5.9 pp, low confidence |
| Kimi K2.5 | 78.5% | 91.7% | estimated ± 5.9 pp, low confidence |
| Kimi K2.5 (Reasoning) | 78.5% | 91.7% | estimated ± 5.9 pp, low confidence |
| Kimi K3 | 81.6% | 91.7% | estimated ± 5.9 pp, low confidence |
| LFM2.5-VL-3B | 30.5% | 0.0% | estimated ± 5.9 pp, low confidence |
| MiMo-V2.5 | 77.9% | 91.7% | estimated ± 5.9 pp, low confidence |
| Muse Spark | 80.4% | 91.7% | estimated ± 5.9 pp, low confidence |
| Pareto 26.9 | 78.0% | 91.7% | estimated ± 5.9 pp, low confidence |
| Qwen3.5 397B | 79.0% | 91.7% | estimated ± 5.9 pp, low confidence |
| Qwen3.6-27B | 75.8% | 91.0% | estimated ± 5.9 pp, low confidence |
| Qwen3.6 Plus | 78.8% | 91.7% | estimated ± 5.9 pp, low confidence |
| Seed 2.1 Pro | 81.6% | 91.7% | estimated ± 5.9 pp, low confidence |
| Seed 2.1 Turbo | 80.1% | 91.7% | estimated ± 5.9 pp, low confidence |
| Step 5 Preview | 76.0% | 91.2% | estimated ± 5.9 pp, low confidence |