Calibration
MMMU-Pro → MathVision
MathVision is estimated from MMMU-Pro with a inverse Michaelis–Menten curve fitted on 13 models measured on both: y = 1.34771·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.64, cross-validated error 3.6 pp. It is used for 15 estimates.
| Estimated model | MMMU-Pro | MathVision | Source |
|---|---|---|---|
| Gemma 4 26B A4B | 73.8% | 78.8% | estimated ± 3.6 pp, high confidence |
| Gemma 4 31B | 76.9% | 84.2% | estimated ± 3.6 pp, high confidence |
| GPT-5.4 mini | 76.6% | 83.7% | estimated ± 3.6 pp, high confidence |
| GPT-5.4 nano | 66.1% | 66.5% | estimated ± 3.6 pp, medium confidence |
| GPT-5.5 | 81.2% | 92.1% | estimated ± 3.6 pp, high confidence |
| GPT-5.6 Luna | 78.4% | 86.9% | estimated ± 3.6 pp, high confidence |
| GPT-5.6 Sol | 83.0% | 95.6% | estimated ± 3.6 pp, medium confidence |
| GPT-5.6 Terra | 80.7% | 91.2% | estimated ± 3.6 pp, high confidence |
| Grok 4.3 | 78.1% | 86.3% | estimated ± 3.6 pp, high confidence |
| Interfaze Beta | 71.1% | 74.3% | estimated ± 3.6 pp, high confidence |
| Kimi K2.5 (Reasoning) | 78.5% | 87.1% | estimated ± 3.6 pp, high confidence |
| LFM2.5-VL-3B | 30.5% | 24.3% | estimated ± 3.6 pp, medium confidence |
| MiniMax M3 | 78.1% | 86.3% | estimated ± 3.6 pp, high confidence |
| Pareto 26.9 | 78.0% | 86.2% | estimated ± 3.6 pp, high confidence |
| Step 5 Preview | 76.0% | 82.6% | estimated ± 3.6 pp, high confidence |