Calibration
MMMU-Pro → VideoMMMU
VideoMMMU is estimated from MMMU-Pro with a offset logistic curve fitted on 11 models measured on both: y = 0.8412 + (0.8952 − 0.8412) / (1 + exp(−83.51·(x − 0.8026))), R² = 0.76, cross-validated error 1.0 pp. It is used for 31 estimates.
| Estimated model | MMMU-Pro | VideoMMMU | Source |
|---|---|---|---|
| Claude Opus 4.6 | 77.3% | 84.5% | estimated ± 1.0 pp, high confidence |
| Command A+ | 63.0% | 84.1% | estimated ± 1.0 pp, medium confidence |
| Gemini 3.1 Pro | 83.9% | 89.3% | estimated ± 1.0 pp, medium confidence |
| Gemini 3.5 Flash | 83.6% | 89.2% | estimated ± 1.0 pp, medium confidence |
| Gemma 4 12B | 69.1% | 84.1% | estimated ± 1.0 pp, medium confidence |
| Gemma 4 26B A4B | 73.8% | 84.1% | estimated ± 1.0 pp, high confidence |
| Gemma 4 31B | 76.9% | 84.4% | estimated ± 1.0 pp, high confidence |
| GPT-5.2 | 79.5% | 86.0% | estimated ± 1.0 pp, high confidence |
| GPT-5.4 | 81.2% | 87.8% | estimated ± 1.0 pp, high confidence |
| GPT-5.4 mini | 76.6% | 84.4% | estimated ± 1.0 pp, high confidence |
| GPT-5.4 nano | 66.1% | 84.1% | estimated ± 1.0 pp, medium confidence |
| GPT-5.5 | 81.2% | 87.8% | estimated ± 1.0 pp, high confidence |
| GPT-5.6 Luna | 78.4% | 85.1% | estimated ± 1.0 pp, high confidence |
| GPT-5.6 Sol | 83.0% | 89.0% | estimated ± 1.0 pp, medium confidence |
| GPT-5.6 Terra | 80.7% | 87.3% | estimated ± 1.0 pp, high confidence |
| Grok 4.20 | 75.2% | 84.2% | estimated ± 1.0 pp, high confidence |
| Grok 4.3 | 78.1% | 84.9% | estimated ± 1.0 pp, high confidence |
| Inkling | 73.5% | 84.1% | estimated ± 1.0 pp, high confidence |
| Inkling-Small | 74.0% | 84.1% | estimated ± 1.0 pp, high confidence |
| Interfaze Beta | 71.1% | 84.1% | estimated ± 1.0 pp, high confidence |
| Kimi K2.6 | 79.4% | 85.9% | estimated ± 1.0 pp, high confidence |
| Kimi K2.5 (Reasoning) | 78.5% | 85.1% | estimated ± 1.0 pp, high confidence |
| Kimi K3 | 81.6% | 88.2% | estimated ± 1.0 pp, high confidence |
| LFM2.5-VL-3B | 30.5% | 84.1% | estimated ± 1.0 pp, medium confidence |
| MiMo-V2.5 | 77.9% | 84.8% | estimated ± 1.0 pp, high confidence |
| Muse Glimmer 30B | 74.0% | 84.1% | estimated ± 1.0 pp, high confidence |
| Muse Spark | 80.4% | 87.0% | estimated ± 1.0 pp, high confidence |
| Pareto 26.9 | 78.0% | 84.8% | estimated ± 1.0 pp, high confidence |
| Seed 2.1 Pro | 81.6% | 88.2% | estimated ± 1.0 pp, high confidence |
| Seed 2.1 Turbo | 80.1% | 86.6% | estimated ± 1.0 pp, high confidence |
| Step 5 Preview | 76.0% | 84.3% | estimated ± 1.0 pp, high confidence |