Calibration
MMMU-Pro → ScreenSpot Pro
ScreenSpot Pro is estimated from MMMU-Pro with a inverse Michaelis–Menten curve fitted on 11 models measured on both: y = 3.11373·(x − 0.4000) / (0.4000 + 2.0000 − x), R² = 0.55, cross-validated error 10.0 pp. It is used for 15 estimates.
| Estimated model | MMMU-Pro | ScreenSpot Pro | Source |
|---|---|---|---|
| Gemma 4 26B A4B | 73.8% | 63.3% | estimated ± 10.0 pp, low confidence |
| Gemma 4 31B | 76.9% | 70.4% | estimated ± 10.0 pp, low confidence |
| GPT-5.4 mini | 76.6% | 69.7% | estimated ± 10.0 pp, low confidence |
| GPT-5.4 nano | 66.1% | 46.7% | estimated ± 10.0 pp, low confidence |
| GPT-5.5 | 81.2% | 80.8% | estimated ± 10.0 pp, low confidence |
| GPT-5.6 Luna | 78.4% | 74.0% | estimated ± 10.0 pp, low confidence |
| GPT-5.6 Sol | 83.0% | 85.3% | estimated ± 10.0 pp, low confidence |
| GPT-5.6 Terra | 80.7% | 79.5% | estimated ± 10.0 pp, low confidence |
| Grok 4.3 | 78.1% | 73.3% | estimated ± 10.0 pp, low confidence |
| Interfaze Beta | 71.1% | 57.3% | estimated ± 10.0 pp, low confidence |
| Kimi K2.5 | 78.5% | 74.2% | estimated ± 10.0 pp, low confidence |
| Kimi K2.5 (Reasoning) | 78.5% | 74.2% | estimated ± 10.0 pp, low confidence |
| MiniMax M3 | 78.1% | 73.3% | estimated ± 10.0 pp, low confidence |
| Pareto 26.9 | 78.0% | 73.0% | estimated ± 10.0 pp, low confidence |
| Step 5 Preview | 76.0% | 68.3% | estimated ± 10.0 pp, low confidence |