benchgap
Calibration

MMMU-Pro → Video-MME (with subtitle)

Video-MME (with subtitle) is estimated from MMMU-Pro with a Hill curve fitted on 6 models measured on both: y = 0.8034 + (1.2000 − 0.8034)·x^6.00 / (1.00186^6.00 + x^6.00), R² = 0.52, cross-validated error 1.8 pp. It is used for 16 estimates.

Estimated modelMMMU-ProVideo-MME (with subtitle)Source
Claude Opus 4.677.3%87.2%estimated ± 1.8 pp, medium confidence
Gemma 4 12B69.1%84.2%estimated ± 1.8 pp, low confidence
Gemma 4 26B A4B73.8%85.8%estimated ± 1.8 pp, low confidence
Gemma 4 31B76.9%87.1%estimated ± 1.8 pp, medium confidence
GPT-5.4 mini76.6%86.9%estimated ± 1.8 pp, medium confidence
GPT-5.4 nano66.1%83.4%estimated ± 1.8 pp, low confidence
GPT-5.581.2%89.1%estimated ± 1.8 pp, medium confidence
GPT-5.6 Luna78.4%87.7%estimated ± 1.8 pp, medium confidence
GPT-5.6 Sol83.0%90.0%estimated ± 1.8 pp, low confidence
GPT-5.6 Terra80.7%88.8%estimated ± 1.8 pp, medium confidence
Grok 4.378.1%87.6%estimated ± 1.8 pp, medium confidence
Interfaze Beta71.1%84.8%estimated ± 1.8 pp, low confidence
Kimi K2.5 (Reasoning)78.5%87.8%estimated ± 1.8 pp, medium confidence
LFM2.5-VL-3B30.5%80.4%estimated ± 1.8 pp, low confidence
Pareto 26.978.0%87.6%estimated ± 1.8 pp, medium confidence
Step 5 Preview76.0%86.7%estimated ± 1.8 pp, medium confidence