benchgap
Calibration

MMMU-Pro → MedXpertQA (MM)

MedXpertQA (MM) is estimated from MMMU-Pro with a Hill curve fitted on 8 models measured on both: y = 0.0007 + (1.2000 − 0.0007)·x^6.00 / (0.73672^6.00 + x^6.00), R² = 0.96, cross-validated error 2.4 pp. It is used for 34 estimates.

Estimated modelMMMU-ProMedXpertQA (MM)Source
Claude Opus 4.570.6%52.4%estimated ± 2.4 pp, high confidence
Command A+63.0%33.8%estimated ± 2.4 pp, medium confidence
dots3-note Preview79.1%72.6%estimated ± 2.4 pp, high confidence
Gemini 3.5 Flash83.6%81.7%estimated ± 2.4 pp, high confidence
Gemini 3 Pro81.0%76.6%estimated ± 2.4 pp, high confidence
Gemma 4 26B A4B73.8%60.3%estimated ± 2.4 pp, high confidence
Gemma 4 31B76.9%67.7%estimated ± 2.4 pp, high confidence
GPT-5.279.5%73.5%estimated ± 2.4 pp, high confidence
GPT-5.4 mini76.6%67.0%estimated ± 2.4 pp, high confidence
GPT-5.4 nano66.1%41.2%estimated ± 2.4 pp, medium confidence
GPT-5.581.2%77.1%estimated ± 2.4 pp, high confidence
GPT-5.6 Luna78.4%71.1%estimated ± 2.4 pp, high confidence
GPT-5.6 Sol83.0%80.6%estimated ± 2.4 pp, high confidence
GPT-5.6 Terra80.7%76.0%estimated ± 2.4 pp, high confidence
Grok 4.378.1%70.4%estimated ± 2.4 pp, high confidence
Inkling73.5%59.6%estimated ± 2.4 pp, high confidence
Inkling-Small74.0%60.8%estimated ± 2.4 pp, high confidence
Interfaze Beta71.1%53.7%estimated ± 2.4 pp, high confidence
Kimi K2.679.4%73.3%estimated ± 2.4 pp, high confidence
Kimi K2.578.5%71.3%estimated ± 2.4 pp, high confidence
Kimi K2.5 (Reasoning)78.5%71.3%estimated ± 2.4 pp, high confidence
Kimi K381.6%77.9%estimated ± 2.4 pp, high confidence
LFM2.5-VL-3B30.5%0.7%estimated ± 2.4 pp, medium confidence
MiMo-V2.577.9%70.0%estimated ± 2.4 pp, high confidence
MiniMax M378.1%70.4%estimated ± 2.4 pp, high confidence
Muse Glimmer 30B74.0%60.8%estimated ± 2.4 pp, high confidence
Pareto 26.978.0%70.2%estimated ± 2.4 pp, high confidence
Qwen3.5 397B79.0%72.4%estimated ± 2.4 pp, high confidence
Qwen3.6-27B75.8%65.1%estimated ± 2.4 pp, high confidence
Qwen3.6-35B-A3B75.3%64.0%estimated ± 2.4 pp, high confidence
Qwen3.6 Plus78.8%72.0%estimated ± 2.4 pp, high confidence
Seed 2.1 Pro81.6%77.9%estimated ± 2.4 pp, high confidence
Seed 2.1 Turbo80.1%74.8%estimated ± 2.4 pp, high confidence
Step 5 Preview76.0%65.6%estimated ± 2.4 pp, high confidence