benchgap
Calibration

AA-MMMU-Pro → MedXpertQA (MM)

MedXpertQA (MM) is estimated from AA-MMMU-Pro with a offset logistic curve fitted on 6 models measured on both: y = 0.0000 + (0.7792 − 0.0000) / (1 + exp(−37.57·(x − 0.6832))), R² = 0.93, cross-validated error 6.5 pp. It is used for 64 estimates.

Estimated modelAA-MMMU-ProMedXpertQA (MM)Source
Apodex 1.179.2%76.6%estimated ± 6.5 pp, low confidence
Apodex 1.1 Mini79.2%76.6%estimated ± 6.5 pp, low confidence
Claude 3 Haiku30.8%0.0%estimated ± 6.5 pp, low confidence
Claude 4.1 Opus Thinking67.9%35.9%estimated ± 6.5 pp, low confidence
Claude 4 Sonnet62.4%7.6%estimated ± 6.5 pp, low confidence
Claude Opus 4.5 Thinking74.0%69.7%estimated ± 6.5 pp, low confidence
Claude Opus 4.6 (Adaptive)75.4%72.8%estimated ± 6.5 pp, low confidence
Claude Opus 4.776.4%74.3%estimated ± 6.5 pp, low confidence
Claude Opus 584.7%77.8%estimated ± 6.5 pp, low confidence
Claude Opus 5.587.7%77.9%estimated ± 6.5 pp, low confidence
DeepSeek V4.1 Flash77.0%75.0%estimated ± 6.5 pp, low confidence
Gemini 1.5 Pro55.0%0.5%estimated ± 6.5 pp, low confidence
Gemini 2.5 Flash65.5%20.1%estimated ± 6.5 pp, low confidence
Gemini 2.5 Pro74.9%71.9%estimated ± 6.5 pp, low confidence
Gemini 3.5 Flash-Lite79.0%76.5%estimated ± 6.5 pp, low confidence
Gemini 3.6 Flash83.2%77.6%estimated ± 6.5 pp, low confidence
Gemini 3.8 Flash85.6%77.8%estimated ± 6.5 pp, low confidence
Gemini 3 Flash78.6%76.3%estimated ± 6.5 pp, low confidence
Gemma 3 27B48.0%0.0%estimated ± 6.5 pp, low confidence
Gemma 4 E2B44.6%0.0%estimated ± 6.5 pp, low confidence
Gemma 4 E4B51.4%0.1%estimated ± 6.5 pp, low confidence
GLM-5V-Turbo72.8%65.7%estimated ± 6.5 pp, low confidence
GPT-4.161.2%5.0%estimated ± 6.5 pp, low confidence
GPT-4.1 mini58.7%2.0%estimated ± 6.5 pp, low confidence
GPT-4.1 nano40.1%0.0%estimated ± 6.5 pp, low confidence
GPT-4o mini41.5%0.0%estimated ± 6.5 pp, low confidence
GPT-5.175.5%73.0%estimated ± 6.5 pp, low confidence
GPT-5.1-Codex72.5%64.5%estimated ± 6.5 pp, low confidence
GPT-5.1-Codex-Max72.5%64.5%estimated ± 6.5 pp, low confidence
GPT-5.2-Codex76.3%74.2%estimated ± 6.5 pp, low confidence
GPT-5.3 Codex78.5%76.3%estimated ± 6.5 pp, low confidence
GPT-5 (high)74.2%70.2%estimated ± 6.5 pp, low confidence
GPT-5 (medium)74.3%70.5%estimated ± 6.5 pp, low confidence
GPT-6.1 Sol86.0%77.8%estimated ± 6.5 pp, low confidence
GPT-6 Astra86.9%77.8%estimated ± 6.5 pp, low confidence
GPT-6 Luna79.7%76.9%estimated ± 6.5 pp, low confidence
GPT-6 Sol82.9%77.6%estimated ± 6.5 pp, low confidence
Grok 468.8%42.5%estimated ± 6.5 pp, low confidence
Grok 4.1 Fast48.4%0.0%estimated ± 6.5 pp, low confidence
Grok 4.1 Fast (Reasoning)63.3%10.3%estimated ± 6.5 pp, low confidence
Grok 4.580.4%77.1%estimated ± 6.5 pp, low confidence
Grok 4 Fast (Reasoning)61.8%6.2%estimated ± 6.5 pp, low confidence
LFM2.5-VL-1.6B-Extract26.5%0.0%estimated ± 6.5 pp, low confidence
Ling 3.0 Flash VL79.0%76.5%estimated ± 6.5 pp, low confidence
Llama 4 Maverick62.1%6.9%estimated ± 6.5 pp, low confidence
Llama 4 Scout52.9%0.2%estimated ± 6.5 pp, low confidence
MiMo-V2.6-Flash73.1%66.8%estimated ± 6.5 pp, low confidence
MiMo-V2-Omni69.9%50.2%estimated ± 6.5 pp, low confidence
Mistral Large 355.7%0.7%estimated ± 6.5 pp, low confidence
Mistral Large 476.4%74.3%estimated ± 6.5 pp, low confidence
Mistral Medium 353.0%0.2%estimated ± 6.5 pp, low confidence
Mistral Medium 3.5 128B64.9%16.9%estimated ± 6.5 pp, low confidence
Mistral Small 456.8%1.0%estimated ± 6.5 pp, low confidence
Mistral Small 4 (Reasoning)56.8%1.0%estimated ± 6.5 pp, low confidence
Nova Pro44.3%0.0%estimated ± 6.5 pp, low confidence
o370.1%51.5%estimated ± 6.5 pp, low confidence
Phi-4 Multimodal Instruct14.5%0.0%estimated ± 6.5 pp, low confidence
Qwen3.5-27B75.0%72.1%estimated ± 6.5 pp, low confidence
Qwen3.5-35B-A3B72.7%65.3%estimated ± 6.5 pp, low confidence
Qwen3.5 397B (Reasoning)52.7%0.2%estimated ± 6.5 pp, low confidence
Qwen3.8 Max Preview82.8%77.6%estimated ± 6.5 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct55.5%0.6%estimated ± 6.5 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking60.2%3.5%estimated ± 6.5 pp, low confidence
Step 3.7 Flash75.3%72.6%estimated ± 6.5 pp, low confidence