benchgap
Calibration

CharXiv w/o tools → MMMU-Pro

MMMU-Pro is estimated from CharXiv w/o tools with a offset logistic curve fitted on 5 models measured on both: y = 0.7372 + (0.8229 − 0.7372) / (1 + exp(−115.50·(x − 0.8265))), R² = 1.00, cross-validated error 1.4 pp. It is used for 9 estimates.

Estimated modelCharXiv w/o toolsMMMU-ProSource
Claude Mythos 588.9%82.3%estimated ± 1.4 pp, low confidence
Claude Opus 4.7 (Adaptive)82.1%76.7%estimated ± 1.4 pp, medium confidence
Claude Opus 4.880.5%74.4%estimated ± 1.4 pp, medium confidence
Claude Sonnet 577.0%73.7%estimated ± 1.4 pp, low confidence
Gemini 3.7 Flash84.5%81.4%estimated ± 1.4 pp, medium confidence
Gemini 3.8 Flash86.2%82.2%estimated ± 1.4 pp, medium confidence
Qwen3.8-27B83.7%80.3%estimated ± 1.4 pp, medium confidence
Qwen3.8-Flash-Next84.6%81.5%estimated ± 1.4 pp, medium confidence
Qwen3.8-Omni-Flash83.5%79.9%estimated ± 1.4 pp, medium confidence