benchgap
Calibration

CharXiv → MedXpertQA (MM)

MedXpertQA (MM) is estimated from CharXiv with a linear curve fitted on 6 models measured on both: y = 0.4023·x + 0.4283, R² = 0.54, cross-validated error 6.4 pp. It is used for 13 estimates.

Estimated modelCharXivMedXpertQA (MM)Source
Claude Mythos 593.5%80.4%estimated ± 6.4 pp, low confidence
Claude Opus 4.7 (Adaptive)91.0%79.4%estimated ± 6.4 pp, low confidence
Claude Opus 4.889.9%79.0%estimated ± 6.4 pp, low confidence
Claude Sonnet 4.677.4%74.0%estimated ± 6.4 pp, low confidence
Claude Sonnet 588.3%78.4%estimated ± 6.4 pp, low confidence
Gemini 3.1 Flash-Lite73.2%72.3%estimated ± 6.4 pp, low confidence
Gemini 3.7 Flash88.7%78.5%estimated ± 6.4 pp, low confidence
GLM-5.3-Flash89.4%78.8%estimated ± 6.4 pp, low confidence
Muse Spark 1.188.4%78.4%estimated ± 6.4 pp, low confidence
Nemotron 3 Nano Omni 30B A3B76.3%73.5%estimated ± 6.4 pp, low confidence
Qwen3.5-122B-A10B77.2%73.9%estimated ± 6.4 pp, low confidence
Sakana Fugu85.1%77.1%estimated ± 6.4 pp, low confidence
Sakana Fugu-Ultra86.6%77.7%estimated ± 6.4 pp, low confidence