benchgap
Calibration

HLE w/o tools → MedXpertQA (Text)

MedXpertQA (Text) is estimated from HLE w/o tools with a Michaelis–Menten curve fitted on 5 models measured on both: y = 2.0000·x / (0.99249 + x), R² = 0.44, cross-validated error 8.9 pp. It is used for 16 estimates.

Estimated modelHLE w/o toolsMedXpertQA (Text)Source
Claude Fable 5.160.9%76.1%estimated ± 8.9 pp, low confidence
Claude Haiku 5.545.9%63.2%estimated ± 8.9 pp, low confidence
Claude Mythos 559.0%74.6%estimated ± 8.9 pp, low confidence
Claude Opus 556.3%72.4%estimated ± 8.9 pp, low confidence
Claude Opus 5.564.4%78.7%estimated ± 8.9 pp, low confidence
Claude Sonnet 543.2%60.7%estimated ± 8.9 pp, low confidence
Claude Sonnet 5.556.9%72.9%estimated ± 8.9 pp, low confidence
Gemma 4 26B A4B8.7%16.1%estimated ± 8.9 pp, low confidence
Gemma 4 31B19.5%32.8%estimated ± 8.9 pp, low confidence
GPT-5.4 mini28.2%44.3%estimated ± 8.9 pp, low confidence
GPT-5.4 nano24.3%39.3%estimated ± 8.9 pp, low confidence
GPT-5.4 Pro42.7%60.2%estimated ± 8.9 pp, low confidence
GPT-5.5 Pro43.1%60.6%estimated ± 8.9 pp, low confidence
MiMo-V2.5-Pro34.0%51.0%estimated ± 8.9 pp, low confidence
Muse Spark 1.152.2%68.9%estimated ± 8.9 pp, low confidence
Pareto 26.949.0%66.1%estimated ± 8.9 pp, low confidence