Calibration
HLE w/o tools → MedXpertQA (Text)
MedXpertQA (Text) is estimated from HLE w/o tools with a Michaelis–Menten curve fitted on 5 models measured on both: y = 2.0000·x / (0.99249 + x), R² = 0.44, cross-validated error 8.9 pp. It is used for 16 estimates.
| Estimated model | HLE w/o tools | MedXpertQA (Text) | Source |
|---|---|---|---|
| Claude Fable 5.1 | 60.9% | 76.1% | estimated ± 8.9 pp, low confidence |
| Claude Haiku 5.5 | 45.9% | 63.2% | estimated ± 8.9 pp, low confidence |
| Claude Mythos 5 | 59.0% | 74.6% | estimated ± 8.9 pp, low confidence |
| Claude Opus 5 | 56.3% | 72.4% | estimated ± 8.9 pp, low confidence |
| Claude Opus 5.5 | 64.4% | 78.7% | estimated ± 8.9 pp, low confidence |
| Claude Sonnet 5 | 43.2% | 60.7% | estimated ± 8.9 pp, low confidence |
| Claude Sonnet 5.5 | 56.9% | 72.9% | estimated ± 8.9 pp, low confidence |
| Gemma 4 26B A4B | 8.7% | 16.1% | estimated ± 8.9 pp, low confidence |
| Gemma 4 31B | 19.5% | 32.8% | estimated ± 8.9 pp, low confidence |
| GPT-5.4 mini | 28.2% | 44.3% | estimated ± 8.9 pp, low confidence |
| GPT-5.4 nano | 24.3% | 39.3% | estimated ± 8.9 pp, low confidence |
| GPT-5.4 Pro | 42.7% | 60.2% | estimated ± 8.9 pp, low confidence |
| GPT-5.5 Pro | 43.1% | 60.6% | estimated ± 8.9 pp, low confidence |
| MiMo-V2.5-Pro | 34.0% | 51.0% | estimated ± 8.9 pp, low confidence |
| Muse Spark 1.1 | 52.2% | 68.9% | estimated ± 8.9 pp, low confidence |
| Pareto 26.9 | 49.0% | 66.1% | estimated ± 8.9 pp, low confidence |