benchgap
Calibration

AIME26 → MMAnswerBench

MMAnswerBench is estimated from AIME26 with a inverse Michaelis–Menten curve fitted on 10 models measured on both: y = 0.84157·(x − 0.0000) / (0.0000 + 1.9159 − x), R² = 0.88, cross-validated error 1.3 pp. It is used for 19 estimates.

Estimated modelAIME26MMAnswerBenchSource
A.X K297.1%86.5%estimated ± 1.3 pp, high confidence
Gemma 4 12B77.5%57.2%estimated ± 1.3 pp, medium confidence
Inkling97.1%86.5%estimated ± 1.3 pp, high confidence
Inkling-Small95.5%83.6%estimated ± 1.3 pp, high confidence
K-EXAONE 2.092.3%78.2%estimated ± 1.3 pp, medium confidence
LFM2.5-8B-A1B50.0%29.7%estimated ± 1.3 pp, medium confidence
Ling 3.0 Flash93.2%79.7%estimated ± 1.3 pp, high confidence
LLaDA2.2-mini35.1%18.8%estimated ± 1.3 pp, medium confidence
LongCat-Flash-Lite-Sparse65.7%44.0%estimated ± 1.3 pp, medium confidence
MAI-Thinking-194.5%81.9%estimated ± 1.3 pp, high confidence
MiniCPM5-1B40.4%22.5%estimated ± 1.3 pp, medium confidence
MiniCPM5-2B86.5%69.3%estimated ± 1.3 pp, medium confidence
Muse Glimmer 30B94.7%82.3%estimated ± 1.3 pp, high confidence
Beam97.8%87.8%estimated ± 1.3 pp, high confidence
Solar Open 295.7%84.0%estimated ± 1.3 pp, high confidence
Solar Pro 495.3%83.3%estimated ± 1.3 pp, high confidence
Ternary Bonsai 2 27B95.8%84.2%estimated ± 1.3 pp, high confidence
ZAYA1-74B-Preview76.4%55.8%estimated ± 1.3 pp, medium confidence
ZAYA1-8B89.1%73.2%estimated ± 1.3 pp, medium confidence