benchgap
Calibration

GPQA → HLE

HLE is estimated from GPQA with a inverse Michaelis–Menten curve fitted on 42 models measured on both: y = 0.05839·(x − 0.0000) / (0.0000 + 1.0375 − x), R² = 0.58, cross-validated error 8.2 pp. It is used for 25 estimates.

Estimated modelGPQAHLESource
Claude 3.5 Sonnet59.4%7.8%estimated ± 8.2 pp, low confidence
DeepSeek V359.1%7.7%estimated ± 8.2 pp, low confidence
Gemma 4 E2B43.4%4.2%estimated ± 8.2 pp, low confidence
Gemma 4 E4B58.6%7.6%estimated ± 8.2 pp, low confidence
GPT-4.166.3%10.3%estimated ± 8.2 pp, low confidence
GPT-4.1 mini64.2%9.5%estimated ± 8.2 pp, low confidence
GPT-4.1 nano50.3%5.5%estimated ± 8.2 pp, low confidence
Granite 4.2 30B66.4%10.4%estimated ± 8.2 pp, low confidence
Granite 4.2 3B54.8%6.5%estimated ± 8.2 pp, low confidence
Granite 4.2 8B64.1%9.5%estimated ± 8.2 pp, low confidence
Interfaze Beta89.9%37.9%estimated ± 8.2 pp, medium confidence
Kimi K2.5 (Reasoning)87.6%31.7%estimated ± 8.2 pp, medium confidence
LFM2.5-230M25.4%1.9%estimated ± 8.2 pp, low confidence
LFM2.5-VL-450M25.7%1.9%estimated ± 8.2 pp, low confidence
Ling 2.6 Flash59.0%7.7%estimated ± 8.2 pp, low confidence
Ling 3.0 Flash FP884.0%24.8%estimated ± 8.2 pp, medium confidence
MAI-Thinking-184.2%25.1%estimated ± 8.2 pp, medium confidence
MiMo-V2-Flash83.7%24.4%estimated ± 8.2 pp, medium confidence
Nemotron 3 Nano Omni 30B A3B72.2%13.4%estimated ± 8.2 pp, low confidence
o175.7%15.8%estimated ± 8.2 pp, low confidence
o1-pro79.0%18.6%estimated ± 8.2 pp, low confidence
o3-mini77.2%17.0%estimated ± 8.2 pp, low confidence
Soofi S 30B-A3B43.4%4.2%estimated ± 8.2 pp, low confidence
ZAYA1-74B-Preview57.3%7.2%estimated ± 8.2 pp, low confidence
ZAYA1-8B71.0%12.7%estimated ± 8.2 pp, low confidence