benchgap
Calibration

CritPt → HLE

HLE is estimated from CritPt with a Hill curve fitted on 51 models measured on both: y = 0.1895 + (1.2000 − 0.1895)·x^0.55 / (0.75804^0.55 + x^0.55), R² = 0.55, cross-validated error 9.5 pp. It is used for 90 estimates.

Estimated modelCritPtHLESource
Apodex 1.1 Mini4.6%36.7%estimated ± 9.5 pp, medium confidence
Celeris-10.0%19.0%estimated ± 9.5 pp, medium confidence
Claude 3 Haiku0.0%19.0%estimated ± 9.5 pp, medium confidence
Claude 4.1 Opus Thinking0.0%19.0%estimated ± 9.5 pp, medium confidence
Claude 4 Sonnet1.1%27.8%estimated ± 9.5 pp, medium confidence
Claude Opus 4.75.1%37.5%estimated ± 9.5 pp, medium confidence
Command A+0.3%23.5%estimated ± 9.5 pp, medium confidence
DeepSeek-R11.4%29.0%estimated ± 9.5 pp, medium confidence
DeepSeek V3 03240.0%19.0%estimated ± 9.5 pp, medium confidence
DeepSeek V3.10.0%19.0%estimated ± 9.5 pp, medium confidence
DeepSeek V3.1 (Reasoning)2.0%30.9%estimated ± 9.5 pp, medium confidence
DeepSeek V3.20.9%27.0%estimated ± 9.5 pp, medium confidence
Exaone 4.0 1.2B0.0%19.0%estimated ± 9.5 pp, medium confidence
Exaone 4.0 32B0.0%19.0%estimated ± 9.5 pp, medium confidence
Gemini 2.5 Flash1.4%29.0%estimated ± 9.5 pp, medium confidence
Gemini 3.5 Flash-Lite0.0%19.0%estimated ± 9.5 pp, medium confidence
Gemini 3 Flash1.4%29.0%estimated ± 9.5 pp, medium confidence
Gemini 4 Argon27.1%55.5%estimated ± 9.5 pp, medium confidence
Gemma 3 27B0.0%19.0%estimated ± 9.5 pp, medium confidence
GLM-4.5-Air0.0%19.0%estimated ± 9.5 pp, medium confidence
GLM-4.60.0%19.0%estimated ± 9.5 pp, medium confidence
GLM-5.319.1%51.1%estimated ± 9.5 pp, medium confidence
GLM-5.3-Flash15.4%48.6%estimated ± 9.5 pp, medium confidence
GLM-5-Turbo0.3%23.5%estimated ± 9.5 pp, medium confidence
GLM-5V-Turbo0.6%25.5%estimated ± 9.5 pp, medium confidence
GPT-4o0.0%19.0%estimated ± 9.5 pp, medium confidence
GPT-5.1-Codex5.7%38.5%estimated ± 9.5 pp, medium confidence
GPT-5.1-Codex-Max5.7%38.5%estimated ± 9.5 pp, medium confidence
GPT-5.2-Codex8.7%42.4%estimated ± 9.5 pp, medium confidence
GPT-5.3 Codex16.9%49.6%estimated ± 9.5 pp, medium confidence
GPT-5 (high)5.7%38.5%estimated ± 9.5 pp, medium confidence
GPT-5 (medium)0.0%19.0%estimated ± 9.5 pp, medium confidence
GPT-OSS 120B1.1%27.8%estimated ± 9.5 pp, medium confidence
GPT-OSS 20B1.4%29.0%estimated ± 9.5 pp, medium confidence
Granite-4.0-350M0.0%19.0%estimated ± 9.5 pp, medium confidence
Granite-4.0-H-1B0.0%19.0%estimated ± 9.5 pp, medium confidence
Granite-4.0-H-350M0.0%19.0%estimated ± 9.5 pp, medium confidence
Grok 42.0%30.9%estimated ± 9.5 pp, medium confidence
Grok 4.1 Fast0.0%19.0%estimated ± 9.5 pp, medium confidence
Grok 4.1 Fast (Reasoning)2.9%33.2%estimated ± 9.5 pp, medium confidence
Grok 4.717.7%50.2%estimated ± 9.5 pp, medium confidence
Grok 4 Fast (Reasoning)2.9%33.2%estimated ± 9.5 pp, medium confidence
Grok Code Fast 10.0%19.0%estimated ± 9.5 pp, medium confidence
Hy34.9%37.2%estimated ± 9.5 pp, medium confidence
K-Exaone1.1%27.8%estimated ± 9.5 pp, medium confidence
Kimi K20.0%19.0%estimated ± 9.5 pp, medium confidence
Kimi K2.7 Code10.0%43.8%estimated ± 9.5 pp, medium confidence
LFM2.5-2.6B0.0%19.0%estimated ± 9.5 pp, medium confidence
LFM2.5-8B-A1B0.0%19.0%estimated ± 9.5 pp, medium confidence
LFM2.5-VL-1.6B-Extract0.0%19.0%estimated ± 9.5 pp, medium confidence
Ling 3.0 Flash VL2.0%30.9%estimated ± 9.5 pp, medium confidence
Ling 3.0 Tiny0.0%19.0%estimated ± 9.5 pp, medium confidence
Ling 3.1 Flash18.0%50.4%estimated ± 9.5 pp, medium confidence
Llama 3.1 405B0.0%19.0%estimated ± 9.5 pp, medium confidence
Llama 4 Maverick0.0%19.0%estimated ± 9.5 pp, medium confidence
Llama 4 Scout0.0%19.0%estimated ± 9.5 pp, medium confidence
MiMo-V2.6-Flash12.0%45.8%estimated ± 9.5 pp, medium confidence
MiMo-V2.6-Pro26.6%55.2%estimated ± 9.5 pp, medium confidence
MiMo-V2-Omni1.1%27.8%estimated ± 9.5 pp, medium confidence
MiMo-V2-Pro0.3%23.5%estimated ± 9.5 pp, medium confidence
MiniMax M33.7%35.0%estimated ± 9.5 pp, medium confidence
Mistral Large 20.0%19.0%estimated ± 9.5 pp, medium confidence
Mistral Large 30.0%19.0%estimated ± 9.5 pp, medium confidence
Mistral Large 410.6%44.4%estimated ± 9.5 pp, medium confidence
Mistral Medium 30.0%19.0%estimated ± 9.5 pp, medium confidence
Mistral Medium 3.5 128B0.0%19.0%estimated ± 9.5 pp, medium confidence
Mistral Small 40.3%23.5%estimated ± 9.5 pp, medium confidence
Mistral Small 4 (Reasoning)0.3%23.5%estimated ± 9.5 pp, medium confidence
Muse Glimmer 30B2.6%32.5%estimated ± 9.5 pp, medium confidence
Muse Spark 1.217.7%50.2%estimated ± 9.5 pp, medium confidence
Muse Spark 1.324.9%54.4%estimated ± 9.5 pp, medium confidence
Nemotron 3 Nano 30B0.9%27.0%estimated ± 9.5 pp, medium confidence
Nemotron 3 Super 100B3.1%33.7%estimated ± 9.5 pp, medium confidence
Nemotron Ultra 253B0.0%19.0%estimated ± 9.5 pp, medium confidence
North Mini Code0.3%23.5%estimated ± 9.5 pp, medium confidence
Nova Pro0.0%19.0%estimated ± 9.5 pp, medium confidence
o31.1%27.8%estimated ± 9.5 pp, medium confidence
Phi-40.0%19.0%estimated ± 9.5 pp, medium confidence
Quasar 438B9.4%43.2%estimated ± 9.5 pp, medium confidence
Qwen3.5 397B (Reasoning)0.9%27.0%estimated ± 9.5 pp, medium confidence
Qwen3.8 Max Preview17.7%50.2%estimated ± 9.5 pp, medium confidence
Qwen3 Max0.0%19.0%estimated ± 9.5 pp, medium confidence
Qwen3-Omni-30B-A3B-Instruct0.0%19.0%estimated ± 9.5 pp, medium confidence
Qwen3-Omni-30B-A3B-Thinking0.0%19.0%estimated ± 9.5 pp, medium confidence
Sarvam 105B0.0%19.0%estimated ± 9.5 pp, medium confidence
Sarvam 30B0.3%23.5%estimated ± 9.5 pp, medium confidence
Solar Pro 20.0%19.0%estimated ± 9.5 pp, medium confidence
Solar Pro 30.0%19.0%estimated ± 9.5 pp, medium confidence
Step 3.7 Flash2.3%31.7%estimated ± 9.5 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B0.0%19.0%estimated ± 9.5 pp, medium confidence