benchgap
Calibration

CritPt → BioMysteryBench (human-solvable)

BioMysteryBench (human-solvable) is estimated from CritPt with a Hill curve fitted on 5 models measured on both: y = 0.8174 + (0.8956 − 0.8174)·x^6.00 / (0.12565^6.00 + x^6.00), R² = 0.90, cross-validated error 0.6 pp. It is used for 185 estimates.

Estimated modelCritPtBioMysteryBench (human-solvable)Source
A.X K28.9%82.6%estimated ± 0.6 pp, low confidence
Apodex 1.14.6%81.8%estimated ± 0.6 pp, low confidence
Apodex 1.1 Mini4.6%81.8%estimated ± 0.6 pp, low confidence
Celeris-10.0%81.7%estimated ± 0.6 pp, low confidence
Claude 3 Haiku0.0%81.7%estimated ± 0.6 pp, low confidence
Claude 4.1 Opus Thinking0.0%81.7%estimated ± 0.6 pp, low confidence
Claude 4 Sonnet1.1%81.7%estimated ± 0.6 pp, low confidence
Claude Fable 528.6%89.5%estimated ± 0.6 pp, medium confidence
Claude Fable 5.129.7%89.5%estimated ± 0.6 pp, medium confidence
Claude Haiku 5.518.9%88.9%estimated ± 0.6 pp, medium confidence
Claude Opus 4.50.3%81.7%estimated ± 0.6 pp, low confidence
Claude Opus 4.5 Thinking4.6%81.8%estimated ± 0.6 pp, low confidence
Claude Opus 4.62.8%81.7%estimated ± 0.6 pp, low confidence
Claude Opus 4.6 (Adaptive)12.6%85.7%estimated ± 0.6 pp, low confidence
Claude Opus 4.75.1%81.8%estimated ± 0.6 pp, low confidence
Claude Opus 4.7 (Adaptive)12.0%85.1%estimated ± 0.6 pp, low confidence
Claude Opus 4.820.9%89.2%estimated ± 0.6 pp, medium confidence
Claude Sonnet 4.60.9%81.7%estimated ± 0.6 pp, low confidence
Claude Sonnet 516.9%88.4%estimated ± 0.6 pp, medium confidence
Command A+0.3%81.7%estimated ± 0.6 pp, low confidence
DeepSeek-R11.4%81.7%estimated ± 0.6 pp, low confidence
DeepSeek V30.0%81.7%estimated ± 0.6 pp, low confidence
DeepSeek V3 03240.0%81.7%estimated ± 0.6 pp, low confidence
DeepSeek V3.10.0%81.7%estimated ± 0.6 pp, low confidence
DeepSeek V3.1 (Reasoning)2.0%81.7%estimated ± 0.6 pp, low confidence
DeepSeek V3.20.9%81.7%estimated ± 0.6 pp, low confidence
DeepSeek V4.1 Flash14.3%87.1%estimated ± 0.6 pp, medium confidence
DeepSeek V4 Flash 073116.6%88.3%estimated ± 0.6 pp, medium confidence
DeepSeek V4 Pro 081318.0%88.8%estimated ± 0.6 pp, medium confidence
Exaone 4.0 1.2B0.0%81.7%estimated ± 0.6 pp, low confidence
Exaone 4.0 32B0.0%81.7%estimated ± 0.6 pp, low confidence
Gemini 2.5 Flash1.4%81.7%estimated ± 0.6 pp, low confidence
Gemini 2.5 Pro2.6%81.7%estimated ± 0.6 pp, low confidence
Gemini 3.1 Pro17.7%88.7%estimated ± 0.6 pp, medium confidence
Gemini 3.5 Flash13.1%86.1%estimated ± 0.6 pp, low confidence
Gemini 3.5 Flash-Lite0.0%81.7%estimated ± 0.6 pp, low confidence
Gemini 3.6 Flash10.6%83.8%estimated ± 0.6 pp, low confidence
Gemini 3 Flash1.4%81.7%estimated ± 0.6 pp, low confidence
Gemini 3 Pro9.1%82.7%estimated ± 0.6 pp, low confidence
Gemini 3 Pro Deep Think25.7%89.5%estimated ± 0.6 pp, medium confidence
Gemini 4 Argon27.1%89.5%estimated ± 0.6 pp, medium confidence
Gemma 3 27B0.0%81.7%estimated ± 0.6 pp, low confidence
Gemma 4 12B0.0%81.7%estimated ± 0.6 pp, low confidence
Gemma 4 26B A4B0.0%81.7%estimated ± 0.6 pp, low confidence
Gemma 4 31B1.4%81.7%estimated ± 0.6 pp, low confidence
Gemma 4 E2B0.0%81.7%estimated ± 0.6 pp, low confidence
Gemma 4 E4B0.6%81.7%estimated ± 0.6 pp, low confidence
GLM-4.5-Air0.0%81.7%estimated ± 0.6 pp, low confidence
GLM-4.60.0%81.7%estimated ± 0.6 pp, low confidence
GLM-4.71.7%81.7%estimated ± 0.6 pp, low confidence
GLM-52.0%81.7%estimated ± 0.6 pp, low confidence
GLM-5.14.6%81.8%estimated ± 0.6 pp, low confidence
GLM-5.220.9%89.2%estimated ± 0.6 pp, medium confidence
GLM-5.319.1%89.0%estimated ± 0.6 pp, medium confidence
GLM-5.3-Flash15.4%87.8%estimated ± 0.6 pp, medium confidence
GLM-5-Turbo0.3%81.7%estimated ± 0.6 pp, low confidence
GLM-5V-Turbo0.6%81.7%estimated ± 0.6 pp, low confidence
GPT-4.10.0%81.7%estimated ± 0.6 pp, low confidence
GPT-4.1 mini0.0%81.7%estimated ± 0.6 pp, low confidence
GPT-4.1 nano0.0%81.7%estimated ± 0.6 pp, low confidence
GPT-4o0.0%81.7%estimated ± 0.6 pp, low confidence
GPT-5.14.9%81.8%estimated ± 0.6 pp, low confidence
GPT-5.1-Codex5.7%81.8%estimated ± 0.6 pp, low confidence
GPT-5.1-Codex-Max5.7%81.8%estimated ± 0.6 pp, low confidence
GPT-5.211.6%84.7%estimated ± 0.6 pp, low confidence
GPT-5.2-Codex8.7%82.5%estimated ± 0.6 pp, low confidence
GPT-5.3 Codex16.9%88.4%estimated ± 0.6 pp, medium confidence
GPT-5.423.4%89.4%estimated ± 0.6 pp, medium confidence
GPT-5.4 mini10.0%83.3%estimated ± 0.6 pp, low confidence
GPT-5.4 nano9.3%82.8%estimated ± 0.6 pp, low confidence
GPT-5.4 Pro30.0%89.5%estimated ± 0.6 pp, medium confidence
GPT-5.527.1%89.5%estimated ± 0.6 pp, medium confidence
GPT-5.5 Pro30.6%89.5%estimated ± 0.6 pp, medium confidence
GPT-5.6 Luna20.6%89.2%estimated ± 0.6 pp, medium confidence
GPT-5.6 Sol32.3%89.5%estimated ± 0.6 pp, low confidence
GPT-5.6 Terra30.0%89.5%estimated ± 0.6 pp, medium confidence
GPT-5 (high)5.7%81.8%estimated ± 0.6 pp, low confidence
GPT-5 (medium)0.0%81.7%estimated ± 0.6 pp, low confidence
GPT-6.1 Sol31.7%89.5%estimated ± 0.6 pp, medium confidence
GPT-6 Astra31.7%89.5%estimated ± 0.6 pp, medium confidence
GPT-6 Luna19.4%89.0%estimated ± 0.6 pp, medium confidence
GPT-6 Sol30.9%89.5%estimated ± 0.6 pp, medium confidence
GPT-OSS 120B1.1%81.7%estimated ± 0.6 pp, low confidence
GPT-OSS 20B1.4%81.7%estimated ± 0.6 pp, low confidence
Granite-4.0-350M0.0%81.7%estimated ± 0.6 pp, low confidence
Granite-4.0-H-1B0.0%81.7%estimated ± 0.6 pp, low confidence
Granite-4.0-H-350M0.0%81.7%estimated ± 0.6 pp, low confidence
Granite 4.2 30B0.3%81.7%estimated ± 0.6 pp, low confidence
Granite 4.2 3B0.0%81.7%estimated ± 0.6 pp, low confidence
Granite 4.2 8B0.3%81.7%estimated ± 0.6 pp, low confidence
Grok 42.0%81.7%estimated ± 0.6 pp, low confidence
Grok 4.1 Fast0.0%81.7%estimated ± 0.6 pp, low confidence
Grok 4.1 Fast (Reasoning)2.9%81.7%estimated ± 0.6 pp, low confidence
Grok 4.38.0%82.2%estimated ± 0.6 pp, low confidence
Grok 4.515.4%87.8%estimated ± 0.6 pp, medium confidence
Grok 4.617.1%88.5%estimated ± 0.6 pp, medium confidence
Grok 4.717.7%88.7%estimated ± 0.6 pp, medium confidence
Grok 4 Fast (Reasoning)2.9%81.7%estimated ± 0.6 pp, low confidence
Grok Code Fast 10.0%81.7%estimated ± 0.6 pp, low confidence
Hy34.9%81.8%estimated ± 0.6 pp, low confidence
Hy3 Preview4.9%81.8%estimated ± 0.6 pp, low confidence
Hy4 preview16.9%88.4%estimated ± 0.6 pp, medium confidence
Inkling5.4%81.8%estimated ± 0.6 pp, low confidence
Inkling-Small8.3%82.3%estimated ± 0.6 pp, low confidence
K-Exaone1.1%81.7%estimated ± 0.6 pp, low confidence
K-EXAONE 2.00.9%81.7%estimated ± 0.6 pp, low confidence
Kimi K2.68.0%82.2%estimated ± 0.6 pp, low confidence
Kimi K20.0%81.7%estimated ± 0.6 pp, low confidence
Kimi K2.53.1%81.7%estimated ± 0.6 pp, low confidence
Kimi K2.5 (Reasoning)3.1%81.7%estimated ± 0.6 pp, low confidence
Kimi K2.7 Code10.0%83.3%estimated ± 0.6 pp, low confidence
Kimi K323.4%89.4%estimated ± 0.6 pp, medium confidence
LFM2.5-2.6B0.0%81.7%estimated ± 0.6 pp, low confidence
LFM2.5-8B-A1B0.0%81.7%estimated ± 0.6 pp, low confidence
LFM2.5-VL-1.6B-Extract0.0%81.7%estimated ± 0.6 pp, low confidence
Ling 2.6 Flash0.0%81.7%estimated ± 0.6 pp, low confidence
Ling 3.0 Flash1.7%81.7%estimated ± 0.6 pp, low confidence
Ling 3.0 Flash FP81.7%81.7%estimated ± 0.6 pp, low confidence
Ling 3.0 Flash VL2.0%81.7%estimated ± 0.6 pp, low confidence
Ling 3.0 Tiny0.0%81.7%estimated ± 0.6 pp, low confidence
Ling 3.1 Flash18.0%88.8%estimated ± 0.6 pp, medium confidence
Llama 3.1 405B0.0%81.7%estimated ± 0.6 pp, low confidence
Llama 4 Maverick0.0%81.7%estimated ± 0.6 pp, low confidence
Llama 4 Scout0.0%81.7%estimated ± 0.6 pp, low confidence
Mercury 2.50.0%81.7%estimated ± 0.6 pp, low confidence
MiMo-V2.5-Pro4.0%81.7%estimated ± 0.6 pp, low confidence
MiMo-V2.6-Flash12.0%85.1%estimated ± 0.6 pp, low confidence
MiMo-V2.6-Pro26.6%89.5%estimated ± 0.6 pp, medium confidence
MiMo-V2-Flash0.0%81.7%estimated ± 0.6 pp, low confidence
MiMo-V2-Omni1.1%81.7%estimated ± 0.6 pp, low confidence
MiMo-V2-Pro0.3%81.7%estimated ± 0.6 pp, low confidence
MiniCPM5-2B0.3%81.7%estimated ± 0.6 pp, low confidence
MiniMax M2.70.6%81.7%estimated ± 0.6 pp, low confidence
MiniMax M33.7%81.7%estimated ± 0.6 pp, low confidence
Mistral Large 20.0%81.7%estimated ± 0.6 pp, low confidence
Mistral Large 30.0%81.7%estimated ± 0.6 pp, low confidence
Mistral Large 410.6%83.8%estimated ± 0.6 pp, low confidence
Mistral Medium 30.0%81.7%estimated ± 0.6 pp, low confidence
Mistral Medium 3.5 128B0.0%81.7%estimated ± 0.6 pp, low confidence
Mistral Small 40.3%81.7%estimated ± 0.6 pp, low confidence
Mistral Small 4 (Reasoning)0.3%81.7%estimated ± 0.6 pp, low confidence
Muse Glimmer 30B2.6%81.7%estimated ± 0.6 pp, low confidence
Muse Spark11.3%84.4%estimated ± 0.6 pp, low confidence
Muse Spark 1.115.1%87.6%estimated ± 0.6 pp, medium confidence
Muse Spark 1.217.7%88.7%estimated ± 0.6 pp, medium confidence
Muse Spark 1.324.9%89.4%estimated ± 0.6 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP40.0%81.7%estimated ± 0.6 pp, low confidence
Nemotron 3 Nano 30B0.9%81.7%estimated ± 0.6 pp, low confidence
Nemotron 3 Nano Omni 30B A3B0.0%81.7%estimated ± 0.6 pp, low confidence
Nemotron 3 Super 100B3.1%81.7%estimated ± 0.6 pp, low confidence
Nemotron 3 Ultra3.1%81.7%estimated ± 0.6 pp, low confidence
Nemotron Ultra 253B0.0%81.7%estimated ± 0.6 pp, low confidence
North Mini Code0.3%81.7%estimated ± 0.6 pp, low confidence
Nova Pro0.0%81.7%estimated ± 0.6 pp, low confidence
o10.3%81.7%estimated ± 0.6 pp, low confidence
o31.1%81.7%estimated ± 0.6 pp, low confidence
Phi-40.0%81.7%estimated ± 0.6 pp, low confidence
Quasar 438B9.4%82.9%estimated ± 0.6 pp, low confidence
Qwen3.5-122B-A10B0.6%81.7%estimated ± 0.6 pp, low confidence
Qwen3.5-27B0.9%81.7%estimated ± 0.6 pp, low confidence
Qwen3.5-35B-A3B0.9%81.7%estimated ± 0.6 pp, low confidence
Qwen3.5 397B0.9%81.7%estimated ± 0.6 pp, low confidence
Qwen3.5 397B (Reasoning)0.9%81.7%estimated ± 0.6 pp, low confidence
Qwen3.6-27B1.1%81.7%estimated ± 0.6 pp, low confidence
Qwen3.6-35B-A3B0.3%81.7%estimated ± 0.6 pp, low confidence
Qwen 3.6 Max (preview)3.7%81.7%estimated ± 0.6 pp, low confidence
Qwen3.6 Plus2.9%81.7%estimated ± 0.6 pp, low confidence
Qwen3.7 Max13.4%86.4%estimated ± 0.6 pp, low confidence
Qwen3.7 Plus9.1%82.7%estimated ± 0.6 pp, low confidence
Qwen3.8-27B5.4%81.8%estimated ± 0.6 pp, low confidence
Qwen3.8-Flash-Next11.1%84.3%estimated ± 0.6 pp, low confidence
Qwen3.8 Max Preview17.7%88.7%estimated ± 0.6 pp, medium confidence
Qwen3 Max0.0%81.7%estimated ± 0.6 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct0.0%81.7%estimated ± 0.6 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking0.0%81.7%estimated ± 0.6 pp, low confidence
Sarvam 105B0.0%81.7%estimated ± 0.6 pp, low confidence
Sarvam 30B0.3%81.7%estimated ± 0.6 pp, low confidence
Solar Pro 20.0%81.7%estimated ± 0.6 pp, low confidence
Solar Pro 30.0%81.7%estimated ± 0.6 pp, low confidence
Solar Pro 45.4%81.8%estimated ± 0.6 pp, low confidence
Step 3.7 Flash2.3%81.7%estimated ± 0.6 pp, low confidence
Step 5 Preview20.9%89.2%estimated ± 0.6 pp, medium confidence
Trinity-Large-Preview0.9%81.7%estimated ± 0.6 pp, low confidence
Trinity-Large-Thinking0.9%81.7%estimated ± 0.6 pp, low confidence
Ultravox v0.6 Llama 3.3 70B0.0%81.7%estimated ± 0.6 pp, low confidence