benchgap
Calibration

AA-Omniscience Accuracy → BioMysteryBench (human-difficult)

BioMysteryBench (human-difficult) is estimated from AA-Omniscience Accuracy with a Michaelis–Menten curve fitted on 6 models measured on both: y = 0.8393·x / (0.42005 + x), R² = 0.58, cross-validated error 5.9 pp. It is used for 177 estimates.

Estimated modelAA-Omniscience AccuracyBioMysteryBench (human-difficult)Source
A.X K218.6%25.8%estimated ± 5.9 pp, low confidence
Apodex 1.1 Mini31.7%36.1%estimated ± 5.9 pp, low confidence
Celeris-111.0%17.4%estimated ± 5.9 pp, low confidence
Claude 3 Haiku17.6%24.8%estimated ± 5.9 pp, low confidence
Claude 4 Sonnet22.7%29.4%estimated ± 5.9 pp, low confidence
Claude Fable 565.4%51.1%estimated ± 5.9 pp, low confidence
Claude Fable 5.167.2%51.6%estimated ± 5.9 pp, low confidence
Claude Opus 4.540.9%41.4%estimated ± 5.9 pp, low confidence
Claude Opus 4.5 Thinking46.6%44.1%estimated ± 5.9 pp, low confidence
Claude Opus 4.645.8%43.8%estimated ± 5.9 pp, low confidence
Claude Opus 4.6 (Adaptive)47.0%44.3%estimated ± 5.9 pp, low confidence
Claude Opus 4.744.7%43.3%estimated ± 5.9 pp, low confidence
Claude Opus 4.7 (Adaptive)48.9%45.1%estimated ± 5.9 pp, low confidence
Claude Opus 4.848.8%45.1%estimated ± 5.9 pp, low confidence
Claude Sonnet 4.638.6%40.2%estimated ± 5.9 pp, low confidence
Claude Sonnet 540.1%41.0%estimated ± 5.9 pp, low confidence
Command A+8.9%14.7%estimated ± 5.9 pp, low confidence
DeepSeek-R130.5%35.3%estimated ± 5.9 pp, low confidence
DeepSeek V325.5%31.7%estimated ± 5.9 pp, low confidence
DeepSeek V3 032424.3%30.8%estimated ± 5.9 pp, low confidence
DeepSeek V3.123.1%29.8%estimated ± 5.9 pp, low confidence
DeepSeek V3.1 (Reasoning)29.0%34.3%estimated ± 5.9 pp, low confidence
DeepSeek V3.224.0%30.5%estimated ± 5.9 pp, low confidence
DeepSeek V4.1 Flash46.4%44.1%estimated ± 5.9 pp, low confidence
DeepSeek V4 Flash 073140.4%41.1%estimated ± 5.9 pp, low confidence
DeepSeek V4 Pro 081349.1%45.2%estimated ± 5.9 pp, low confidence
Exaone 4.0 1.2B5.0%8.9%estimated ± 5.9 pp, low confidence
Exaone 4.0 32B10.6%16.9%estimated ± 5.9 pp, low confidence
Gemini 2.5 Flash26.1%32.2%estimated ± 5.9 pp, low confidence
Gemini 2.5 Pro39.1%40.5%estimated ± 5.9 pp, low confidence
Gemini 3.1 Pro54.9%47.5%estimated ± 5.9 pp, low confidence
Gemini 3.5 Flash51.9%46.4%estimated ± 5.9 pp, low confidence
Gemini 3.5 Flash-Lite29.5%34.6%estimated ± 5.9 pp, low confidence
Gemini 3.6 Flash50.0%45.6%estimated ± 5.9 pp, low confidence
Gemini 3 Flash45.8%43.8%estimated ± 5.9 pp, low confidence
Gemini 3 Pro55.8%47.9%estimated ± 5.9 pp, low confidence
Gemini 4 Argon49.9%45.6%estimated ± 5.9 pp, low confidence
Gemma 3 27B13.0%19.8%estimated ± 5.9 pp, low confidence
Gemma 4 12B15.6%22.7%estimated ± 5.9 pp, low confidence
Gemma 4 26B A4B19.1%26.2%estimated ± 5.9 pp, low confidence
Gemma 4 31B20.0%27.1%estimated ± 5.9 pp, low confidence
Gemma 4 E2B6.6%11.4%estimated ± 5.9 pp, low confidence
Gemma 4 E4B8.6%14.3%estimated ± 5.9 pp, low confidence
GLM-4.5-Air16.3%23.5%estimated ± 5.9 pp, low confidence
GLM-4.621.4%28.3%estimated ± 5.9 pp, low confidence
GLM-4.729.3%34.5%estimated ± 5.9 pp, low confidence
GLM-526.3%32.3%estimated ± 5.9 pp, low confidence
GLM-5.123.7%30.3%estimated ± 5.9 pp, low confidence
GLM-5.224.3%30.8%estimated ± 5.9 pp, low confidence
GLM-5.333.9%37.5%estimated ± 5.9 pp, low confidence
GLM-5-Turbo28.4%33.9%estimated ± 5.9 pp, low confidence
GLM-5V-Turbo29.3%34.5%estimated ± 5.9 pp, low confidence
GPT-4.127.8%33.4%estimated ± 5.9 pp, low confidence
GPT-4.1 mini20.3%27.3%estimated ± 5.9 pp, low confidence
GPT-4.1 nano13.7%20.6%estimated ± 5.9 pp, low confidence
GPT-4o19.9%27.0%estimated ± 5.9 pp, low confidence
GPT-5.137.7%39.7%estimated ± 5.9 pp, low confidence
GPT-5.1-Codex39.9%40.9%estimated ± 5.9 pp, low confidence
GPT-5.1-Codex-Max39.9%40.9%estimated ± 5.9 pp, low confidence
GPT-5.244.3%43.1%estimated ± 5.9 pp, low confidence
GPT-5.2-Codex41.1%41.5%estimated ± 5.9 pp, low confidence
GPT-5.3 Codex52.9%46.8%estimated ± 5.9 pp, low confidence
GPT-5.450.8%45.9%estimated ± 5.9 pp, low confidence
GPT-5.4 mini37.5%39.6%estimated ± 5.9 pp, low confidence
GPT-5.4 nano25.7%31.9%estimated ± 5.9 pp, low confidence
GPT-5.558.0%48.7%estimated ± 5.9 pp, low confidence
GPT-5.6 Luna42.7%42.3%estimated ± 5.9 pp, low confidence
GPT-5.6 Sol59.4%49.2%estimated ± 5.9 pp, low confidence
GPT-5.6 Terra46.8%44.2%estimated ± 5.9 pp, low confidence
GPT-5 (high)40.3%41.1%estimated ± 5.9 pp, low confidence
GPT-5 (medium)39.5%40.7%estimated ± 5.9 pp, low confidence
GPT-6.1 Sol62.1%50.1%estimated ± 5.9 pp, low confidence
GPT-6 Astra62.6%50.2%estimated ± 5.9 pp, low confidence
GPT-6 Luna43.8%42.8%estimated ± 5.9 pp, low confidence
GPT-6 Sol54.5%47.4%estimated ± 5.9 pp, low confidence
GPT-OSS 120B21.8%28.7%estimated ± 5.9 pp, low confidence
GPT-OSS 20B16.0%23.2%estimated ± 5.9 pp, low confidence
Granite-4.0-350M3.9%7.1%estimated ± 5.9 pp, low confidence
Granite-4.0-H-1B5.2%9.2%estimated ± 5.9 pp, low confidence
Granite-4.0-H-350M3.8%7.0%estimated ± 5.9 pp, low confidence
Granite 4.2 30B10.1%16.3%estimated ± 5.9 pp, low confidence
Granite 4.2 3B9.2%15.1%estimated ± 5.9 pp, low confidence
Granite 4.2 8B11.2%17.7%estimated ± 5.9 pp, low confidence
Grok 440.5%41.2%estimated ± 5.9 pp, low confidence
Grok 4.1 Fast17.2%24.4%estimated ± 5.9 pp, low confidence
Grok 4.1 Fast (Reasoning)25.1%31.4%estimated ± 5.9 pp, low confidence
Grok 4.334.6%37.9%estimated ± 5.9 pp, low confidence
Grok 4.551.6%46.3%estimated ± 5.9 pp, low confidence
Grok 4.648.2%44.8%estimated ± 5.9 pp, low confidence
Grok 4.747.4%44.5%estimated ± 5.9 pp, low confidence
Grok 4 Fast (Reasoning)22.8%29.5%estimated ± 5.9 pp, low confidence
Grok Code Fast 123.5%30.1%estimated ± 5.9 pp, low confidence
Hy332.0%36.3%estimated ± 5.9 pp, low confidence
Hy3 Preview31.5%36.0%estimated ± 5.9 pp, low confidence
Inkling41.6%41.8%estimated ± 5.9 pp, low confidence
Inkling-Small33.2%37.1%estimated ± 5.9 pp, low confidence
K-Exaone16.4%23.6%estimated ± 5.9 pp, low confidence
K-EXAONE 2.013.1%20.0%estimated ± 5.9 pp, low confidence
Kimi K2.632.6%36.7%estimated ± 5.9 pp, low confidence
Kimi K227.4%33.1%estimated ± 5.9 pp, low confidence
Kimi K2.535.2%38.3%estimated ± 5.9 pp, low confidence
Kimi K2.5 (Reasoning)35.2%38.3%estimated ± 5.9 pp, low confidence
Kimi K2.7 Code39.6%40.7%estimated ± 5.9 pp, low confidence
Kimi K347.6%44.6%estimated ± 5.9 pp, low confidence
LFM2.5-2.6B4.4%8.0%estimated ± 5.9 pp, low confidence
LFM2.5-8B-A1B9.4%15.3%estimated ± 5.9 pp, low confidence
LFM2.5-VL-1.6B-Extract5.8%10.2%estimated ± 5.9 pp, low confidence
Ling 2.6 Flash15.6%22.7%estimated ± 5.9 pp, low confidence
Ling 3.0 Flash18.2%25.4%estimated ± 5.9 pp, low confidence
Ling 3.0 Flash FP818.2%25.4%estimated ± 5.9 pp, low confidence
Ling 3.0 Flash VL14.4%21.4%estimated ± 5.9 pp, low confidence
Ling 3.0 Tiny8.5%14.1%estimated ± 5.9 pp, low confidence
Ling 3.1 Flash29.1%34.3%estimated ± 5.9 pp, low confidence
Llama 3.1 405B23.2%29.9%estimated ± 5.9 pp, low confidence
Llama 4 Maverick24.9%31.2%estimated ± 5.9 pp, low confidence
Llama 4 Scout15.2%22.3%estimated ± 5.9 pp, low confidence
Mercury 2.522.0%28.8%estimated ± 5.9 pp, low confidence
MiMo-V2.5-Pro22.4%29.2%estimated ± 5.9 pp, low confidence
MiMo-V2.6-Flash27.0%32.8%estimated ± 5.9 pp, low confidence
MiMo-V2.6-Pro34.8%38.0%estimated ± 5.9 pp, low confidence
MiMo-V2-Flash15.6%22.7%estimated ± 5.9 pp, low confidence
MiMo-V2-Omni19.3%26.4%estimated ± 5.9 pp, low confidence
MiMo-V2-Pro26.6%32.5%estimated ± 5.9 pp, low confidence
MiniCPM5-2B8.4%14.0%estimated ± 5.9 pp, low confidence
MiniMax M2.726.8%32.7%estimated ± 5.9 pp, low confidence
MiniMax M316.7%23.9%estimated ± 5.9 pp, low confidence
Mistral Large 219.9%27.0%estimated ± 5.9 pp, low confidence
Mistral Large 325.0%31.3%estimated ± 5.9 pp, low confidence
Mistral Large 425.8%31.9%estimated ± 5.9 pp, low confidence
Mistral Medium 318.3%25.5%estimated ± 5.9 pp, low confidence
Mistral Medium 3.5 128B24.7%31.1%estimated ± 5.9 pp, low confidence
Mistral Small 421.7%28.6%estimated ± 5.9 pp, low confidence
Mistral Small 4 (Reasoning)21.7%28.6%estimated ± 5.9 pp, low confidence
Muse Glimmer 30B27.0%32.8%estimated ± 5.9 pp, low confidence
Muse Spark49.6%45.4%estimated ± 5.9 pp, low confidence
Muse Spark 1.152.1%46.5%estimated ± 5.9 pp, low confidence
Muse Spark 1.245.4%43.6%estimated ± 5.9 pp, low confidence
Muse Spark 1.343.6%42.7%estimated ± 5.9 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP414.4%21.4%estimated ± 5.9 pp, low confidence
Nemotron 3 Nano 30B17.3%24.5%estimated ± 5.9 pp, low confidence
Nemotron 3 Nano Omni 30B A3B15.2%22.3%estimated ± 5.9 pp, low confidence
Nemotron 3 Super 100B24.3%30.8%estimated ± 5.9 pp, low confidence
Nemotron 3 Ultra21.6%28.5%estimated ± 5.9 pp, low confidence
Nemotron Ultra 253B20.1%27.2%estimated ± 5.9 pp, low confidence
North Mini Code18.9%26.0%estimated ± 5.9 pp, low confidence
Nova Pro16.9%24.1%estimated ± 5.9 pp, low confidence
o134.5%37.8%estimated ± 5.9 pp, low confidence
o338.6%40.2%estimated ± 5.9 pp, low confidence
Phi-414.1%21.1%estimated ± 5.9 pp, low confidence
Quasar 438B15.5%22.6%estimated ± 5.9 pp, low confidence
Qwen3.5-122B-A10B24.4%30.8%estimated ± 5.9 pp, low confidence
Qwen3.5-27B20.7%27.7%estimated ± 5.9 pp, low confidence
Qwen3.5-35B-A3B20.1%27.2%estimated ± 5.9 pp, low confidence
Qwen3.5 397B24.5%30.9%estimated ± 5.9 pp, low confidence
Qwen3.5 397B (Reasoning)24.5%30.9%estimated ± 5.9 pp, low confidence
Qwen3.6-27B19.6%26.7%estimated ± 5.9 pp, low confidence
Qwen3.6-35B-A3B18.8%25.9%estimated ± 5.9 pp, low confidence
Qwen 3.6 Max (preview)37.9%39.8%estimated ± 5.9 pp, low confidence
Qwen3.6 Plus26.4%32.4%estimated ± 5.9 pp, low confidence
Qwen3.7 Max31.1%35.7%estimated ± 5.9 pp, low confidence
Qwen3.7 Plus22.5%29.3%estimated ± 5.9 pp, low confidence
Qwen3.8-27B15.6%22.7%estimated ± 5.9 pp, low confidence
Qwen3.8-Flash-Next24.5%30.9%estimated ± 5.9 pp, low confidence
Qwen3.8 Max Preview31.7%36.1%estimated ± 5.9 pp, low confidence
Qwen3 Max24.4%30.8%estimated ± 5.9 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct14.3%21.3%estimated ± 5.9 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking14.6%21.6%estimated ± 5.9 pp, low confidence
Sarvam 105B17.6%24.8%estimated ± 5.9 pp, low confidence
Sarvam 30B12.6%19.4%estimated ± 5.9 pp, low confidence
Solar Pro 216.1%23.3%estimated ± 5.9 pp, low confidence
Solar Pro 318.5%25.7%estimated ± 5.9 pp, low confidence
Solar Pro 418.9%26.0%estimated ± 5.9 pp, low confidence
Step 3.7 Flash25.8%31.9%estimated ± 5.9 pp, low confidence
Step 5 Preview41.5%41.7%estimated ± 5.9 pp, low confidence
Trinity-Large-Preview22.5%29.3%estimated ± 5.9 pp, low confidence
Trinity-Large-Thinking22.5%29.3%estimated ± 5.9 pp, low confidence
Ultravox v0.6 Llama 3.3 70B19.0%26.1%estimated ± 5.9 pp, low confidence