benchgap
Calibration

AA-Omniscience Accuracy → SuperGPQA

SuperGPQA is estimated from AA-Omniscience Accuracy with a Michaelis–Menten curve fitted on 16 models measured on both: y = 1.1039·x / (0.14409 + x), R² = 0.72, cross-validated error 7.5 pp. It is used for 167 estimates.

Estimated modelAA-Omniscience AccuracySuperGPQASource
A.X K218.6%62.2%estimated ± 7.5 pp, medium confidence
Apodex 1.131.7%75.9%estimated ± 7.5 pp, medium confidence
Apodex 1.1 Mini31.7%75.9%estimated ± 7.5 pp, medium confidence
Celeris-111.0%47.8%estimated ± 7.5 pp, medium confidence
Claude 3 Haiku17.6%60.7%estimated ± 7.5 pp, medium confidence
Claude 4 Sonnet22.7%67.5%estimated ± 7.5 pp, medium confidence
Claude Fable 565.4%90.5%estimated ± 7.5 pp, low confidence
Claude Fable 5.167.2%90.9%estimated ± 7.5 pp, low confidence
Claude Opus 4.5 Thinking46.6%84.3%estimated ± 7.5 pp, low confidence
Claude Opus 4.6 (Adaptive)47.0%84.5%estimated ± 7.5 pp, low confidence
Claude Opus 4.744.7%83.5%estimated ± 7.5 pp, medium confidence
Claude Opus 4.7 (Adaptive)48.9%85.3%estimated ± 7.5 pp, low confidence
Claude Opus 4.848.8%85.2%estimated ± 7.5 pp, low confidence
Claude Opus 560.9%89.3%estimated ± 7.5 pp, low confidence
Claude Opus 5.566.2%90.7%estimated ± 7.5 pp, low confidence
Claude Sonnet 540.1%81.2%estimated ± 7.5 pp, medium confidence
Claude Sonnet 5.554.0%87.1%estimated ± 7.5 pp, low confidence
Command A+8.9%42.2%estimated ± 7.5 pp, medium confidence
DeepSeek-R130.5%75.0%estimated ± 7.5 pp, medium confidence
DeepSeek V325.5%70.5%estimated ± 7.5 pp, medium confidence
DeepSeek V3 032424.3%69.3%estimated ± 7.5 pp, medium confidence
DeepSeek V3.123.1%68.0%estimated ± 7.5 pp, medium confidence
DeepSeek V3.1 (Reasoning)29.0%73.7%estimated ± 7.5 pp, medium confidence
DeepSeek V3.224.0%69.0%estimated ± 7.5 pp, medium confidence
DeepSeek V4.1 Flash46.4%84.2%estimated ± 7.5 pp, low confidence
DeepSeek V4 Flash 073140.4%81.4%estimated ± 7.5 pp, medium confidence
DeepSeek V4 Pro 081349.1%85.3%estimated ± 7.5 pp, low confidence
Exaone 4.0 1.2B5.0%28.4%estimated ± 7.5 pp, low confidence
Exaone 4.0 32B10.6%46.8%estimated ± 7.5 pp, medium confidence
Gemini 2.5 Flash26.1%71.1%estimated ± 7.5 pp, medium confidence
Gemini 2.5 Pro39.1%80.7%estimated ± 7.5 pp, medium confidence
Gemini 3.1 Pro54.9%87.4%estimated ± 7.5 pp, low confidence
Gemini 3.5 Flash51.9%86.4%estimated ± 7.5 pp, low confidence
Gemini 3.5 Flash-Lite29.5%74.2%estimated ± 7.5 pp, medium confidence
Gemini 3.6 Flash50.0%85.7%estimated ± 7.5 pp, low confidence
Gemini 3.7 Flash55.3%87.6%estimated ± 7.5 pp, low confidence
Gemini 3.8 Flash54.6%87.3%estimated ± 7.5 pp, low confidence
Gemini 3 Flash45.8%84.0%estimated ± 7.5 pp, medium confidence
Gemini 3 Pro55.8%87.7%estimated ± 7.5 pp, low confidence
Gemini 4 Argon49.9%85.7%estimated ± 7.5 pp, low confidence
Gemma 3 27B13.0%52.4%estimated ± 7.5 pp, medium confidence
Gemma 4 12B15.6%57.4%estimated ± 7.5 pp, medium confidence
Gemma 4 26B A4B19.1%62.9%estimated ± 7.5 pp, medium confidence
Gemma 4 31B20.0%64.2%estimated ± 7.5 pp, medium confidence
Gemma 4 E2B6.6%34.7%estimated ± 7.5 pp, low confidence
Gemma 4 E4B8.6%41.3%estimated ± 7.5 pp, medium confidence
GLM-4.5-Air16.3%58.6%estimated ± 7.5 pp, medium confidence
GLM-4.621.4%66.0%estimated ± 7.5 pp, medium confidence
GLM-4.729.3%74.0%estimated ± 7.5 pp, medium confidence
GLM-5.123.7%68.7%estimated ± 7.5 pp, medium confidence
GLM-5.224.3%69.3%estimated ± 7.5 pp, medium confidence
GLM-5.333.9%77.5%estimated ± 7.5 pp, medium confidence
GLM-5-Turbo28.4%73.2%estimated ± 7.5 pp, medium confidence
GLM-5V-Turbo29.3%74.0%estimated ± 7.5 pp, medium confidence
GPT-4.127.8%72.7%estimated ± 7.5 pp, medium confidence
GPT-4.1 mini20.3%64.6%estimated ± 7.5 pp, medium confidence
GPT-4.1 nano13.7%53.8%estimated ± 7.5 pp, medium confidence
GPT-4o19.9%64.0%estimated ± 7.5 pp, medium confidence
GPT-5.137.7%79.9%estimated ± 7.5 pp, medium confidence
GPT-5.1-Codex39.9%81.1%estimated ± 7.5 pp, medium confidence
GPT-5.1-Codex-Max39.9%81.1%estimated ± 7.5 pp, medium confidence
GPT-5.244.3%83.3%estimated ± 7.5 pp, medium confidence
GPT-5.2-Codex41.1%81.7%estimated ± 7.5 pp, medium confidence
GPT-5.3 Codex52.9%86.8%estimated ± 7.5 pp, low confidence
GPT-5.450.8%86.0%estimated ± 7.5 pp, low confidence
GPT-5.4 mini37.5%79.7%estimated ± 7.5 pp, medium confidence
GPT-5.4 nano25.7%70.7%estimated ± 7.5 pp, medium confidence
GPT-5.558.0%88.4%estimated ± 7.5 pp, low confidence
GPT-5.6 Luna42.7%82.5%estimated ± 7.5 pp, medium confidence
GPT-5.6 Sol59.4%88.8%estimated ± 7.5 pp, low confidence
GPT-5.6 Terra46.8%84.4%estimated ± 7.5 pp, low confidence
GPT-5 (high)40.3%81.3%estimated ± 7.5 pp, medium confidence
GPT-5 (medium)39.5%80.9%estimated ± 7.5 pp, medium confidence
GPT-6.1 Sol62.1%89.6%estimated ± 7.5 pp, low confidence
GPT-6 Astra62.6%89.7%estimated ± 7.5 pp, low confidence
GPT-6 Luna43.8%83.1%estimated ± 7.5 pp, medium confidence
GPT-6 Sol54.5%87.3%estimated ± 7.5 pp, low confidence
GPT-OSS 120B21.8%66.5%estimated ± 7.5 pp, medium confidence
GPT-OSS 20B16.0%58.1%estimated ± 7.5 pp, medium confidence
Granite-4.0-350M3.9%23.5%estimated ± 7.5 pp, low confidence
Granite-4.0-H-1B5.2%29.3%estimated ± 7.5 pp, low confidence
Granite-4.0-H-350M3.8%23.0%estimated ± 7.5 pp, low confidence
Granite 4.2 30B10.1%45.5%estimated ± 7.5 pp, medium confidence
Granite 4.2 3B9.2%43.0%estimated ± 7.5 pp, medium confidence
Granite 4.2 8B11.2%48.3%estimated ± 7.5 pp, medium confidence
Grok 440.5%81.4%estimated ± 7.5 pp, medium confidence
Grok 4.1 Fast17.2%60.1%estimated ± 7.5 pp, medium confidence
Grok 4.1 Fast (Reasoning)25.1%70.1%estimated ± 7.5 pp, medium confidence
Grok 4.334.6%77.9%estimated ± 7.5 pp, medium confidence
Grok 4.551.6%86.3%estimated ± 7.5 pp, low confidence
Grok 4.648.2%85.0%estimated ± 7.5 pp, low confidence
Grok 4.747.4%84.7%estimated ± 7.5 pp, low confidence
Grok 4 Fast (Reasoning)22.8%67.6%estimated ± 7.5 pp, medium confidence
Grok Code Fast 123.5%68.4%estimated ± 7.5 pp, medium confidence
Hy332.0%76.1%estimated ± 7.5 pp, medium confidence
Hy3 Preview31.5%75.7%estimated ± 7.5 pp, medium confidence
Inkling41.6%82.0%estimated ± 7.5 pp, medium confidence
Inkling-Small33.2%77.0%estimated ± 7.5 pp, medium confidence
K-Exaone16.4%58.8%estimated ± 7.5 pp, medium confidence
K-EXAONE 2.013.1%52.6%estimated ± 7.5 pp, medium confidence
Kimi K2.632.6%76.6%estimated ± 7.5 pp, medium confidence
Kimi K227.4%72.3%estimated ± 7.5 pp, medium confidence
Kimi K2.5 (Reasoning)35.2%78.3%estimated ± 7.5 pp, medium confidence
Kimi K2.7 Code39.6%80.9%estimated ± 7.5 pp, medium confidence
Kimi K347.6%84.7%estimated ± 7.5 pp, low confidence
LFM2.5-2.6B4.4%25.8%estimated ± 7.5 pp, low confidence
LFM2.5-8B-A1B9.4%43.6%estimated ± 7.5 pp, medium confidence
LFM2.5-VL-1.6B-Extract5.8%31.7%estimated ± 7.5 pp, low confidence
Ling 2.6 Flash15.6%57.4%estimated ± 7.5 pp, medium confidence
Ling 3.0 Flash18.2%61.6%estimated ± 7.5 pp, medium confidence
Ling 3.0 Flash FP818.2%61.6%estimated ± 7.5 pp, medium confidence
Ling 3.0 Flash VL14.4%55.2%estimated ± 7.5 pp, medium confidence
Ling 3.0 Tiny8.5%41.0%estimated ± 7.5 pp, medium confidence
Ling 3.1 Flash29.1%73.8%estimated ± 7.5 pp, medium confidence
Llama 3.1 405B23.2%68.1%estimated ± 7.5 pp, medium confidence
Llama 4 Maverick24.9%69.9%estimated ± 7.5 pp, medium confidence
Llama 4 Scout15.2%56.7%estimated ± 7.5 pp, medium confidence
Mercury 2.522.0%66.7%estimated ± 7.5 pp, medium confidence
MiMo-V2.5-Pro22.4%67.2%estimated ± 7.5 pp, medium confidence
MiMo-V2.6-Flash27.0%72.0%estimated ± 7.5 pp, medium confidence
MiMo-V2.6-Pro34.8%78.1%estimated ± 7.5 pp, medium confidence
MiMo-V2-Flash15.6%57.4%estimated ± 7.5 pp, medium confidence
MiMo-V2-Omni19.3%63.2%estimated ± 7.5 pp, medium confidence
MiMo-V2-Pro26.6%71.6%estimated ± 7.5 pp, medium confidence
MiniMax M2.726.8%71.8%estimated ± 7.5 pp, medium confidence
MiniMax M316.7%59.3%estimated ± 7.5 pp, medium confidence
Mistral Large 219.9%64.0%estimated ± 7.5 pp, medium confidence
Mistral Large 325.0%70.0%estimated ± 7.5 pp, medium confidence
Mistral Large 425.8%70.8%estimated ± 7.5 pp, medium confidence
Mistral Medium 318.3%61.8%estimated ± 7.5 pp, medium confidence
Mistral Medium 3.5 128B24.7%69.7%estimated ± 7.5 pp, medium confidence
Mistral Small 421.7%66.3%estimated ± 7.5 pp, medium confidence
Mistral Small 4 (Reasoning)21.7%66.3%estimated ± 7.5 pp, medium confidence
Muse Glimmer 30B27.0%72.0%estimated ± 7.5 pp, medium confidence
Muse Spark49.6%85.5%estimated ± 7.5 pp, low confidence
Muse Spark 1.152.1%86.5%estimated ± 7.5 pp, low confidence
Muse Spark 1.245.4%83.8%estimated ± 7.5 pp, medium confidence
Muse Spark 1.343.6%83.0%estimated ± 7.5 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP414.4%55.2%estimated ± 7.5 pp, medium confidence
Nemotron 3 Nano 30B17.3%60.2%estimated ± 7.5 pp, medium confidence
Nemotron 3 Nano Omni 30B A3B15.2%56.7%estimated ± 7.5 pp, medium confidence
Nemotron 3 Super 100B24.3%69.3%estimated ± 7.5 pp, medium confidence
Nemotron 3 Ultra21.6%66.2%estimated ± 7.5 pp, medium confidence
Nemotron Ultra 253B20.1%64.3%estimated ± 7.5 pp, medium confidence
North Mini Code18.9%62.6%estimated ± 7.5 pp, medium confidence
Nova Pro16.9%59.6%estimated ± 7.5 pp, medium confidence
o134.5%77.9%estimated ± 7.5 pp, medium confidence
o338.6%80.4%estimated ± 7.5 pp, medium confidence
Phi-414.1%54.6%estimated ± 7.5 pp, medium confidence
Quasar 438B15.5%57.2%estimated ± 7.5 pp, medium confidence
Qwen3.5 397B (Reasoning)24.5%69.5%estimated ± 7.5 pp, medium confidence
Qwen3.8-27B15.6%57.4%estimated ± 7.5 pp, medium confidence
Qwen3.8-Flash-Next24.5%69.5%estimated ± 7.5 pp, medium confidence
Qwen3.8 Max Preview31.7%75.9%estimated ± 7.5 pp, medium confidence
Qwen3 Max24.4%69.4%estimated ± 7.5 pp, medium confidence
Qwen3-Omni-30B-A3B-Instruct14.3%55.0%estimated ± 7.5 pp, medium confidence
Qwen3-Omni-30B-A3B-Thinking14.6%55.6%estimated ± 7.5 pp, medium confidence
Sarvam 105B17.6%60.7%estimated ± 7.5 pp, medium confidence
Sarvam 30B12.6%51.5%estimated ± 7.5 pp, medium confidence
Solar Pro 216.1%58.3%estimated ± 7.5 pp, medium confidence
Solar Pro 318.5%62.1%estimated ± 7.5 pp, medium confidence
Solar Pro 418.9%62.6%estimated ± 7.5 pp, medium confidence
Step 3.7 Flash25.8%70.8%estimated ± 7.5 pp, medium confidence
Step 5 Preview41.5%81.9%estimated ± 7.5 pp, medium confidence
Trinity-Large-Preview22.5%67.3%estimated ± 7.5 pp, medium confidence
Trinity-Large-Thinking22.5%67.3%estimated ± 7.5 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B19.0%62.8%estimated ± 7.5 pp, medium confidence