benchgap
Calibration

AA-Omniscience Accuracy → MMLU-Pro

MMLU-Pro is estimated from AA-Omniscience Accuracy with a offset logistic curve fitted on 36 models measured on both: y = 0.0000 + (0.8574 − 0.0000) / (1 + exp(−28.24·(x − 0.0340))), R² = 0.77, cross-validated error 3.5 pp. It is used for 108 estimates.

Estimated modelAA-Omniscience AccuracyMMLU-ProSource
A.X K218.6%84.6%estimated ± 3.5 pp, high confidence
Apodex 1.131.7%85.7%estimated ± 3.5 pp, high confidence
Apodex 1.1 Mini31.7%85.7%estimated ± 3.5 pp, high confidence
Claude 3 Haiku17.6%84.2%estimated ± 3.5 pp, high confidence
Claude 4 Sonnet22.7%85.4%estimated ± 3.5 pp, high confidence
Claude Fable 565.4%85.7%estimated ± 3.5 pp, medium confidence
Claude Opus 4.5 Thinking46.6%85.7%estimated ± 3.5 pp, high confidence
Claude Opus 4.6 (Adaptive)47.0%85.7%estimated ± 3.5 pp, high confidence
Claude Opus 4.744.7%85.7%estimated ± 3.5 pp, high confidence
Command A+8.9%70.8%estimated ± 3.5 pp, high confidence
DeepSeek-R130.5%85.7%estimated ± 3.5 pp, high confidence
DeepSeek V3 032424.3%85.5%estimated ± 3.5 pp, high confidence
DeepSeek V3.123.1%85.4%estimated ± 3.5 pp, high confidence
DeepSeek V3.1 (Reasoning)29.0%85.7%estimated ± 3.5 pp, high confidence
DeepSeek V3.224.0%85.5%estimated ± 3.5 pp, high confidence
Exaone 4.0 1.2B5.0%52.4%estimated ± 3.5 pp, medium confidence
Gemini 2.5 Flash26.1%85.6%estimated ± 3.5 pp, high confidence
Gemini 3.5 Flash-Lite29.5%85.7%estimated ± 3.5 pp, high confidence
Gemini 3.6 Flash50.0%85.7%estimated ± 3.5 pp, medium confidence
Gemini 3.7 Flash55.3%85.7%estimated ± 3.5 pp, medium confidence
Gemini 3.8 Flash54.6%85.7%estimated ± 3.5 pp, medium confidence
Gemini 3 Flash45.8%85.7%estimated ± 3.5 pp, high confidence
Gemini 3 Pro55.8%85.7%estimated ± 3.5 pp, medium confidence
Gemini 4 Argon49.9%85.7%estimated ± 3.5 pp, medium confidence
Gemma 3 27B13.0%80.4%estimated ± 3.5 pp, high confidence
GLM-4.5-Air16.3%83.6%estimated ± 3.5 pp, high confidence
GLM-4.621.4%85.2%estimated ± 3.5 pp, high confidence
GLM-5.123.7%85.5%estimated ± 3.5 pp, high confidence
GLM-5.333.9%85.7%estimated ± 3.5 pp, high confidence
GLM-5-Turbo28.4%85.7%estimated ± 3.5 pp, high confidence
GLM-5V-Turbo29.3%85.7%estimated ± 3.5 pp, high confidence
GPT-4o19.9%84.9%estimated ± 3.5 pp, high confidence
GPT-5.137.7%85.7%estimated ± 3.5 pp, high confidence
GPT-5.1-Codex39.9%85.7%estimated ± 3.5 pp, high confidence
GPT-5.1-Codex-Max39.9%85.7%estimated ± 3.5 pp, high confidence
GPT-5.2-Codex41.1%85.7%estimated ± 3.5 pp, high confidence
GPT-5.3 Codex52.9%85.7%estimated ± 3.5 pp, medium confidence
GPT-5 (high)40.3%85.7%estimated ± 3.5 pp, high confidence
GPT-5 (medium)39.5%85.7%estimated ± 3.5 pp, high confidence
GPT-6.1 Sol62.1%85.7%estimated ± 3.5 pp, medium confidence
GPT-6 Luna43.8%85.7%estimated ± 3.5 pp, high confidence
GPT-6 Sol54.5%85.7%estimated ± 3.5 pp, medium confidence
GPT-OSS 120B21.8%85.3%estimated ± 3.5 pp, high confidence
GPT-OSS 20B16.0%83.4%estimated ± 3.5 pp, high confidence
Granite-4.0-350M3.9%45.9%estimated ± 3.5 pp, medium confidence
Granite-4.0-H-1B5.2%53.5%estimated ± 3.5 pp, medium confidence
Granite-4.0-H-350M3.8%45.3%estimated ± 3.5 pp, medium confidence
Grok 440.5%85.7%estimated ± 3.5 pp, high confidence
Grok 4.1 Fast17.2%84.0%estimated ± 3.5 pp, high confidence
Grok 4.1 Fast (Reasoning)25.1%85.6%estimated ± 3.5 pp, high confidence
Grok 4.551.6%85.7%estimated ± 3.5 pp, medium confidence
Grok 4.648.2%85.7%estimated ± 3.5 pp, high confidence
Grok 4.747.4%85.7%estimated ± 3.5 pp, high confidence
Grok 4 Fast (Reasoning)22.8%85.4%estimated ± 3.5 pp, high confidence
Grok Code Fast 123.5%85.5%estimated ± 3.5 pp, high confidence
Hy332.0%85.7%estimated ± 3.5 pp, high confidence
K-Exaone16.4%83.6%estimated ± 3.5 pp, high confidence
Kimi K227.4%85.6%estimated ± 3.5 pp, high confidence
Kimi K2.7 Code39.6%85.7%estimated ± 3.5 pp, high confidence
LFM2.5-2.6B4.4%48.9%estimated ± 3.5 pp, medium confidence
LFM2.5-8B-A1B9.4%72.4%estimated ± 3.5 pp, high confidence
LFM2.5-VL-1.6B-Extract5.8%56.9%estimated ± 3.5 pp, medium confidence
Ling 3.0 Flash VL14.4%82.1%estimated ± 3.5 pp, high confidence
Ling 3.0 Tiny8.5%69.3%estimated ± 3.5 pp, high confidence
Ling 3.1 Flash29.1%85.7%estimated ± 3.5 pp, high confidence
Llama 3.1 405B23.2%85.4%estimated ± 3.5 pp, high confidence
Llama 4 Maverick24.9%85.5%estimated ± 3.5 pp, high confidence
Llama 4 Scout15.2%82.8%estimated ± 3.5 pp, high confidence
Mercury 2.522.0%85.3%estimated ± 3.5 pp, high confidence
MiMo-V2.6-Flash27.0%85.6%estimated ± 3.5 pp, high confidence
MiMo-V2.6-Pro34.8%85.7%estimated ± 3.5 pp, high confidence
MiMo-V2-Omni19.3%84.8%estimated ± 3.5 pp, high confidence
MiMo-V2-Pro26.6%85.6%estimated ± 3.5 pp, high confidence
MiniMax M2.726.8%85.6%estimated ± 3.5 pp, high confidence
MiniMax M316.7%83.8%estimated ± 3.5 pp, high confidence
Mistral Large 219.9%84.9%estimated ± 3.5 pp, high confidence
Mistral Large 325.0%85.6%estimated ± 3.5 pp, high confidence
Mistral Large 425.8%85.6%estimated ± 3.5 pp, high confidence
Mistral Medium 318.3%84.5%estimated ± 3.5 pp, high confidence
Mistral Medium 3.5 128B24.7%85.5%estimated ± 3.5 pp, high confidence
Mistral Small 421.7%85.3%estimated ± 3.5 pp, high confidence
Mistral Small 4 (Reasoning)21.7%85.3%estimated ± 3.5 pp, high confidence
Muse Glimmer 30B27.0%85.6%estimated ± 3.5 pp, high confidence
Muse Spark 1.245.4%85.7%estimated ± 3.5 pp, high confidence
Muse Spark 1.343.6%85.7%estimated ± 3.5 pp, high confidence
Nemotron 3 Nano 30B17.3%84.1%estimated ± 3.5 pp, high confidence
Nemotron 3 Super 100B24.3%85.5%estimated ± 3.5 pp, high confidence
Nemotron Ultra 253B20.1%85.0%estimated ± 3.5 pp, high confidence
North Mini Code18.9%84.7%estimated ± 3.5 pp, high confidence
Nova Pro16.9%83.9%estimated ± 3.5 pp, high confidence
o338.6%85.7%estimated ± 3.5 pp, high confidence
Phi-414.1%81.8%estimated ± 3.5 pp, high confidence
Quasar 438B15.5%83.0%estimated ± 3.5 pp, high confidence
Qwen3.5 397B (Reasoning)24.5%85.5%estimated ± 3.5 pp, high confidence
Qwen 3.6 Max (preview)37.9%85.7%estimated ± 3.5 pp, high confidence
Qwen3.8 Max Preview31.7%85.7%estimated ± 3.5 pp, high confidence
Qwen3 Max24.4%85.5%estimated ± 3.5 pp, high confidence
Qwen3-Omni-30B-A3B-Instruct14.3%82.0%estimated ± 3.5 pp, high confidence
Qwen3-Omni-30B-A3B-Thinking14.6%82.3%estimated ± 3.5 pp, high confidence
Sarvam 105B17.6%84.2%estimated ± 3.5 pp, high confidence
Sarvam 30B12.6%79.8%estimated ± 3.5 pp, high confidence
Solar Pro 216.1%83.4%estimated ± 3.5 pp, high confidence
Solar Pro 318.5%84.6%estimated ± 3.5 pp, high confidence
Step 3.7 Flash25.8%85.6%estimated ± 3.5 pp, high confidence
Step 5 Preview41.5%85.7%estimated ± 3.5 pp, high confidence
Trinity-Large-Preview22.5%85.4%estimated ± 3.5 pp, high confidence
Trinity-Large-Thinking22.5%85.4%estimated ± 3.5 pp, high confidence
Ultravox v0.6 Llama 3.3 70B19.0%84.7%estimated ± 3.5 pp, high confidence