benchgap
Calibration

AA-Omniscience Accuracy → ARC-AGI-3

ARC-AGI-3 is estimated from AA-Omniscience Accuracy with a offset logistic curve fitted on 16 models measured on both: y = 0.0168 + (0.7350 − 0.0168) / (1 + exp(−123.71·(x − 0.6128))), R² = 0.98, cross-validated error 3.4 pp. It is used for 115 estimates.

Estimated modelAA-Omniscience AccuracyARC-AGI-3Source
A.X K218.6%1.7%estimated ± 3.4 pp, medium confidence
Apodex 1.131.7%1.7%estimated ± 3.4 pp, medium confidence
Apodex 1.1 Mini31.7%1.7%estimated ± 3.4 pp, medium confidence
Celeris-111.0%1.7%estimated ± 3.4 pp, medium confidence
Claude 3 Haiku17.6%1.7%estimated ± 3.4 pp, medium confidence
Claude 4 Sonnet22.7%1.7%estimated ± 3.4 pp, medium confidence
Claude Fable 565.4%73.1%estimated ± 3.4 pp, medium confidence
Claude Fable 5.167.2%73.5%estimated ± 3.4 pp, medium confidence
Claude Opus 4.5 Thinking46.6%1.7%estimated ± 3.4 pp, high confidence
Claude Opus 4.6 (Adaptive)47.0%1.7%estimated ± 3.4 pp, high confidence
Claude Opus 4.744.7%1.7%estimated ± 3.4 pp, high confidence
Claude Opus 5.566.2%73.3%estimated ± 3.4 pp, medium confidence
Claude Sonnet 540.1%1.7%estimated ± 3.4 pp, medium confidence
Claude Sonnet 5.554.0%1.7%estimated ± 3.4 pp, high confidence
Command A+8.9%1.7%estimated ± 3.4 pp, medium confidence
DeepSeek-R130.5%1.7%estimated ± 3.4 pp, medium confidence
DeepSeek V3 032424.3%1.7%estimated ± 3.4 pp, medium confidence
DeepSeek V3.123.1%1.7%estimated ± 3.4 pp, medium confidence
DeepSeek V3.1 (Reasoning)29.0%1.7%estimated ± 3.4 pp, medium confidence
DeepSeek V3.224.0%1.7%estimated ± 3.4 pp, medium confidence
Exaone 4.0 1.2B5.0%1.7%estimated ± 3.4 pp, medium confidence
Exaone 4.0 32B10.6%1.7%estimated ± 3.4 pp, medium confidence
Gemini 2.5 Flash26.1%1.7%estimated ± 3.4 pp, medium confidence
Gemini 3.5 Flash-Lite29.5%1.7%estimated ± 3.4 pp, medium confidence
Gemini 3.6 Flash50.0%1.7%estimated ± 3.4 pp, high confidence
Gemini 3.7 Flash55.3%1.7%estimated ± 3.4 pp, high confidence
Gemini 3 Flash45.8%1.7%estimated ± 3.4 pp, high confidence
Gemini 3 Pro55.8%1.8%estimated ± 3.4 pp, high confidence
Gemini 4 Argon49.9%1.7%estimated ± 3.4 pp, high confidence
Gemma 3 27B13.0%1.7%estimated ± 3.4 pp, medium confidence
Gemma 4 26B A4B19.1%1.7%estimated ± 3.4 pp, medium confidence
GLM-4.5-Air16.3%1.7%estimated ± 3.4 pp, medium confidence
GLM-4.621.4%1.7%estimated ± 3.4 pp, medium confidence
GLM-5.123.7%1.7%estimated ± 3.4 pp, medium confidence
GLM-5.333.9%1.7%estimated ± 3.4 pp, medium confidence
GLM-5-Turbo28.4%1.7%estimated ± 3.4 pp, medium confidence
GLM-5V-Turbo29.3%1.7%estimated ± 3.4 pp, medium confidence
GPT-4o19.9%1.7%estimated ± 3.4 pp, medium confidence
GPT-5.137.7%1.7%estimated ± 3.4 pp, medium confidence
GPT-5.1-Codex39.9%1.7%estimated ± 3.4 pp, medium confidence
GPT-5.1-Codex-Max39.9%1.7%estimated ± 3.4 pp, medium confidence
GPT-5.2-Codex41.1%1.7%estimated ± 3.4 pp, medium confidence
GPT-5.3 Codex52.9%1.7%estimated ± 3.4 pp, high confidence
GPT-5 (high)40.3%1.7%estimated ± 3.4 pp, medium confidence
GPT-5 (medium)39.5%1.7%estimated ± 3.4 pp, medium confidence
GPT-OSS 120B21.8%1.7%estimated ± 3.4 pp, medium confidence
GPT-OSS 20B16.0%1.7%estimated ± 3.4 pp, medium confidence
Granite-4.0-350M3.9%1.7%estimated ± 3.4 pp, medium confidence
Granite-4.0-H-1B5.2%1.7%estimated ± 3.4 pp, medium confidence
Granite-4.0-H-350M3.8%1.7%estimated ± 3.4 pp, medium confidence
Grok 440.5%1.7%estimated ± 3.4 pp, medium confidence
Grok 4.1 Fast17.2%1.7%estimated ± 3.4 pp, medium confidence
Grok 4.1 Fast (Reasoning)25.1%1.7%estimated ± 3.4 pp, medium confidence
Grok 4.747.4%1.7%estimated ± 3.4 pp, high confidence
Grok 4 Fast (Reasoning)22.8%1.7%estimated ± 3.4 pp, medium confidence
Grok Code Fast 123.5%1.7%estimated ± 3.4 pp, medium confidence
Hy332.0%1.7%estimated ± 3.4 pp, medium confidence
K-Exaone16.4%1.7%estimated ± 3.4 pp, medium confidence
K-EXAONE 2.013.1%1.7%estimated ± 3.4 pp, medium confidence
Kimi K227.4%1.7%estimated ± 3.4 pp, medium confidence
Kimi K2.7 Code39.6%1.7%estimated ± 3.4 pp, medium confidence
LFM2.5-2.6B4.4%1.7%estimated ± 3.4 pp, medium confidence
LFM2.5-8B-A1B9.4%1.7%estimated ± 3.4 pp, medium confidence
LFM2.5-VL-1.6B-Extract5.8%1.7%estimated ± 3.4 pp, medium confidence
Ling 3.0 Flash VL14.4%1.7%estimated ± 3.4 pp, medium confidence
Ling 3.0 Tiny8.5%1.7%estimated ± 3.4 pp, medium confidence
Ling 3.1 Flash29.1%1.7%estimated ± 3.4 pp, medium confidence
Llama 3.1 405B23.2%1.7%estimated ± 3.4 pp, medium confidence
Llama 4 Maverick24.9%1.7%estimated ± 3.4 pp, medium confidence
Llama 4 Scout15.2%1.7%estimated ± 3.4 pp, medium confidence
Mercury 2.522.0%1.7%estimated ± 3.4 pp, medium confidence
MiMo-V2.5-Pro22.4%1.7%estimated ± 3.4 pp, medium confidence
MiMo-V2.6-Flash27.0%1.7%estimated ± 3.4 pp, medium confidence
MiMo-V2.6-Pro34.8%1.7%estimated ± 3.4 pp, medium confidence
MiMo-V2-Omni19.3%1.7%estimated ± 3.4 pp, medium confidence
MiMo-V2-Pro26.6%1.7%estimated ± 3.4 pp, medium confidence
MiniCPM5-2B8.4%1.7%estimated ± 3.4 pp, medium confidence
MiniMax M2.726.8%1.7%estimated ± 3.4 pp, medium confidence
MiniMax M316.7%1.7%estimated ± 3.4 pp, medium confidence
Mistral Large 219.9%1.7%estimated ± 3.4 pp, medium confidence
Mistral Large 325.0%1.7%estimated ± 3.4 pp, medium confidence
Mistral Large 425.8%1.7%estimated ± 3.4 pp, medium confidence
Mistral Medium 318.3%1.7%estimated ± 3.4 pp, medium confidence
Mistral Medium 3.5 128B24.7%1.7%estimated ± 3.4 pp, medium confidence
Mistral Small 421.7%1.7%estimated ± 3.4 pp, medium confidence
Mistral Small 4 (Reasoning)21.7%1.7%estimated ± 3.4 pp, medium confidence
Muse Glimmer 30B27.0%1.7%estimated ± 3.4 pp, medium confidence
Muse Spark49.6%1.7%estimated ± 3.4 pp, high confidence
Muse Spark 1.152.1%1.7%estimated ± 3.4 pp, high confidence
Muse Spark 1.245.4%1.7%estimated ± 3.4 pp, high confidence
Muse Spark 1.343.6%1.7%estimated ± 3.4 pp, high confidence
Nemotron 3 Nano 30B17.3%1.7%estimated ± 3.4 pp, medium confidence
Nemotron 3 Super 100B24.3%1.7%estimated ± 3.4 pp, medium confidence
Nemotron Ultra 253B20.1%1.7%estimated ± 3.4 pp, medium confidence
North Mini Code18.9%1.7%estimated ± 3.4 pp, medium confidence
Nova Pro16.9%1.7%estimated ± 3.4 pp, medium confidence
o338.6%1.7%estimated ± 3.4 pp, medium confidence
Phi-414.1%1.7%estimated ± 3.4 pp, medium confidence
Quasar 438B15.5%1.7%estimated ± 3.4 pp, medium confidence
Qwen3.5 397B (Reasoning)24.5%1.7%estimated ± 3.4 pp, medium confidence
Qwen 3.6 Max (preview)37.9%1.7%estimated ± 3.4 pp, medium confidence
Qwen3.8 Max Preview31.7%1.7%estimated ± 3.4 pp, medium confidence
Qwen3 Max24.4%1.7%estimated ± 3.4 pp, medium confidence
Qwen3-Omni-30B-A3B-Instruct14.3%1.7%estimated ± 3.4 pp, medium confidence
Qwen3-Omni-30B-A3B-Thinking14.6%1.7%estimated ± 3.4 pp, medium confidence
Sarvam 105B17.6%1.7%estimated ± 3.4 pp, medium confidence
Sarvam 30B12.6%1.7%estimated ± 3.4 pp, medium confidence
Solar Pro 216.1%1.7%estimated ± 3.4 pp, medium confidence
Solar Pro 318.5%1.7%estimated ± 3.4 pp, medium confidence
Solar Pro 418.9%1.7%estimated ± 3.4 pp, medium confidence
Step 3.7 Flash25.8%1.7%estimated ± 3.4 pp, medium confidence
Step 5 Preview41.5%1.7%estimated ± 3.4 pp, medium confidence
Trinity-Large-Preview22.5%1.7%estimated ± 3.4 pp, medium confidence
Trinity-Large-Thinking22.5%1.7%estimated ± 3.4 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B19.0%1.7%estimated ± 3.4 pp, medium confidence