benchgap
Calibration

Artificial Analysis Intelligence Index → C-Eval

C-Eval is estimated from Artificial Analysis Intelligence Index with a Michaelis–Menten curve fitted on 5 models measured on both: y = 1.0056·x / (0.02053 + x), R² = 0.73, cross-validated error 1.0 pp. It is used for 123 estimates.

Estimated modelArtificial Analysis Intelligence IndexC-EvalSource
Apodex 1.1 Mini26.4%93.3%estimated ± 1.0 pp, medium confidence
Celeris-16.4%76.0%estimated ± 1.0 pp, low confidence
Claude 3 Haiku5.6%73.5%estimated ± 1.0 pp, low confidence
Claude 3 Opus8.7%81.4%estimated ± 1.0 pp, low confidence
Claude 4.1 Opus18.6%90.5%estimated ± 1.0 pp, medium confidence
Claude 4.1 Opus Thinking22.9%92.3%estimated ± 1.0 pp, medium confidence
Claude 4 Sonnet16.6%89.5%estimated ± 1.0 pp, low confidence
Claude Fable 549.6%96.6%estimated ± 1.0 pp, low confidence
Claude Haiku 5.543.4%96.0%estimated ± 1.0 pp, low confidence
Claude Opus 4.5 Thinking29.1%93.9%estimated ± 1.0 pp, low confidence
Claude Opus 4.6 (Adaptive)32.0%94.5%estimated ± 1.0 pp, low confidence
Claude Opus 4.730.9%94.3%estimated ± 1.0 pp, low confidence
Claude Opus 5.557.6%97.1%estimated ± 1.0 pp, low confidence
Claude Sonnet 5.556.0%97.0%estimated ± 1.0 pp, low confidence
Command A+22.5%92.2%estimated ± 1.0 pp, medium confidence
DeepSeek-R113.1%87.0%estimated ± 1.0 pp, low confidence
DeepSeek R1 Distill Qwen 32B8.4%80.8%estimated ± 1.0 pp, low confidence
DeepSeek V3 03249.7%83.0%estimated ± 1.0 pp, low confidence
DeepSeek V3.113.7%87.5%estimated ± 1.0 pp, low confidence
DeepSeek V3.1 (Reasoning)13.5%87.3%estimated ± 1.0 pp, low confidence
DeepSeek V3.216.0%89.1%estimated ± 1.0 pp, low confidence
Exaone 4.0 1.2B5.2%72.2%estimated ± 1.0 pp, low confidence
Exaone 4.0 32B6.3%75.9%estimated ± 1.0 pp, low confidence
Gemini 1.0 Pro5.3%72.6%estimated ± 1.0 pp, low confidence
Gemini 1.5 Pro7.9%79.8%estimated ± 1.0 pp, low confidence
Gemini 2.5 Flash9.9%83.2%estimated ± 1.0 pp, low confidence
Gemini 3.1 Pro29.7%94.1%estimated ± 1.0 pp, low confidence
Gemini 3.5 Flash-Lite22.2%92.0%estimated ± 1.0 pp, medium confidence
Gemini 3.6 Flash34.0%94.8%estimated ± 1.0 pp, low confidence
Gemini 3.7 Flash39.1%95.5%estimated ± 1.0 pp, low confidence
Gemini 3.8 Flash40.9%95.8%estimated ± 1.0 pp, low confidence
Gemini 3 Flash17.9%90.2%estimated ± 1.0 pp, low confidence
Gemini 3 Pro28.0%93.7%estimated ± 1.0 pp, low confidence
Gemini 4 Argon52.6%96.8%estimated ± 1.0 pp, low confidence
Gemma 3 27B4.9%70.7%estimated ± 1.0 pp, low confidence
GLM-4.5-Air11.1%84.9%estimated ± 1.0 pp, low confidence
GLM-4.614.9%88.4%estimated ± 1.0 pp, low confidence
GLM-5.344.8%96.1%estimated ± 1.0 pp, low confidence
GLM-5.3-Flash41.8%95.8%estimated ± 1.0 pp, low confidence
GLM-5-Turbo26.6%93.4%estimated ± 1.0 pp, medium confidence
GLM-5V-Turbo23.5%92.5%estimated ± 1.0 pp, medium confidence
GPT-4 Turbo7.0%77.9%estimated ± 1.0 pp, low confidence
GPT-4o8.4%80.9%estimated ± 1.0 pp, low confidence
GPT-4o mini6.7%76.9%estimated ± 1.0 pp, low confidence
GPT-5.124.7%92.9%estimated ± 1.0 pp, medium confidence
GPT-5.1-Codex23.7%92.5%estimated ± 1.0 pp, medium confidence
GPT-5.1-Codex-Max23.7%92.5%estimated ± 1.0 pp, medium confidence
GPT-5.2-Codex28.5%93.8%estimated ± 1.0 pp, low confidence
GPT-5.3 Codex32.5%94.6%estimated ± 1.0 pp, low confidence
GPT-5 (high)23.0%92.3%estimated ± 1.0 pp, medium confidence
GPT-5 (medium)22.9%92.3%estimated ± 1.0 pp, medium confidence
GPT-6.1 Sol51.8%96.7%estimated ± 1.0 pp, low confidence
GPT-6 Luna38.1%95.4%estimated ± 1.0 pp, low confidence
GPT-6 Sol47.6%96.4%estimated ± 1.0 pp, low confidence
GPT-OSS 120B11.6%85.4%estimated ± 1.0 pp, low confidence
GPT-OSS 20B9.0%81.8%estimated ± 1.0 pp, low confidence
Granite-4.0-350M4.8%70.6%estimated ± 1.0 pp, low confidence
Granite-4.0-H-1B5.2%72.1%estimated ± 1.0 pp, low confidence
Granite-4.0-H-350M4.8%70.6%estimated ± 1.0 pp, low confidence
Grok 422.5%92.1%estimated ± 1.0 pp, medium confidence
Grok 4.1 Fast11.3%85.1%estimated ± 1.0 pp, low confidence
Grok 4.1 Fast (Reasoning)20.4%91.3%estimated ± 1.0 pp, medium confidence
Grok 4.538.8%95.5%estimated ± 1.0 pp, low confidence
Grok 4.644.3%96.1%estimated ± 1.0 pp, low confidence
Grok 4.746.5%96.3%estimated ± 1.0 pp, low confidence
Grok 4 Fast (Reasoning)17.9%90.2%estimated ± 1.0 pp, low confidence
Grok Code Fast 114.1%87.7%estimated ± 1.0 pp, low confidence
Hy325.3%93.0%estimated ± 1.0 pp, medium confidence
K-Exaone14.4%88.0%estimated ± 1.0 pp, low confidence
Kimi K212.7%86.6%estimated ± 1.0 pp, low confidence
Kimi K2.7 Code25.8%93.1%estimated ± 1.0 pp, medium confidence
LFM2.5-2.6B8.4%80.8%estimated ± 1.0 pp, low confidence
LFM2.5-8B-A1B7.2%78.3%estimated ± 1.0 pp, low confidence
LFM2.5-VL-1.6B-Extract4.8%70.6%estimated ± 1.0 pp, low confidence
Ling 3.0 Flash VL24.6%92.8%estimated ± 1.0 pp, medium confidence
Ling 3.0 Tiny11.1%84.8%estimated ± 1.0 pp, low confidence
Ling 3.1 Flash41.1%95.8%estimated ± 1.0 pp, low confidence
Llama 3.1 405B7.3%78.5%estimated ± 1.0 pp, low confidence
Llama 4 Maverick10.0%83.4%estimated ± 1.0 pp, low confidence
Llama 4 Scout8.1%80.2%estimated ± 1.0 pp, low confidence
Mercury 2.512.3%86.2%estimated ± 1.0 pp, low confidence
MiMo-V2.6-Flash37.9%95.4%estimated ± 1.0 pp, low confidence
MiMo-V2.6-Pro46.3%96.3%estimated ± 1.0 pp, low confidence
MiMo-V2-Omni23.9%92.6%estimated ± 1.0 pp, medium confidence
MiMo-V2-Pro28.6%93.8%estimated ± 1.0 pp, low confidence
MiniMax M2.722.8%92.2%estimated ± 1.0 pp, medium confidence
MiniMax M329.2%94.0%estimated ± 1.0 pp, low confidence
Mistral Large 27.6%79.1%estimated ± 1.0 pp, low confidence
Mistral Large 39.3%82.3%estimated ± 1.0 pp, low confidence
Mistral Large 438.4%95.5%estimated ± 1.0 pp, low confidence
Mistral Medium 39.0%82.0%estimated ± 1.0 pp, low confidence
Mistral Medium 3.5 128B14.2%87.8%estimated ± 1.0 pp, low confidence
Mistral Small 411.3%85.1%estimated ± 1.0 pp, low confidence
Mistral Small 4 (Reasoning)11.3%85.1%estimated ± 1.0 pp, low confidence
Muse Glimmer 30B17.5%90.0%estimated ± 1.0 pp, low confidence
Muse Spark 1.239.6%95.6%estimated ± 1.0 pp, low confidence
Muse Spark 1.348.1%96.4%estimated ± 1.0 pp, low confidence
Nemotron 3 Nano 30B8.9%81.7%estimated ± 1.0 pp, low confidence
Nemotron 3 Super 100B12.8%86.7%estimated ± 1.0 pp, low confidence
Nemotron Ultra 253B7.5%79.0%estimated ± 1.0 pp, low confidence
North Mini Code9.9%83.3%estimated ± 1.0 pp, low confidence
Nova Pro7.0%77.7%estimated ± 1.0 pp, low confidence
o1-preview11.4%85.2%estimated ± 1.0 pp, low confidence
o320.2%91.3%estimated ± 1.0 pp, medium confidence
o3-pro21.9%91.9%estimated ± 1.0 pp, medium confidence
Phi-45.9%74.7%estimated ± 1.0 pp, low confidence
Phi-4 Multimodal Instruct5.8%74.3%estimated ± 1.0 pp, low confidence
Quasar 438B26.7%93.4%estimated ± 1.0 pp, medium confidence
Qwen2.5 Coder 32B Instruct6.7%77.1%estimated ± 1.0 pp, low confidence
Qwen3.5 397B (Reasoning)21.4%91.8%estimated ± 1.0 pp, medium confidence
Qwen3.8 Max Preview45.4%96.2%estimated ± 1.0 pp, low confidence
Qwen3 Max15.6%88.9%estimated ± 1.0 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct6.0%75.0%estimated ± 1.0 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking7.8%79.5%estimated ± 1.0 pp, low confidence
Sarvam 105B8.8%81.5%estimated ± 1.0 pp, low confidence
Sarvam 30B6.6%76.6%estimated ± 1.0 pp, low confidence
Solar Pro 27.0%77.8%estimated ± 1.0 pp, low confidence
Solar Pro 37.8%79.6%estimated ± 1.0 pp, low confidence
Solar Pro 428.2%93.7%estimated ± 1.0 pp, low confidence
Step 3.7 Flash19.5%91.0%estimated ± 1.0 pp, medium confidence
Trinity-Large-Preview10.8%84.5%estimated ± 1.0 pp, low confidence
Trinity-Large-Thinking10.8%84.5%estimated ± 1.0 pp, low confidence
Ultravox v0.6 Llama 3.3 70B7.7%79.3%estimated ± 1.0 pp, low confidence