Score matrix
Every model × every benchmark
Every model on every benchmark: measured LLM scores and calibrated estimates for the missing ones, side by side.
Models
- Claude 4 Sonnet (1 measured, 0 estimated)
- Claude Fable 5 (with fallback) (7 measured, 4 estimated)
- Claude Fable 5.1 (high with fallback) (3 measured, 7 estimated)
- Claude Fable 5.1 (max with fallback) (13 measured, 2 estimated)
- Claude Fable 5.1 (xhigh with fallback) (6 measured, 6 estimated)
- Claude Opus 4.5 (1 measured, 0 estimated)
- Claude Opus 4.5 (non-reasoning) (1 measured, 0 estimated)
- Claude Opus 4.6 (max) (1 measured, 0 estimated)
- Claude Opus 4.7 (max) (1 measured, 2 estimated)
- Claude Opus 5 (max) (2 measured, 3 estimated)
- Claude Opus 5.5 (high with fallback) (5 measured, 4 estimated)
- Claude Opus 5.5 (max with fallback) (11 measured, 4 estimated)
- Claude Opus 5.5 (medium with fallback) (2 measured, 3 estimated)
- Claude Opus 5.5 (xhigh with fallback) (6 measured, 3 estimated)
- Claude Sonnet 5.5 (high with fallback) (1 measured, 2 estimated)
- Claude Sonnet 5.5 (max with fallback) (9 measured, 5 estimated)
- Claude Sonnet 5.5 (xhigh with fallback) (2 measured, 1 estimated)
- DeepSeek V3.2 Speciale (2 measured, 0 estimated)
- DeepSeek V4 Pro 0813 (max) (1 measured, 3 estimated)
- DeepSeek V4.1 Flash (max) (9 measured, 6 estimated)
- GLM-4.7 (1 measured, 0 estimated)
- GLM-4.7-Flash (1 measured, 0 estimated)
- GLM-5-Turbo (1 measured, 0 estimated)
- GLM-5.2 (max) (3 measured, 2 estimated)
- GLM-5.3 (max) (12 measured, 3 estimated)
- GLM-5.3-Flash (12 measured, 3 estimated)
- GPT-5 (high) (1 measured, 0 estimated)
- GPT-5 (medium) (1 measured, 0 estimated)
- GPT-5 Codex (high) (1 measured, 0 estimated)
- GPT-5.2 (medium) (2 measured, 0 estimated)
- GPT-5.2 (xhigh) (1 measured, 0 estimated)
- GPT-5.5 (xhigh) (3 measured, 4 estimated)
- GPT-5.6 Luna (max) (1 measured, 0 estimated)
- GPT-5.6 Sol (high) (1 measured, 0 estimated)
- GPT-5.6 Sol (max) (5 measured, 5 estimated)
- GPT-5.6 Sol (medium) (1 measured, 0 estimated)
- GPT-5.6 Sol (xhigh) (1 measured, 1 estimated)
- GPT-5.6 Terra (max) (2 measured, 2 estimated)
- GPT-5.6 Terra (xhigh) (1 measured, 0 estimated)
- GPT-6 Astra (high) (5 measured, 7 estimated)
- GPT-6 Astra (max) (13 measured, 2 estimated)
- GPT-6 Astra (medium) (2 measured, 2 estimated)
- GPT-6 Astra (xhigh) (6 measured, 7 estimated)
- GPT-6 Luna (max) (8 measured, 6 estimated)
- GPT-6 Sol (max) (9 measured, 6 estimated)
- GPT-6.1 Sol (high) (1 measured, 1 estimated)
- GPT-6.1 Sol (max) (9 measured, 5 estimated)
- GPT-6.1 Sol (xhigh) (3 measured, 7 estimated)
- Gemini 3 Flash (4 measured, 0 estimated)
- Gemini 3 Flash (non-reasoning) (1 measured, 0 estimated)
- Gemini 3 Pro Preview (high) (3 measured, 0 estimated)
- Gemini 3 Pro Preview (low) (1 measured, 0 estimated)
- Gemini 3.1 Pro Preview (1 measured, 0 estimated)
- Gemini 3.5 Flash (high) (2 measured, 3 estimated)
- Gemini 3.7 Flash (high) (2 measured, 6 estimated)
- Gemini 3.7 Flash (medium) (2 measured, 3 estimated)
- Gemini 3.8 Flash (high) (12 measured, 3 estimated)
- Gemini 4 Argon (high) (6 measured, 8 estimated)
- Grok 3 mini Reasoning (high) (1 measured, 0 estimated)
- Grok 4.20 0309 (1 measured, 0 estimated)
- Grok 4.20 0309 v2 (1 measured, 0 estimated)
- Grok 4.3 (high) (1 measured, 0 estimated)
- Grok 4.3 (low) (1 measured, 0 estimated)
- Grok 4.3 (medium) (1 measured, 0 estimated)
- Grok 4.6 (high) (4 measured, 4 estimated)
- Grok 4.6 (xhigh) (1 measured, 4 estimated)
- Grok 4.7 (xhigh) (7 measured, 8 estimated)
- Inkling (xhigh) (13 measured, 1 estimated)
- JT-35B-Flash (1 measured, 0 estimated)
- K2 Horizon 375B A23B (9 measured, 5 estimated)
- Kimi K3 (max) (15 measured, 1 estimated)
- MiMo-V2.6-Pro (7 measured, 7 estimated)
- MiniMax-M3 (17 measured, 0 estimated)
- Mistral Large 4 Preview (7 measured, 7 estimated)
- Muse Glimmer (high) (11 measured, 3 estimated)
- Muse Spark 1.3 (max) (11 measured, 4 estimated)
- Muse Spark 1.3 (xhigh) (2 measured, 3 estimated)
- Nemotron 3 Ultra (15 measured, 2 estimated)
- Qwen3.8 2.4T A95B (1 measured, 4 estimated)
- Qwen3.8 27B (medium) (1 measured, 4 estimated)
- Qwen3.8 27B (xhigh) (12 measured, 2 estimated)
- Qwen3.8 Max (1 measured, 4 estimated)
- Qwen3.8 Max (0902) (16 measured, 0 estimated)
- Step 5 Preview (13 measured, 3 estimated)
- o3 (1 measured, 0 estimated)
Benchmarks
Agentic · terminal
Agentic · tools
- AA-AnalystAgent
- APEX-Agents
- AutomationBench
- EnterpriseOps-Gym
- Harvey LAB
- ITBench SRE
- τ³-Bench Banking
- τ²-Bench Telecom