Model compare
Two models, head to head
Measured scores decide who leads; estimates fill in the rest, hatched, with their confidence. Pick two models, or one and the frontier model to hold it against: its closest rival in general performance, or for a model below the frontier the closest of the frontier models it shares the most benchmarks with.
| Frontier model | General performance | Measured on |
|---|---|---|
| Claude Mythos 5 | 91.2 | 13 |
| Sakana Fugu-Ultra | 90.0 | 10 |
| Gemini 4 Argon | 87.5 | 20 |
| GPT-6 Astra | 87.1 | 49 |
| Claude Fable 5 | 86.7 | 34 |
| Claude Opus 5.5 | 85.3 | 41 |
| Ember-1 | 84.8 | 3 |
| Claude Sonnet 5.5 | 83.7 | 34 |
| Claude Opus 5 | 82.9 | 56 |
| Claude Fable 5.1 | 82.1 | 42 |
| Qwen3.8 Max Preview | 81.2 | 13 |
| GPT-5.6 Sol | 81.0 | 53 |
| Sakana Fugu | 80.1 | 10 |
| GPT-5.5 Pro | 78.0 | 9 |
| Ling 3.1 Flash | 77.9 | 13 |
| GPT-6.1 Sol | 76.9 | 27 |
| Claude Opus 4.8 | 76.7 | 48 |
| Qwen3.8 Max | 76.1 | 46 |
| Qwen3.7 Max | 75.5 | 50 |
| Muse Spark 1.3 | 75.1 | 27 |
| Ornith-1.0-397B | 74.9 | 6 |
| Atria Dawn Preview | 74.5 | 11 |
| Grok 4.6 | 73.7 | 32 |
| Gemini 3.8 Flash | 73.4 | 39 |
| Step 3.5 Flash | 72.9 | 3 |
| GPT-5.4 Pro | 72.2 | 9 |
| Qwen3.8-Omni-Flash | 72.0 | 15 |
| GPT-5.3 Codex | 71.5 | 19 |
Pick two models on the interactive page, or one and a frontier model is picked for it - of one provider, or of any.