Anthropic · model
Claude Haiku 4.5 Thinking benchmark scores
As of 2026-10-07, Claude Haiku 4.5 Thinking (Anthropic) has measured scores on 1 benchmark and estimated scores on 8 more.
| Benchmark | Score | Source |
|---|---|---|
| AA-SciCode | 47.4% | estimated ± 3.8 pp, high confidence |
| AA Coding Index | 44.9% | estimated ± 6.8 pp, medium confidence |
| SWE-bench Verified | 73.8% | estimated ± 4.4 pp, high confidence |
| FrontierCode 1.1 Main | 4.5% | estimated ± 4.3 pp, low confidence |
| Vals SWE-bench | 68.5% | estimated ± 6.3 pp, medium confidence |
| PostTrainBench v1.1 | 19.9% | estimated ± 10.5 pp, low confidence |
| Vibe Code Bench | 11.4% | measured |
| React Native Evals | 57.0% | estimated ± 3.0 pp, low confidence |
| SWE-Rebench | 2.2% | estimated ± 6.5 pp, low confidence |