DeepSeek · model
DeepSeek V3.2 (Thinking) benchmark scores
As of 2026-10-07, DeepSeek V3.2 (Thinking) (DeepSeek) has measured scores on 1 benchmark and estimated scores on 8 more.
| Benchmark | Score | Source |
|---|---|---|
| AA-SciCode | 46.5% | estimated ± 3.8 pp, high confidence |
| AA Coding Index | 41.4% | estimated ± 6.8 pp, medium confidence |
| SWE-bench Verified | 72.6% | estimated ± 4.4 pp, high confidence |
| FrontierCode 1.1 Main | 1.9% | estimated ± 4.3 pp, low confidence |
| Vals SWE-bench | 65.0% | estimated ± 6.3 pp, medium confidence |
| PostTrainBench v1.1 | 19.9% | estimated ± 10.5 pp, low confidence |
| Vibe Code Bench | 5.1% | measured |
| React Native Evals | 38.5% | estimated ± 3.0 pp, low confidence |
| SWE-Rebench | 0.0% | estimated ± 6.5 pp, low confidence |