benchgap
DeepSeek · model

DeepSeek V3.2 (Thinking) benchmark scores

As of 2026-10-07, DeepSeek V3.2 (Thinking) (DeepSeek) has measured scores on 1 benchmark and estimated scores on 8 more.

BenchmarkScoreSource
AA-SciCode46.5%estimated ± 3.8 pp, high confidence
AA Coding Index41.4%estimated ± 6.8 pp, medium confidence
SWE-bench Verified72.6%estimated ± 4.4 pp, high confidence
FrontierCode 1.1 Main1.9%estimated ± 4.3 pp, low confidence
Vals SWE-bench65.0%estimated ± 6.3 pp, medium confidence
PostTrainBench v1.119.9%estimated ± 10.5 pp, low confidence
Vibe Code Bench5.1%measured
React Native Evals38.5%estimated ± 3.0 pp, low confidence
SWE-Rebench0.0%estimated ± 6.5 pp, low confidence