benchgap
DeepSeek · model

DeepSeek R1 Distill Qwen 32B benchmark scores

As of 2026-10-07, DeepSeek R1 Distill Qwen 32B (DeepSeek) has measured scores on 5 benchmarks and estimated scores on 29 more.

BenchmarkScoreSource
AA-GPQA Diamond61.5%measured
AA-HLE4.6%measured
Artificial Analysis Intelligence Index8.4%measured
CritPt0.3%estimated ± 2.3 pp, high confidence
GPQA Diamond69.2%estimated ± 2.2 pp, medium confidence
HLE19.4%estimated ± 9.9 pp, low confidence
IFBench55.3%estimated ± 6.9 pp, low confidence
AA-LCR8.7%measured
AA-Omniscience Accuracy17.4%estimated ± 6.9 pp, medium confidence
IFEval70.3%estimated ± 1.3 pp, medium confidence
LongBench v20.0%estimated ± 5.9 pp, low confidence
BioMysteryBench (human-difficult)35.2%estimated ± 8.1 pp, low confidence
MMLU-Pro75.2%estimated ± 3.9 pp, high confidence
GPQA53.8%estimated ± 3.2 pp, high confidence
AA-IFBench22.9%measured
ARC-AGI-11.0%estimated ± 5.4 pp, low confidence
ARC-AGI-20.4%estimated ± 14.2 pp, low confidence
Vals GPQA Diamond48.1%estimated ± 7.4 pp, low confidence
Vals MMLU-Pro83.1%estimated ± 2.5 pp, medium confidence
HLE w/o tools3.6%estimated ± 5.7 pp, low confidence
C-Eval80.8%estimated ± 1.0 pp, low confidence
MMLU-Redux73.0%estimated ± 2.1 pp, low confidence
SuperGPQA39.1%estimated ± 10.4 pp, low confidence
HealthBench Hard0.1%estimated ± 7.7 pp, low confidence
MMLU-Pro (Arcee)78.0%estimated ± 5.4 pp, low confidence
ARC-AGI-31.6%estimated ± 10.7 pp, low confidence
BioMysteryBench (human-solvable)0.3%estimated ± 1.0 pp, low confidence
HealthBench (raw)7.2%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional57.7%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)21.4%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench241.8%estimated ± 1.8 pp, low confidence
MMLU82.3%estimated ± 1.9 pp, medium confidence