benchgap
Google · model

Gemini 3 Pro Deep Think benchmark scores

As of 2026-10-07, Gemini 3 Pro Deep Think (Google) has measured scores on 2 benchmarks and estimated scores on 21 more.

BenchmarkScoreSource
AA-GPQA Diamond89.1%estimated ± 2.7 pp, high confidence
AA-HLE48.8%estimated ± 4.6 pp, high confidence
Artificial Analysis Intelligence Index46.8%estimated ± 5.1 pp, medium confidence
CritPt25.7%measured
GPQA Diamond89.4%estimated ± 1.5 pp, high confidence
HLE44.2%estimated ± 6.3 pp, medium confidence
AA-Omniscience Accuracy43.4%estimated ± 6.2 pp, medium confidence
BioMysteryBench (human-difficult)49.1%estimated ± 6.8 pp, low confidence
MMLU-Pro88.5%estimated ± 5.8 pp, low confidence
GPQA89.9%estimated ± 1.6 pp, high confidence
ARC-AGI-183.3%estimated ± 1.9 pp, high confidence
ARC-AGI-245.1%measured
Vals GPQA Diamond88.8%estimated ± 2.2 pp, high confidence
Vals MMLU-Pro85.7%estimated ± 2.0 pp, high confidence
HLE w/o tools36.3%estimated ± 6.9 pp, medium confidence
C-Eval100.0%estimated ± 1.4 pp, low confidence
HealthBench Hard33.7%estimated ± 9.0 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
ARC-AGI-30.0%estimated ± 13.8 pp, low confidence
BioMysteryBench (human-solvable)89.5%estimated ± 0.6 pp, medium confidence
HealthBench Professional55.1%estimated ± 4.6 pp, medium confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench280.7%estimated ± 2.5 pp, low confidence