benchgap
OpenAI · model

o3-pro benchmark scores

As of 2026-10-07, o3-pro (OpenAI) has measured scores on 2 benchmarks and estimated scores on 26 more.

BenchmarkScoreSource
AA-GPQA Diamond84.5%measured
AA-HLE23.8%estimated ± 4.0 pp, high confidence
Artificial Analysis Intelligence Index21.9%measured
CritPt1.6%estimated ± 2.9 pp, high confidence
GPQA Diamond86.0%estimated ± 2.4 pp, high confidence
HLE27.5%estimated ± 9.9 pp, medium confidence
AA-Omniscience Accuracy27.9%estimated ± 7.5 pp, medium confidence
BioMysteryBench (human-difficult)35.2%estimated ± 8.1 pp, low confidence
MMLU-Pro85.2%estimated ± 3.9 pp, high confidence
GPQA86.7%estimated ± 3.2 pp, high confidence
ARC-AGI-160.6%estimated ± 5.4 pp, medium confidence
ARC-AGI-222.7%estimated ± 14.2 pp, low confidence
Vals GPQA Diamond79.3%estimated ± 7.4 pp, medium confidence
Vals MMLU-Pro84.2%estimated ± 2.6 pp, medium confidence
HLE w/o tools27.5%estimated ± 6.3 pp, medium confidence
C-Eval91.9%estimated ± 1.0 pp, medium confidence
MMLU-Redux94.6%estimated ± 2.1 pp, medium confidence
SuperGPQA69.5%estimated ± 10.4 pp, low confidence
HealthBench Hard11.2%estimated ± 7.7 pp, low confidence
MMLU-Pro (Arcee)84.3%estimated ± 5.4 pp, low confidence
BioMysteryBench (human-solvable)49.1%estimated ± 1.0 pp, low confidence
HealthBench (raw)30.9%estimated ± 6.7 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional55.1%estimated ± 4.9 pp, low confidence
HealthBench Professional (raw)40.8%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench268.9%estimated ± 1.8 pp, low confidence
MMLU96.0%estimated ± 1.9 pp, low confidence