benchgap
OpenAI · model

o1-pro benchmark scores

As of 2026-10-07, o1-pro (OpenAI) has measured scores on 2 benchmarks and estimated scores on 25 more.

BenchmarkScoreSource
AA-GPQA Diamond77.5%estimated ± 5.4 pp, medium confidence
AA-HLE10.2%estimated ± 4.4 pp, high confidence
Artificial Analysis Intelligence Index12.4%measured
CritPt0.4%estimated ± 3.7 pp, high confidence
GPQA Diamond79.0%estimated ± 0.4 pp, high confidence
HLE18.6%estimated ± 8.2 pp, low confidence
AA-Omniscience Accuracy19.5%estimated ± 7.5 pp, medium confidence
BioMysteryBench (human-difficult)35.2%estimated ± 8.1 pp, low confidence
MMLU-Pro83.0%estimated ± 3.0 pp, high confidence
GPQA79.0%measured
ARC-AGI-17.5%estimated ± 5.4 pp, low confidence
ARC-AGI-20.0%estimated ± 11.6 pp, low confidence
Vals GPQA Diamond68.8%estimated ± 4.8 pp, medium confidence
Vals MMLU-Pro65.7%estimated ± 1.2 pp, medium confidence
HLE w/o tools11.4%estimated ± 4.9 pp, high confidence
C-Eval86.7%estimated ± 0.9 pp, low confidence
MMLU-Redux92.1%estimated ± 2.1 pp, high confidence
SuperGPQA58.5%estimated ± 9.2 pp, low confidence
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
MMLU-Pro (Arcee)79.9%estimated ± 5.4 pp, low confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
BioMysteryBench (human-solvable)3.5%estimated ± 1.0 pp, low confidence
HealthBench (raw)18.8%estimated ± 6.7 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional29.7%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)28.8%estimated ± 5.0 pp, low confidence
MMLU89.0%estimated ± 1.9 pp, medium confidence