benchgap
OpenAI · model

GPT-6 Sol (max) benchmark scores

As of 2026-10-07, GPT-6 Sol (max) (OpenAI) has measured scores on 9 benchmarks and estimated scores on 6 more.

BenchmarkScoreSource
Terminal-Bench 2.187.8%estimated ± 4.0 pp, high confidence
Terminal-Bench 4.043.9%measured
Terminal-Bench Science 0.130.0%measured
GPQA Diamond93.6%estimated ± 1.3 pp, high confidence
SciCode57.6%measured
CritPt30.9%measured
GDP.pdf25.2%measured
τ³-Bench Banking46.0%estimated ± 2.6 pp, medium confidence
MMMU-Pro83.0%measured
AA-AnalystAgent49.4%estimated ± 7.8 pp, medium confidence
Harvey LAB93.7%estimated ± 1.6 pp, low confidence
Humanity's Last Exam47.9%measured
AutomationBench61.6%measured
EnterpriseOps-Gym42.3%estimated ± 6.2 pp, low confidence
ITBench SRE49.4%measured