benchgap
OpenAI · model

GPT-6.1 Sol (xhigh) benchmark scores

As of 2026-10-07, GPT-6.1 Sol (xhigh) (OpenAI) has measured scores on 3 benchmarks and estimated scores on 7 more.

BenchmarkScoreSource
GPQA Diamond94.2%estimated ± 2.3 pp, high confidence
CritPt31.7%measured
GDP.pdf31.8%measured
τ³-Bench Banking47.4%estimated ± 5.0 pp, medium confidence
MMMU-Pro86.7%estimated ± 2.6 pp, high confidence
AA-AnalystAgent53.4%estimated ± 7.8 pp, medium confidence
Harvey LAB93.9%estimated ± 3.2 pp, high confidence
Humanity's Last Exam56.1%estimated ± 4.1 pp, high confidence
AutomationBench66.6%measured
EnterpriseOps-Gym42.3%estimated ± 6.2 pp, low confidence