benchgap
OpenAI · model

GPT-6.1 Sol (max) benchmark scores

As of 2026-10-07, GPT-6.1 Sol (max) (OpenAI) has measured scores on 9 benchmarks and estimated scores on 5 more.

BenchmarkScoreSource
Terminal-Bench 2.188.2%estimated ± 4.0 pp, high confidence
Terminal-Bench 4.056.1%measured
Terminal-Bench Science 0.158.1%measured
GPQA Diamond94.6%estimated ± 1.3 pp, high confidence
SciCode54.2%measured
CritPt31.7%measured
GDP.pdf31.0%measured
τ³-Bench Banking45.8%estimated ± 4.5 pp, medium confidence
MMMU-Pro86.0%measured
AA-AnalystAgent50.0%measured
Harvey LAB93.4%estimated ± 1.8 pp, high confidence
Humanity's Last Exam52.9%measured
AutomationBench64.9%measured
EnterpriseOps-Gym51.6%estimated ± 1.8 pp, low confidence