benchgap
OpenAI · model

GPT-5.6 Sol (max) benchmark scores

As of 2026-10-07, GPT-5.6 Sol (max) (OpenAI) has measured scores on 5 benchmarks and estimated scores on 5 more.

BenchmarkScoreSource
Terminal-Bench Hard65.9%measured
GPQA Diamond94.1%measured
CritPt32.3%measured
τ³-Bench Banking44.2%estimated ± 2.6 pp, low confidence
AA-AnalystAgent47.5%measured
Harvey LAB93.9%estimated ± 1.6 pp, low confidence
Humanity's Last Exam54.5%estimated ± 3.4 pp, high confidence
AutomationBench57.2%estimated ± 11.2 pp, low confidence
EnterpriseOps-Gym50.4%estimated ± 1.8 pp, medium confidence
ITBench SRE56.2%measured