OpenAI · model
GPT-5.6 Sol (max) benchmark scores
As of 2026-10-07, GPT-5.6 Sol (max) (OpenAI) has measured scores on 5 benchmarks and estimated scores on 5 more.
| Benchmark | Score | Source |
|---|---|---|
| Terminal-Bench Hard | 65.9% | measured |
| GPQA Diamond | 94.1% | measured |
| CritPt | 32.3% | measured |
| τ³-Bench Banking | 44.2% | estimated ± 2.6 pp, low confidence |
| AA-AnalystAgent | 47.5% | measured |
| Harvey LAB | 93.9% | estimated ± 1.6 pp, low confidence |
| Humanity's Last Exam | 54.5% | estimated ± 3.4 pp, high confidence |
| AutomationBench | 57.2% | estimated ± 11.2 pp, low confidence |
| EnterpriseOps-Gym | 50.4% | estimated ± 1.8 pp, medium confidence |
| ITBench SRE | 56.2% | measured |