OpenAI · model
GPT-6.1 Sol (xhigh) benchmark scores
As of 2026-10-07, GPT-6.1 Sol (xhigh) (OpenAI) has measured scores on 3 benchmarks and estimated scores on 7 more.
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 94.2% | estimated ± 2.3 pp, high confidence |
| CritPt | 31.7% | measured |
| GDP.pdf | 31.8% | measured |
| τ³-Bench Banking | 47.4% | estimated ± 5.0 pp, medium confidence |
| MMMU-Pro | 86.7% | estimated ± 2.6 pp, high confidence |
| AA-AnalystAgent | 53.4% | estimated ± 7.8 pp, medium confidence |
| Harvey LAB | 93.9% | estimated ± 3.2 pp, high confidence |
| Humanity's Last Exam | 56.1% | estimated ± 4.1 pp, high confidence |
| AutomationBench | 66.6% | measured |
| EnterpriseOps-Gym | 42.3% | estimated ± 6.2 pp, low confidence |