OpenAI · model
GPT-6 Luna (max) benchmark scores
As of 2026-10-07, GPT-6 Luna (max) (OpenAI) has measured scores on 8 benchmarks and estimated scores on 6 more.
| Benchmark | Score | Source |
|---|---|---|
| Terminal-Bench 2.1 | 84.3% | estimated ± 4.0 pp, high confidence |
| Terminal-Bench 4.0 | 12.6% | measured |
| Terminal-Bench Science 0.1 | 8.6% | measured |
| GPQA Diamond | 91.2% | estimated ± 1.3 pp, high confidence |
| SciCode | 54.6% | measured |
| CritPt | 19.4% | measured |
| GDP.pdf | 22.8% | measured |
| τ³-Bench Banking | 47.4% | estimated ± 5.0 pp, medium confidence |
| MMMU-Pro | 80.0% | measured |
| AA-AnalystAgent | 39.6% | estimated ± 7.8 pp, medium confidence |
| Harvey LAB | 91.8% | estimated ± 3.2 pp, high confidence |
| Humanity's Last Exam | 38.5% | measured |
| AutomationBench | 53.2% | measured |
| EnterpriseOps-Gym | 42.1% | estimated ± 6.2 pp, low confidence |