benchgap
OpenAI · model

GPT-6 Luna (max) benchmark scores

As of 2026-10-07, GPT-6 Luna (max) (OpenAI) has measured scores on 8 benchmarks and estimated scores on 6 more.

BenchmarkScoreSource
Terminal-Bench 2.184.3%estimated ± 4.0 pp, high confidence
Terminal-Bench 4.012.6%measured
Terminal-Bench Science 0.18.6%measured
GPQA Diamond91.2%estimated ± 1.3 pp, high confidence
SciCode54.6%measured
CritPt19.4%measured
GDP.pdf22.8%measured
τ³-Bench Banking47.4%estimated ± 5.0 pp, medium confidence
MMMU-Pro80.0%measured
AA-AnalystAgent39.6%estimated ± 7.8 pp, medium confidence
Harvey LAB91.8%estimated ± 3.2 pp, high confidence
Humanity's Last Exam38.5%measured
AutomationBench53.2%measured
EnterpriseOps-Gym42.1%estimated ± 6.2 pp, low confidence