benchgap
OpenAI · model

GPT-6 Luna benchmark scores

As of 2026-10-07, GPT-6 Luna (OpenAI) has measured scores on 22 benchmarks and estimated scores on 77 more.

BenchmarkScoreSource
AA-GPQA Diamond89.3%estimated ± 1.2 pp, medium confidence
AA-HLE38.5%measured
AA-SciCode54.6%measured
Artificial Analysis Intelligence Index38.1%measured
CritPt19.4%measured
GDPval-AA46.9%measured
GPQA Diamond94.9%estimated ± 1.0 pp, medium confidence
HLE46.0%estimated ± 5.5 pp, low confidence
AA-LCR83.3%measured
LiveCodeBench v695.5%estimated ± 3.3 pp, medium confidence
AA-Omniscience Accuracy43.8%measured
SciCode54.2%estimated ± 3.2 pp, high confidence
Terminal-Bench 2.184.0%estimated ± 3.9 pp, high confidence
BrowseComp84.0%estimated ± 1.8 pp, medium confidence
HLE w/ tools55.7%estimated ± 6.1 pp, medium confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
AA Coding Index75.1%estimated ± 2.5 pp, high confidence
AA-MMMU-Pro79.7%measured
APEX-Agents39.1%estimated ± 10.0 pp, low confidence
APEX-Agents-AA35.1%estimated ± 9.1 pp, medium confidence
BioMysteryBench (human-difficult)42.8%estimated ± 5.9 pp, low confidence
DeepSearchQA90.2%estimated ± 8.4 pp, medium confidence
SWE-bench Verified88.4%estimated ± 5.2 pp, low confidence
AutomationBench42.6%estimated ± 7.6 pp, low confidence
BFCL v477.2%estimated ± 2.6 pp, medium confidence
CyberGym79.5%estimated ± 6.2 pp, medium confidence
JobBench48.1%estimated ± 8.6 pp, medium confidence
SWE-bench Pro66.3%estimated ± 6.8 pp, medium confidence
WideResearch97.1%estimated ± 9.1 pp, low confidence
AA Agentic Index42.1%estimated ± 2.8 pp, high confidence
MMLU-Pro85.7%estimated ± 3.5 pp, high confidence
GPQA93.1%estimated ± 0.8 pp, medium confidence
Gert Labs70.0%estimated ± 8.9 pp, medium confidence
AA-AnalystAgent50.1%estimated ± 6.5 pp, medium confidence
AA EnterpriseOps-Gym47.9%estimated ± 4.4 pp, medium confidence
ARC-AGI-186.7%measured
ARC-AGI-259.3%measured
CursorBench 3.155.3%estimated ± 3.4 pp, medium confidence
CursorBench 3.261.0%estimated ± 5.3 pp, medium confidence
FrontierCode 1.1 Main45.3%estimated ± 3.7 pp, low confidence
FrontierSWE v225.8%estimated ± 14.6 pp, low confidence
OfficeQA Pro54.1%estimated ± 7.2 pp, medium confidence
OSWorld-Verified80.2%estimated ± 8.7 pp, low confidence
Vals GPQA Diamond96.2%estimated ± 1.5 pp, low confidence
Vals MMLU-Pro90.3%estimated ± 1.5 pp, low confidence
Vals SWE-bench85.0%estimated ± 7.6 pp, low confidence
Terminal-Bench 2.1 (Vals AI)70.3%estimated ± 8.5 pp, medium confidence
AA AutomationBench53.2%measured
GDP.pdf22.8%measured
AA ITBench46.3%estimated ± 3.9 pp, low confidence
AA Tau3 Banking47.2%estimated ± 5.8 pp, medium confidence
AA Terminal-Bench 2.185.1%estimated ± 2.6 pp, high confidence
AA Terminal-Bench 4.012.6%measured
CursorBench 4.038.1%estimated ± 3.5 pp, high confidence
CWE-bench v150.5%estimated ± 6.0 pp, medium confidence
DeepSWE66.6%measured
HLE w/o tools33.8%estimated ± 5.0 pp, medium confidence
OSWorld 2.050.6%estimated ± 13.3 pp, low confidence
PostTrainBench v1.135.5%estimated ± 4.8 pp, high confidence
Terminal-Bench 4.016.5%estimated ± 5.1 pp, low confidence
Terminal-Bench-Science 0.10.1%estimated ± 4.1 pp, low confidence
Toolathlon-Verified76.9%estimated ± 2.2 pp, medium confidence
CharXiv85.7%estimated ± 7.5 pp, low confidence
CharXiv w/o tools82.5%estimated ± 2.4 pp, high confidence
ExploitBench12.5%estimated ± 12.2 pp, low confidence
ExploitGym11.6%measured
C-Eval95.4%estimated ± 1.0 pp, low confidence
MathVision89.7%estimated ± 4.9 pp, high confidence
MCP Atlas84.0%estimated ± 2.7 pp, medium confidence
MMLU-Redux97.0%estimated ± 1.1 pp, low confidence
MMMU-Pro79.1%estimated ± 4.4 pp, medium confidence
NL2Repo57.1%estimated ± 6.5 pp, medium confidence
QwenClawBench100.0%estimated ± 6.0 pp, low confidence
ScreenSpot Pro78.8%estimated ± 10.7 pp, low confidence
SuperGPQA83.1%estimated ± 7.5 pp, medium confidence
Toolathlon55.9%estimated ± 6.3 pp, medium confidence
ERQA68.3%estimated ± 3.9 pp, high confidence
HealthBench Hard31.4%measured
MedXpertQA (MM)76.9%estimated ± 6.5 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals78.6%estimated ± 4.7 pp, medium confidence
ResearchClawBench19.2%estimated ± 2.2 pp, medium confidence
ARC-AGI-30.1%measured
BioMysteryBench (human-solvable)89.0%estimated ± 0.6 pp, medium confidence
DRACO80.8%estimated ± 3.0 pp, medium confidence
HealthBench (raw)50.0%measured
HealthBench (length-adjusted)54.5%measured
HealthBench Professional60.8%measured
HealthBench Professional (raw)61.2%measured
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench265.1%estimated ± 1.3 pp, low confidence
MMMU87.0%estimated ± 0.9 pp, medium confidence
Agents' Last Exam45.7%estimated ± 6.3 pp, medium confidence
OpenHarmony Bench58.8%estimated ± 2.4 pp, high confidence
SEC-Bench Pro34.2%measured
MMLU94.0%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python81.6%estimated ± 1.6 pp, high confidence
RealWorldQA87.5%estimated ± 1.2 pp, medium confidence
OmniDocBench 1.592.0%estimated ± 6.4 pp, low confidence