benchgap
OpenAI · model

GPT-5.2 benchmark scores

As of 2026-10-07, GPT-5.2 (OpenAI) has measured scores on 24 benchmarks and estimated scores on 91 more.

BenchmarkScoreSource
AA-GPQA Diamond90.3%measured
AA-HLE37.7%measured
AA-SciCode48.1%estimated ± 3.8 pp, high confidence
AIME26100.0%estimated ± 0.6 pp, low confidence
Artificial Analysis Intelligence Index30.5%measured
CritPt11.6%measured
GDPval-AA37.3%estimated ± 6.2 pp, medium confidence
GPQA Diamond92.4%estimated ± 0.4 pp, high confidence
HLE44.8%estimated ± 6.3 pp, medium confidence
IFBench74.0%estimated ± 6.9 pp, medium confidence
AA-LCR82.7%measured
LiveCodeBench v687.1%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy44.3%measured
SciCode48.5%estimated ± 3.8 pp, high confidence
BrowseComp65.8%measured
HLE w/ tools35.2%estimated ± 4.6 pp, high confidence
IFEval93.7%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
VITA-Bench38.4%estimated ± 9.9 pp, low confidence
τ²-bench results84.8%measured
AA Coding Index60.2%estimated ± 5.1 pp, medium confidence
AA-MMMU-Pro77.6%estimated ± 4.9 pp, medium confidence
APEX-Agents-AA31.9%estimated ± 3.6 pp, medium confidence
BioMysteryBench (human-difficult)43.1%estimated ± 5.9 pp, low confidence
DeepSearchQA78.6%estimated ± 8.1 pp, low confidence
SWE-bench Verified80.0%measured
AutomationBench19.8%estimated ± 10.3 pp, low confidence
BFCL v460.5%estimated ± 9.8 pp, low confidence
CyberGym68.1%estimated ± 4.6 pp, medium confidence
JobBench34.3%measured
SWE-bench Pro55.6%measured
τ³-bench results68.9%estimated ± 8.9 pp, medium confidence
WideResearch77.8%estimated ± 1.2 pp, medium confidence
AA Agentic Index14.0%estimated ± 6.4 pp, medium confidence
MMLU-Pro86.9%estimated ± 3.0 pp, high confidence
FrontierMath v2 (Tier 4)18.8%measured
FrontierMath v2 (Tiers 1-3)40.7%measured
GPQA92.4%measured
AA-IFBench75.4%measured
Gert Labs46.5%measured
AA-AnalystAgent12.0%estimated ± 12.5 pp, low confidence
AA EnterpriseOps-Gym29.5%estimated ± 3.6 pp, medium confidence
AA Harvey LAB86.9%estimated ± 9.8 pp, low confidence
ApprenticeBench0.0%estimated ± 3.1 pp, low confidence
ARC-AGI-186.2%measured
ARC-AGI-252.9%measured
CursorBench 3.151.7%estimated ± 5.6 pp, low confidence
CursorBench 3.251.4%estimated ± 3.6 pp, high confidence
FrontierCode 1.1 Main25.1%estimated ± 2.7 pp, medium confidence
FrontierSWE v27.3%estimated ± 9.0 pp, medium confidence
OSWorld-Verified47.3%measured
Vals GPQA Diamond90.0%estimated ± 2.2 pp, high confidence
Vals LiveCodeBench82.5%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro88.0%estimated ± 1.2 pp, high confidence
Vals SWE-bench80.2%estimated ± 6.3 pp, medium confidence
AA AutomationBench49.2%estimated ± 5.0 pp, medium confidence
GDP.pdf17.8%estimated ± 5.2 pp, low confidence
AA ITBench53.3%estimated ± 3.6 pp, low confidence
AA Tau3 Banking19.2%estimated ± 10.2 pp, low confidence
CursorBench 4.032.1%estimated ± 3.6 pp, medium confidence
HLE w/o tools41.9%estimated ± 4.9 pp, high confidence
OSWorld 2.00.0%estimated ± 3.6 pp, medium confidence
PostTrainBench v1.123.0%estimated ± 4.7 pp, medium confidence
Toolathlon-Verified73.2%estimated ± 1.3 pp, low confidence
Vibe Code Bench53.5%measured
CharXiv82.1%measured
CharXiv w/o tools83.7%estimated ± 1.7 pp, low confidence
ExploitGym0.0%estimated ± 6.1 pp, low confidence
C-Eval94.5%estimated ± 0.9 pp, low confidence
Claw-Eval66.0%estimated ± 5.6 pp, low confidence
HMMT Nov 202593.2%estimated ± 2.6 pp, low confidence
MathVision83.0%measured
MCP Atlas59.1%estimated ± 8.2 pp, medium confidence
MCP-Tasks72.8%estimated ± 8.2 pp, low confidence
MMAnswerBench86.2%estimated ± 1.3 pp, low confidence
MMLU-Redux97.0%estimated ± 1.1 pp, low confidence
MMMU-Pro79.5%measured
NL2Repo47.3%estimated ± 6.1 pp, medium confidence
ScreenSpot Pro61.6%estimated ± 4.5 pp, medium confidence
SuperGPQA83.3%estimated ± 7.5 pp, medium confidence
Toolathlon38.6%estimated ± 3.5 pp, high confidence
V*75.9%measured
VideoMMMU86.0%estimated ± 1.0 pp, high confidence
ERQA59.1%estimated ± 2.4 pp, low confidence
HealthBench Hard32.1%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro80.2%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)73.5%estimated ± 2.4 pp, high confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals82.2%estimated ± 3.0 pp, medium confidence
ResearchClawBench15.6%estimated ± 2.2 pp, medium confidence
SWE-Rebench61.8%estimated ± 3.6 pp, medium confidence
ARC-AGI-30.1%estimated ± 2.1 pp, high confidence
FrontierMath (legacy)41.3%estimated ± 0.7 pp, low confidence
BioMysteryBench (human-solvable)84.7%estimated ± 0.6 pp, low confidence
FrontierCode 1.1 Extended54.5%estimated ± 3.0 pp, low confidence
HealthBench (raw)46.9%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.6%estimated ± 2.5 pp, low confidence
HealthBench Professional55.1%estimated ± 4.6 pp, medium confidence
HealthBench Professional (raw)57.4%estimated ± 5.0 pp, low confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench264.1%estimated ± 1.3 pp, low confidence
MMMU85.3%estimated ± 1.1 pp, low confidence
LiveCodeBench86.5%estimated ± 10.2 pp, low confidence
Agents' Last Exam48.9%estimated ± 8.7 pp, low confidence
OpenHarmony Bench53.3%estimated ± 4.9 pp, low confidence
BabyVision33.5%estimated ± 4.3 pp, low confidence
MMVU73.9%estimated ± 2.9 pp, low confidence
SimpleVQA67.7%estimated ± 7.8 pp, low confidence
ZeroBench11.0%estimated ± 4.7 pp, low confidence
MMLU94.1%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python80.8%estimated ± 0.6 pp, high confidence
MathVision w/ Python91.1%estimated ± 0.5 pp, low confidence
RealWorldQA0.0%estimated ± 0.8 pp, low confidence
Video-MME (with subtitle)87.8%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.591.7%estimated ± 5.9 pp, low confidence