benchgap
OpenAI · model

GPT-5.3 Codex benchmark scores

As of 2026-10-07, GPT-5.3 Codex (OpenAI) has measured scores on 19 benchmarks and estimated scores on 80 more.

BenchmarkScoreSource
AA-GPQA Diamond91.5%measured
AA-HLE42.5%measured
AA-SciCode48.7%estimated ± 3.8 pp, high confidence
Artificial Analysis Intelligence Index32.5%measured
CritPt16.9%measured
GDPval-AA36.8%estimated ± 6.2 pp, medium confidence
GPQA Diamond92.3%estimated ± 2.2 pp, high confidence
HLE49.6%estimated ± 9.5 pp, medium confidence
IFBench74.0%estimated ± 6.9 pp, medium confidence
AA-LCR83.3%measured
LiveCodeBench v687.5%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy52.9%measured
SciCode51.3%estimated ± 2.9 pp, medium confidence
BrowseComp74.3%estimated ± 3.1 pp, high confidence
IFEval93.7%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
τ²-bench results86.0%measured
AA Coding Index68.6%estimated ± 5.1 pp, medium confidence
AA-MMMU-Pro78.5%measured
APEX-Agents-AA31.6%estimated ± 3.6 pp, medium confidence
BioMysteryBench (human-difficult)46.8%estimated ± 5.9 pp, low confidence
DeepSearchQA75.1%estimated ± 8.5 pp, low confidence
SWE-bench Verified85.0%measured
BFCL v461.6%estimated ± 9.8 pp, low confidence
CyberGym56.4%estimated ± 10.0 pp, low confidence
JobBench33.7%measured
SWE-bench Pro56.8%measured
τ³-bench results68.9%estimated ± 8.9 pp, medium confidence
WideResearch77.7%estimated ± 1.2 pp, medium confidence
AA Agentic Index21.1%estimated ± 6.4 pp, medium confidence
MMLU-Pro85.7%estimated ± 3.5 pp, medium confidence
GPQA90.8%estimated ± 3.2 pp, high confidence
AA-IFBench75.4%measured
Gert Labs57.5%measured
AA EnterpriseOps-Gym31.3%estimated ± 3.6 pp, medium confidence
ApprenticeBench0.0%estimated ± 3.1 pp, low confidence
ARC-AGI-189.1%estimated ± 5.4 pp, medium confidence
ARC-AGI-258.9%estimated ± 14.2 pp, low confidence
CursorBench 3.152.5%estimated ± 5.6 pp, low confidence
CursorBench 3.253.5%estimated ± 3.6 pp, high confidence
FrontierCode 1.1 Main40.7%estimated ± 2.7 pp, medium confidence
FrontierSWE v29.4%estimated ± 9.0 pp, medium confidence
OfficeQA Pro50.7%estimated ± 7.2 pp, medium confidence
OSWorld-Verified64.7%measured
Vals GPQA Diamond91.4%estimated ± 7.4 pp, medium confidence
Vals LiveCodeBench87.3%measured
Vals MMLU-Pro89.2%estimated ± 2.2 pp, high confidence
Vals SWE-bench78.0%measured
Terminal-Bench 2.1 (Vals AI)68.0%estimated ± 6.3 pp, medium confidence
VulcanBench v385.2%estimated ± 5.6 pp, low confidence
AA AutomationBench49.0%estimated ± 5.0 pp, medium confidence
GDP.pdf17.8%estimated ± 5.2 pp, low confidence
AA ITBench52.2%estimated ± 3.6 pp, medium confidence
CursorBench 4.033.0%estimated ± 3.6 pp, medium confidence
DeepSWE56.3%estimated ± 7.2 pp, low confidence
HLE w/o tools41.4%estimated ± 5.7 pp, medium confidence
OSWorld 2.00.0%estimated ± 3.6 pp, medium confidence
PostTrainBench v1.123.0%estimated ± 2.7 pp, high confidence
Toolathlon-Verified73.2%estimated ± 1.3 pp, low confidence
Vibe Code Bench61.8%measured
CharXiv85.4%estimated ± 7.5 pp, low confidence
CharXiv w/o tools81.8%estimated ± 2.4 pp, high confidence
C-Eval94.6%estimated ± 1.0 pp, low confidence
Claw-Eval65.5%estimated ± 5.6 pp, low confidence
MathVision89.6%estimated ± 4.9 pp, high confidence
MCP Atlas58.4%estimated ± 8.2 pp, medium confidence
MCP-Tasks72.1%estimated ± 8.2 pp, low confidence
MMLU-Redux97.5%estimated ± 1.1 pp, low confidence
MMMU-Pro78.5%estimated ± 4.4 pp, medium confidence
NL2Repo44.6%estimated ± 4.7 pp, high confidence
ScreenSpot Pro77.5%estimated ± 10.7 pp, low confidence
SuperGPQA86.8%estimated ± 7.5 pp, low confidence
Terminal-Bench 2.077.3%measured
Toolathlon46.6%estimated ± 7.2 pp, medium confidence
ERQA67.8%estimated ± 3.9 pp, high confidence
HealthBench Hard29.5%estimated ± 7.7 pp, low confidence
LiveCodeBench Pro81.2%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)76.3%estimated ± 6.5 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals78.0%estimated ± 2.4 pp, medium confidence
ResearchClawBench17.0%estimated ± 2.2 pp, medium confidence
SWE-Rebench58.2%measured
ARC-AGI-31.7%estimated ± 3.4 pp, high confidence
BioMysteryBench (human-solvable)88.4%estimated ± 0.6 pp, medium confidence
FrontierCode 1.1 Extended27.1%estimated ± 2.4 pp, low confidence
HealthBench (raw)51.4%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional57.7%estimated ± 4.8 pp, medium confidence
HealthBench Professional (raw)61.6%estimated ± 5.0 pp, medium confidence
HLE-Verified38.7%estimated ± 2.0 pp, medium confidence
LABBench279.4%estimated ± 1.8 pp, low confidence
MMMU85.9%estimated ± 0.9 pp, medium confidence
LiveCodeBench89.5%estimated ± 10.2 pp, low confidence
Agents' Last Exam48.9%estimated ± 8.7 pp, low confidence
OpenHarmony Bench53.7%estimated ± 3.2 pp, high confidence
MMLU95.3%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python80.8%estimated ± 1.6 pp, high confidence
RealWorldQA86.8%estimated ± 1.2 pp, medium confidence
OmniDocBench 1.592.0%estimated ± 6.4 pp, low confidence