benchgap
OpenAI · model

GPT-5.4 mini benchmark scores

As of 2026-10-07, GPT-5.4 mini (OpenAI) has measured scores on 34 benchmarks and estimated scores on 89 more.

BenchmarkScoreSource
AA-GPQA Diamond87.5%measured
AA-HLE28.1%measured
AA-SciCode52.1%measured
AIME2695.5%estimated ± 0.6 pp, low confidence
Artificial Analysis Intelligence Index24.1%measured
CritPt10.0%measured
GDPval-AA25.8%measured
GPQA Diamond88.0%estimated ± 0.1 pp, high confidence
HLE41.5%measured
IFBench71.8%estimated ± 6.9 pp, medium confidence
AA-LCR77.0%measured
LiveCodeBench v687.7%estimated ± 3.1 pp, high confidence
AA-Omniscience Accuracy37.5%measured
SciCode39.7%estimated ± 2.9 pp, low confidence
Terminal-Bench 2.171.0%estimated ± 6.9 pp, medium confidence
BrowseComp83.7%estimated ± 3.1 pp, high confidence
HLE w/ tools44.8%estimated ± 6.1 pp, medium confidence
IFEval93.6%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
VITA-Bench38.5%estimated ± 9.9 pp, low confidence
τ²-bench results93.4%measured
AA Coding Index56.1%measured
AA-MMMU-Pro73.3%measured
APEX-Agents16.0%estimated ± 0.6 pp, low confidence
APEX-Agents-AA28.2%measured
BioMysteryBench (human-difficult)39.6%estimated ± 5.9 pp, low confidence
DeepSearchQA84.0%estimated ± 4.5 pp, high confidence
SWE-bench Verified80.2%estimated ± 0.8 pp, medium confidence
AutomationBench28.4%estimated ± 9.7 pp, low confidence
BFCL v474.3%estimated ± 2.6 pp, high confidence
CyberGym60.6%estimated ± 5.3 pp, low confidence
JobBench30.7%estimated ± 6.8 pp, medium confidence
SWE-bench Pro57.1%estimated ± 2.1 pp, high confidence
τ³-bench results69.5%estimated ± 1.8 pp, medium confidence
WideResearch76.0%estimated ± 5.6 pp, low confidence
AA Agentic Index19.7%measured
MMLU-Pro85.1%estimated ± 2.5 pp, medium confidence
FrontierMath v2 (Tier 4)2.1%measured
FrontierMath v2 (Tiers 1-3)28.3%measured
GPQA88.0%measured
AA-IFBench73.3%measured
Gert Labs50.7%estimated ± 7.2 pp, medium confidence
AA-AnalystAgent18.1%estimated ± 12.0 pp, low confidence
AA EnterpriseOps-Gym43.7%estimated ± 3.6 pp, medium confidence
AA Harvey LAB80.8%estimated ± 1.2 pp, low confidence
ApprenticeBench3.5%estimated ± 5.0 pp, low confidence
ARC-AGI-163.7%measured
ARC-AGI-218.9%measured
CursorBench 3.149.2%estimated ± 3.4 pp, medium confidence
CursorBench 3.242.5%estimated ± 1.6 pp, low confidence
Terminal-Bench 3.03.3%estimated ± 6.9 pp, low confidence
FrontierCode 1.1 Main27.0%measured
FrontierSWE v27.3%estimated ± 9.5 pp, low confidence
OfficeQA Pro50.3%estimated ± 7.2 pp, low confidence
OSWorld-Verified72.1%measured
Vals GPQA Diamond83.1%measured
Vals LiveCodeBench81.5%measured
Vals MMLU-Pro84.6%measured
Vals SWE-bench73.0%measured
Terminal-Bench 2.1 (Vals AI)54.7%measured
VulcanBench v381.6%estimated ± 5.6 pp, low confidence
AA AutomationBench4.3%estimated ± 7.0 pp, medium confidence
GDP.pdf9.4%estimated ± 5.4 pp, medium confidence
AA ITBench29.0%estimated ± 3.6 pp, low confidence
AA Tau3 Banking22.0%estimated ± 5.8 pp, medium confidence
AA Terminal-Bench 2.168.7%estimated ± 4.4 pp, high confidence
AA Terminal-Bench 4.07.7%estimated ± 14.0 pp, low confidence
CursorBench 4.019.8%estimated ± 1.6 pp, medium confidence
CWE-bench v135.7%estimated ± 6.9 pp, low confidence
DeepSWE22.2%estimated ± 3.4 pp, low confidence
HLE w/o tools28.2%measured
OSWorld 2.05.9%estimated ± 3.6 pp, high confidence
PostTrainBench v1.121.8%estimated ± 2.7 pp, high confidence
Toolathlon-Verified47.7%estimated ± 2.3 pp, low confidence
Vibe Code Bench48.0%measured
CharXiv79.2%estimated ± 5.4 pp, medium confidence
CharXiv w/o tools80.7%estimated ± 1.8 pp, medium confidence
ExploitGym0.0%estimated ± 12.2 pp, low confidence
C-Eval93.2%estimated ± 0.8 pp, low confidence
Claw-Eval63.9%estimated ± 5.7 pp, medium confidence
DeepPlanning38.8%estimated ± 13.1 pp, low confidence
HMMT Nov 202596.8%estimated ± 2.6 pp, low confidence
MathVision83.7%estimated ± 3.6 pp, high confidence
MCP Atlas57.7%measured
MCP-Tasks76.2%estimated ± 0.6 pp, low confidence
MMAnswerBench82.4%estimated ± 1.3 pp, low confidence
MMLU-Redux96.4%estimated ± 1.1 pp, medium confidence
MMMU-Pro76.6%measured
NL2Repo42.0%estimated ± 4.7 pp, high confidence
QwenClawBench55.2%estimated ± 2.6 pp, high confidence
ScreenSpot Pro69.7%estimated ± 10.0 pp, low confidence
SuperGPQA79.7%estimated ± 7.5 pp, medium confidence
Terminal-Bench 2.060.0%measured
Toolathlon42.9%measured
VideoMMMU84.4%estimated ± 1.0 pp, high confidence
ERQA56.4%estimated ± 3.9 pp, high confidence
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
MedXpertQA (MM)67.0%estimated ± 2.4 pp, high confidence
MedXpertQA (Text)44.3%estimated ± 8.9 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals70.1%estimated ± 2.4 pp, low confidence
ResearchClawBench17.0%estimated ± 1.8 pp, medium confidence
SWE-Rebench62.0%estimated ± 6.5 pp, low confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
FrontierMath (legacy)29.0%estimated ± 0.7 pp, low confidence
BioMysteryBench (human-solvable)83.3%estimated ± 0.6 pp, low confidence
DRACO72.1%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended19.6%estimated ± 2.4 pp, low confidence
HealthBench (raw)37.2%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)43.7%estimated ± 2.5 pp, low confidence
HealthBench Professional51.7%estimated ± 1.3 pp, low confidence
HealthBench Professional (raw)48.4%estimated ± 5.0 pp, low confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench222.8%estimated ± 1.3 pp, low confidence
MMMU81.6%estimated ± 0.9 pp, high confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
OpenHarmony Bench58.6%estimated ± 2.4 pp, high confidence
SimpleVQA66.2%estimated ± 8.2 pp, medium confidence
MMLU92.8%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python78.0%measured
RealWorldQA84.0%estimated ± 1.2 pp, low confidence
Video-MME (with subtitle)86.9%estimated ± 1.8 pp, medium confidence
OmniDocBench 1.591.5%estimated ± 5.9 pp, low confidence