benchgap
Anthropic · model

Claude Opus 4.7 benchmark scores

As of 2026-10-07, Claude Opus 4.7 (Anthropic) has measured scores on 24 benchmarks and estimated scores on 86 more.

BenchmarkScoreSource
AA-GPQA Diamond88.5%measured
AA-HLE33.3%measured
AA-SciCode55.8%estimated ± 3.5 pp, high confidence
AIME26100.0%estimated ± 0.6 pp, low confidence
Artificial Analysis Intelligence Index30.9%measured
CritPt5.1%measured
GDPval-AA43.3%estimated ± 7.5 pp, medium confidence
GPQA Diamond90.8%estimated ± 1.7 pp, high confidence
HLE37.5%estimated ± 9.5 pp, medium confidence
IFBench56.3%estimated ± 6.9 pp, medium confidence
AA-LCR75.7%measured
AA-Omniscience Accuracy44.7%measured
SciCode46.9%estimated ± 2.9 pp, medium confidence
Terminal-Bench 2.183.3%estimated ± 6.9 pp, medium confidence
BrowseComp82.8%estimated ± 2.2 pp, high confidence
IFEval89.8%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
τ²-bench results74.0%measured
AA Coding Index73.5%estimated ± 3.1 pp, high confidence
AA-MMMU-Pro76.4%measured
APEX-Agents-AA35.1%estimated ± 2.9 pp, medium confidence
BioMysteryBench (human-difficult)43.3%estimated ± 5.9 pp, low confidence
DeepSearchQA77.2%estimated ± 7.5 pp, low confidence
SWE-bench Verified83.1%estimated ± 0.8 pp, medium confidence
BFCL v452.7%estimated ± 9.8 pp, low confidence
CyberGym72.5%estimated ± 5.0 pp, medium confidence
JobBench47.3%estimated ± 4.9 pp, medium confidence
SWE-bench Pro57.2%estimated ± 2.1 pp, high confidence
τ³-bench results68.9%estimated ± 8.9 pp, low confidence
WideResearch79.4%estimated ± 5.9 pp, low confidence
AA Agentic Index28.7%estimated ± 6.4 pp, medium confidence
MMLU-Pro85.7%estimated ± 3.5 pp, high confidence
FrontierMath v2 (Tier 4)22.9%measured
FrontierMath v2 (Tiers 1-3)43.8%measured
GPQA93.5%estimated ± 1.6 pp, medium confidence
AA-IFBench43.6%measured
Gert Labs65.6%measured
AA-AnalystAgent51.4%estimated ± 5.5 pp, medium confidence
AA EnterpriseOps-Gym15.6%estimated ± 3.6 pp, low confidence
AA Harvey LAB89.3%estimated ± 1.3 pp, medium confidence
ApprenticeBench7.0%measured
ARC-AGI-187.4%estimated ± 5.4 pp, medium confidence
ARC-AGI-261.8%estimated ± 12.6 pp, low confidence
CursorBench 3.156.2%estimated ± 5.8 pp, low confidence
CursorBench 3.255.6%estimated ± 1.6 pp, low confidence
Terminal-Bench 3.011.9%estimated ± 6.9 pp, medium confidence
FrontierCode 1.1 Main38.5%measured
FrontierSWE v211.6%estimated ± 9.5 pp, medium confidence
OfficeQA Pro50.3%estimated ± 7.2 pp, medium confidence
OSWorld-Verified74.2%estimated ± 1.9 pp, medium confidence
Vals GPQA Diamond90.2%measured
Vals LiveCodeBench85.1%measured
Vals MMLU-Pro89.9%measured
Vals SWE-bench82.0%measured
Terminal-Bench 2.1 (Vals AI)68.5%measured
VulcanBench v383.7%estimated ± 5.6 pp, low confidence
AA AutomationBench41.7%estimated ± 7.0 pp, medium confidence
GDP.pdf22.0%estimated ± 4.4 pp, low confidence
AA ITBench65.8%estimated ± 3.6 pp, low confidence
AA Tau3 Banking31.0%estimated ± 14.6 pp, low confidence
AA Terminal-Bench 2.183.8%estimated ± 4.4 pp, high confidence
AA Terminal-Bench 4.023.3%estimated ± 14.0 pp, low confidence
CursorBench 4.032.5%estimated ± 4.4 pp, medium confidence
DeepSWE58.9%estimated ± 3.4 pp, low confidence
HLE w/o tools51.3%estimated ± 5.1 pp, medium confidence
OSWorld 2.013.9%measured
PostTrainBench v1.128.6%measured
ProgramBench0.5%estimated ± 4.8 pp, low confidence
Toolathlon-Verified69.6%estimated ± 1.0 pp, low confidence
Vibe Code Bench71.0%measured
CharXiv83.8%estimated ± 7.5 pp, low confidence
CharXiv w/o tools80.5%estimated ± 2.4 pp, high confidence
ExploitGym3.5%estimated ± 6.9 pp, low confidence
C-Eval94.3%estimated ± 1.0 pp, low confidence
Claw-Eval73.3%estimated ± 3.2 pp, medium confidence
HMMT Nov 202593.2%estimated ± 2.6 pp, low confidence
MathVision89.1%estimated ± 4.9 pp, high confidence
MCP Atlas79.6%estimated ± 10.4 pp, low confidence
MCP-Tasks79.7%estimated ± 8.2 pp, low confidence
MMAnswerBench87.1%estimated ± 1.3 pp, low confidence
MMLU-Redux97.0%estimated ± 1.1 pp, low confidence
MMMU-Pro77.6%estimated ± 4.4 pp, medium confidence
NL2Repo47.2%estimated ± 4.7 pp, high confidence
ScreenSpot Pro75.2%estimated ± 10.7 pp, low confidence
SuperGPQA83.5%estimated ± 7.5 pp, medium confidence
Terminal-Bench 2.072.5%estimated ± 5.9 pp, medium confidence
Toolathlon53.0%estimated ± 1.7 pp, medium confidence
ERQA65.9%estimated ± 3.9 pp, high confidence
HealthBench Hard27.8%estimated ± 7.7 pp, low confidence
MedXpertQA (MM)74.3%estimated ± 6.5 pp, low confidence
MMLU-Pro (Arcee)96.3%estimated ± 4.0 pp, low confidence
React Native Evals82.8%measured
ResearchClawBench20.7%measured
SWE-Rebench62.2%estimated ± 6.5 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, high confidence
FrontierMath (legacy)44.3%estimated ± 0.7 pp, low confidence
BioMysteryBench (human-solvable)81.8%estimated ± 0.6 pp, low confidence
FrontierCode 1.1 Extended33.8%estimated ± 2.4 pp, low confidence
HealthBench (raw)42.6%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional59.9%estimated ± 1.3 pp, medium confidence
HealthBench Professional (raw)53.4%estimated ± 5.0 pp, low confidence
HLE-Verified15.5%estimated ± 2.0 pp, low confidence
LABBench283.0%estimated ± 1.8 pp, low confidence
MMMU84.2%estimated ± 0.9 pp, high confidence
OpenHarmony Bench54.9%estimated ± 3.2 pp, high confidence
MMLU94.2%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python79.5%estimated ± 1.6 pp, high confidence
RealWorldQA85.7%estimated ± 1.2 pp, medium confidence
OmniDocBench 1.591.7%estimated ± 6.4 pp, low confidence