benchgap
Anthropic · model

Claude Opus 4.7 (Adaptive) benchmark scores

As of 2026-10-07, Claude Opus 4.7 (Adaptive) (Anthropic) has measured scores on 32 benchmarks and estimated scores on 87 more.

BenchmarkScoreSource
AA-GPQA Diamond91.4%measured
AA-HLE42.3%measured
AA-SciCode55.4%estimated ± 3.2 pp, high confidence
Artificial Analysis Intelligence Index40.7%measured
CritPt12.0%measured
GDPval-AA42.8%measured
GPQA Diamond94.2%measured
HLE54.7%measured
IFBench60.6%estimated ± 6.9 pp, medium confidence
AA-LCR78.7%measured
LiveCodeBench v690.4%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy48.9%measured
SciCode51.9%estimated ± 3.8 pp, medium confidence
BrowseComp79.3%measured
HLE w/ tools53.1%estimated ± 4.6 pp, high confidence
IFEval92.6%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
VITA-Bench47.1%estimated ± 9.9 pp, low confidence
τ²-bench results88.6%measured
AA Coding Index73.6%measured
AA-MMMU-Pro78.8%measured
APEX-Agents31.7%estimated ± 0.6 pp, medium confidence
APEX-Agents-AA36.1%estimated ± 2.9 pp, medium confidence
BioMysteryBench (human-difficult)45.1%estimated ± 5.9 pp, low confidence
DeepSearchQA91.6%estimated ± 4.5 pp, high confidence
SWE-bench Verified87.6%measured
AutomationBench28.4%estimated ± 5.8 pp, low confidence
BFCL v476.8%estimated ± 2.6 pp, medium confidence
CyberGym73.1%measured
JobBench45.9%measured
SWE-bench Pro64.3%measured
τ³-bench results69.6%estimated ± 1.8 pp, low confidence
WideResearch80.4%estimated ± 1.2 pp, medium confidence
AA Agentic Index39.5%measured
MMLU-Pro85.8%estimated ± 2.5 pp, low confidence
FrontierMath v2 (Tier 4)35.2%estimated ± 5.0 pp, low confidence
FrontierMath v2 (Tiers 1-3)43.3%estimated ± 0.7 pp, low confidence
GPQA94.2%measured
AA-IFBench58.6%measured
Gert Labs71.5%estimated ± 3.7 pp, medium confidence
AA-AnalystAgent53.7%estimated ± 5.5 pp, medium confidence
AA EnterpriseOps-Gym35.5%estimated ± 3.6 pp, medium confidence
AA Harvey LAB92.1%estimated ± 1.2 pp, medium confidence
ApprenticeBench4.0%estimated ± 4.0 pp, medium confidence
ARC-AGI-193.7%estimated ± 1.9 pp, high confidence
ARC-AGI-275.8%measured
CursorBench 3.155.2%estimated ± 4.8 pp, medium confidence
CursorBench 3.262.9%estimated ± 3.6 pp, high confidence
FrontierCode 1.1 Main46.3%estimated ± 2.7 pp, medium confidence
FrontierSWE v222.8%estimated ± 9.0 pp, medium confidence
OfficeQA Pro43.6%measured
OSWorld-Verified78.0%measured
Vals GPQA Diamond91.7%estimated ± 1.1 pp, high confidence
Vals LiveCodeBench86.3%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro88.5%estimated ± 1.2 pp, high confidence
Vals SWE-bench87.7%estimated ± 5.4 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)61.0%estimated ± 6.3 pp, medium confidence
AA AutomationBench52.4%estimated ± 5.0 pp, medium confidence
GDP.pdf19.8%estimated ± 4.9 pp, high confidence
AA ITBench46.7%measured
AA Tau3 Banking46.9%estimated ± 3.0 pp, low confidence
CursorBench 4.037.3%estimated ± 1.6 pp, high confidence
CWE-bench v150.7%estimated ± 6.9 pp, medium confidence
DeepSWE62.8%estimated ± 6.4 pp, medium confidence
HLE w/o tools46.9%measured
OSWorld 2.018.2%measured
PostTrainBench v1.128.4%estimated ± 4.7 pp, high confidence
Toolathlon-Verified72.3%estimated ± 1.0 pp, low confidence
Vibe Code Bench62.4%estimated ± 8.1 pp, low confidence
CharXiv91.0%measured
CharXiv w/o tools82.1%measured
ExploitGym0.0%estimated ± 6.1 pp, low confidence
C-Eval93.3%estimated ± 0.8 pp, low confidence
Claw-Eval66.2%estimated ± 5.6 pp, low confidence
DeepPlanning48.8%estimated ± 13.1 pp, low confidence
MathVision88.3%estimated ± 2.1 pp, low confidence
MCP Atlas77.3%measured
MCP-Tasks77.0%estimated ± 0.6 pp, low confidence
MMLU-Redux97.3%estimated ± 1.1 pp, low confidence
MMMU-Pro76.7%estimated ± 1.4 pp, medium confidence
NL2Repo54.5%estimated ± 5.1 pp, medium confidence
QwenClawBench63.9%estimated ± 2.6 pp, medium confidence
ScreenSpot Pro87.3%estimated ± 7.1 pp, medium confidence
SuperGPQA85.3%estimated ± 7.5 pp, low confidence
Terminal-Bench 2.069.4%measured
Toolathlon51.7%estimated ± 3.5 pp, high confidence
VideoMMMU87.4%estimated ± 1.5 pp, high confidence
ERQA71.8%estimated ± 3.7 pp, high confidence
HealthBench Hard36.5%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro87.0%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)79.4%estimated ± 6.4 pp, low confidence
MedXpertQA (Text)68.9%estimated ± 4.5 pp, medium confidence
MMLU-Pro (Arcee)88.2%estimated ± 3.4 pp, low confidence
React Native Evals82.6%estimated ± 3.5 pp, high confidence
ResearchClawBench18.8%estimated ± 1.8 pp, medium confidence
SWE-Rebench61.8%estimated ± 3.6 pp, low confidence
ARC-AGI-30.2%measured
FrontierMath (legacy)43.8%measured
BioMysteryBench (human-solvable)85.1%estimated ± 0.6 pp, low confidence
DRACO79.1%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended57.4%estimated ± 3.0 pp, medium confidence
HealthBench (raw)51.2%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.8%estimated ± 0.9 pp, medium confidence
HealthBench Professional56.8%estimated ± 4.3 pp, medium confidence
HealthBench Professional (raw)61.4%estimated ± 5.0 pp, medium confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench279.3%estimated ± 1.8 pp, low confidence
MMMU86.2%estimated ± 0.9 pp, medium confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
Agents' Last Exam24.0%estimated ± 2.1 pp, medium confidence
OpenHarmony Bench58.1%estimated ± 2.8 pp, high confidence
SimpleVQA72.4%estimated ± 7.8 pp, low confidence
MMMLU91.1%estimated ± 1.1 pp, low confidence
MMLU94.8%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python81.0%estimated ± 1.6 pp, high confidence
MathVision w/ Python96.6%estimated ± 0.9 pp, medium confidence
RealWorldQA87.5%estimated ± 1.1 pp, medium confidence
Video-MME (with subtitle)89.6%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.592.0%estimated ± 6.4 pp, low confidence