benchgap
Google · model

Gemini 4 Argon benchmark scores

As of 2026-10-07, Gemini 4 Argon (Google) has measured scores on 20 benchmarks and estimated scores on 67 more.

BenchmarkScoreSource
AA-GPQA Diamond96.6%estimated ± 1.2 pp, low confidence
AA-HLE57.1%measured
AA-SciCode61.8%measured
Artificial Analysis Intelligence Index52.6%measured
CritPt27.1%measured
GDPval-AA56.3%measured
GPQA Diamond97.1%estimated ± 2.2 pp, medium confidence
HLE55.5%estimated ± 9.5 pp, medium confidence
AA-LCR79.7%measured
LiveCodeBench v6100.0%estimated ± 3.3 pp, medium confidence
AA-Omniscience Accuracy49.9%measured
SciCode61.1%estimated ± 3.2 pp, medium confidence
Terminal-Bench 2.189.7%estimated ± 3.9 pp, medium confidence
BrowseComp91.5%estimated ± 2.2 pp, high confidence
HLE w/ tools60.5%estimated ± 6.1 pp, medium confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
AA Coding Index78.6%estimated ± 2.5 pp, medium confidence
APEX-Agents45.1%estimated ± 10.0 pp, low confidence
APEX-Agents-AA38.9%estimated ± 2.9 pp, low confidence
BioMysteryBench (human-difficult)45.6%estimated ± 5.9 pp, low confidence
DeepSearchQA90.4%estimated ± 8.4 pp, medium confidence
SWE-bench Verified89.0%estimated ± 4.4 pp, medium confidence
AutomationBench51.3%measured
BFCL v477.9%estimated ± 2.6 pp, medium confidence
CyberGym88.8%estimated ± 3.7 pp, high confidence
JobBench65.7%estimated ± 4.9 pp, low confidence
SWE-bench Pro81.5%estimated ± 5.1 pp, medium confidence
WideResearch100.0%estimated ± 9.1 pp, low confidence
AA Agentic Index51.1%estimated ± 2.8 pp, high confidence
MMLU-Pro85.7%estimated ± 3.5 pp, medium confidence
GPQA92.8%estimated ± 3.2 pp, high confidence
Gert Labs76.7%estimated ± 3.7 pp, low confidence
AA-AnalystAgent54.4%estimated ± 5.5 pp, medium confidence
AA EnterpriseOps-Gym46.9%estimated ± 4.4 pp, high confidence
AA Harvey LAB93.3%estimated ± 1.3 pp, medium confidence
ARC-AGI-199.7%estimated ± 0.7 pp, low confidence
ARC-AGI-283.0%estimated ± 14.2 pp, low confidence
CursorBench 3.173.1%estimated ± 3.4 pp, low confidence
CursorBench 3.272.4%estimated ± 4.1 pp, low confidence
FrontierCode 1.1 Main53.8%estimated ± 2.8 pp, medium confidence
FrontierSWE v255.0%measured
OSWorld-Verified100.0%estimated ± 3.0 pp, medium confidence
Vals GPQA Diamond96.4%estimated ± 2.2 pp, low confidence
Vals MMLU-Pro92.8%estimated ± 1.4 pp, low confidence
Vals SWE-bench100.0%estimated ± 4.5 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)86.3%estimated ± 4.6 pp, medium confidence
AA AutomationBench77.5%measured
GDP.pdf21.8%measured
AA Tau3 Banking47.5%estimated ± 5.8 pp, medium confidence
AA Terminal-Bench 2.186.8%estimated ± 2.6 pp, high confidence
AA Terminal-Bench 4.057.1%measured
CursorBench 4.049.9%estimated ± 2.9 pp, high confidence
CWE-bench v168.0%measured
DeepSWE77.9%measured
HLE w/o tools61.3%estimated ± 5.7 pp, medium confidence
OSWorld 2.069.2%measured
PostTrainBench v1.145.3%measured
ProgramBench90.9%estimated ± 4.8 pp, medium confidence
Terminal-Bench 4.057.4%measured
Terminal-Bench-Science 0.158.4%estimated ± 4.1 pp, medium confidence
Toolathlon-Verified79.8%estimated ± 1.0 pp, medium confidence
Vibe Code Bench91.9%measured
ExploitGym17.3%estimated ± 5.0 pp, low confidence
C-Eval96.8%estimated ± 1.0 pp, low confidence
MCP Atlas89.2%estimated ± 2.7 pp, low confidence
MMLU-Redux97.4%estimated ± 1.1 pp, low confidence
NL2Repo68.6%estimated ± 6.5 pp, low confidence
QwenClawBench100.0%estimated ± 6.0 pp, low confidence
SuperGPQA85.7%estimated ± 7.5 pp, low confidence
Toolathlon55.9%estimated ± 6.3 pp, low confidence
HealthBench Hard35.0%estimated ± 7.7 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals86.5%estimated ± 3.0 pp, low confidence
ResearchClawBench19.9%estimated ± 2.2 pp, low confidence
SWE-Rebench62.3%estimated ± 6.5 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, high confidence
BioMysteryBench (human-solvable)89.5%estimated ± 0.6 pp, medium confidence
DRACO84.6%estimated ± 3.0 pp, medium confidence
HealthBench (raw)63.4%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)59.5%estimated ± 3.7 pp, medium confidence
HealthBench Professional64.9%estimated ± 4.8 pp, medium confidence
HealthBench Professional (raw)72.9%estimated ± 5.0 pp, medium confidence
HLE-Verified54.8%estimated ± 1.2 pp, low confidence
LABBench288.8%measured
Agents' Last Exam39.5%measured
OpenHarmony Bench58.8%estimated ± 2.4 pp, medium confidence
MMLU94.9%estimated ± 1.1 pp, low confidence