benchgap
Other · model

Beam benchmark scores

As of 2026-10-07, Beam (Other) has measured scores on 14 benchmarks and estimated scores on 81 more.

BenchmarkScoreSource
AA-GPQA Diamond89.3%estimated ± 2.7 pp, high confidence
AA-HLE36.6%estimated ± 4.0 pp, high confidence
AA-SciCode49.2%estimated ± 3.1 pp, high confidence
AIME2697.8%measured
Artificial Analysis Intelligence Index30.8%estimated ± 5.1 pp, medium confidence
CritPt12.2%estimated ± 4.9 pp, high confidence
GDPval-AA35.5%estimated ± 8.4 pp, medium confidence
GPQA Diamond90.5%measured
HLE45.9%estimated ± 5.7 pp, medium confidence
IFBench79.7%measured
AA-LCR73.3%estimated ± 7.0 pp, low confidence
LiveCodeBench v690.9%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy38.2%estimated ± 8.6 pp, medium confidence
SciCode49.7%measured
Terminal-Bench 2.180.1%measured
BrowseComp77.4%measured
HLE w/ tools50.3%estimated ± 4.6 pp, high confidence
IFEval92.8%estimated ± 3.5 pp, high confidence
LongBench v265.5%measured
VITA-Bench47.7%estimated ± 9.9 pp, low confidence
AA Coding Index68.3%estimated ± 2.5 pp, high confidence
APEX-Agents32.8%estimated ± 0.6 pp, medium confidence
APEX-Agents-AA23.2%estimated ± 4.6 pp, high confidence
DeepSearchQA80.1%measured
SWE-bench Verified80.9%measured
AutomationBench27.1%estimated ± 10.3 pp, low confidence
CyberGym75.6%estimated ± 4.6 pp, high confidence
JobBench49.1%estimated ± 6.8 pp, medium confidence
SWE-bench Pro65.5%measured
τ³-bench results69.6%estimated ± 1.8 pp, low confidence
WideResearch76.6%estimated ± 4.6 pp, high confidence
AA Agentic Index31.1%estimated ± 8.2 pp, medium confidence
MMLU-Pro85.5%estimated ± 2.5 pp, medium confidence
GPQA90.5%estimated ± 0.4 pp, high confidence
AA-IFBench80.3%estimated ± 8.5 pp, medium confidence
Gert Labs58.5%estimated ± 7.4 pp, medium confidence
AA-AnalystAgent23.5%estimated ± 12.5 pp, low confidence
AA EnterpriseOps-Gym40.1%estimated ± 6.9 pp, low confidence
AA Harvey LAB89.9%estimated ± 9.8 pp, low confidence
ApprenticeBench11.8%estimated ± 10.1 pp, low confidence
ARC-AGI-189.9%estimated ± 2.7 pp, high confidence
ARC-AGI-257.9%estimated ± 10.1 pp, low confidence
CursorBench 3.158.2%estimated ± 5.6 pp, low confidence
CursorBench 3.264.0%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.013.5%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main28.1%estimated ± 2.7 pp, medium confidence
FrontierSWE v224.9%estimated ± 9.0 pp, medium confidence
OSWorld-Verified69.1%estimated ± 5.4 pp, medium confidence
Vals GPQA Diamond89.8%estimated ± 2.0 pp, high confidence
Vals LiveCodeBench86.2%estimated ± 0.8 pp, medium confidence
Vals MMLU-Pro86.9%estimated ± 1.6 pp, high confidence
Vals SWE-bench74.6%estimated ± 5.3 pp, low confidence
Terminal-Bench 2.1 (Vals AI)67.8%estimated ± 7.4 pp, medium confidence
AA AutomationBench5.9%estimated ± 5.7 pp, low confidence
GDP.pdf9.3%estimated ± 6.6 pp, low confidence
AA ITBench46.4%estimated ± 3.6 pp, medium confidence
AA Tau3 Banking25.7%estimated ± 10.2 pp, low confidence
AA Terminal-Bench 2.184.6%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.018.7%estimated ± 8.9 pp, medium confidence
CursorBench 4.039.7%estimated ± 3.6 pp, high confidence
DeepSWE44.4%measured
HLE w/o tools36.2%measured
OSWorld 2.017.7%estimated ± 13.8 pp, low confidence
PostTrainBench v1.129.1%estimated ± 4.7 pp, high confidence
SWE Multilingual78.0%measured
Toolathlon-Verified61.9%estimated ± 7.7 pp, medium confidence
Vibe Code Bench54.1%estimated ± 8.1 pp, medium confidence
ExploitGym0.0%estimated ± 6.1 pp, low confidence
Claw-Eval70.8%estimated ± 5.7 pp, medium confidence
DeepPlanning49.4%estimated ± 13.1 pp, low confidence
HMMT Feb 202598.4%estimated ± 1.9 pp, low confidence
HMMT Feb 202689.8%estimated ± 4.0 pp, high confidence
MCP Atlas78.7%measured
MCP-Tasks77.1%estimated ± 0.6 pp, low confidence
MMAnswerBench87.8%estimated ± 1.3 pp, high confidence
MMLU-ProX85.9%estimated ± 1.7 pp, medium confidence
MMLU-Redux93.5%estimated ± 3.0 pp, medium confidence
NL2Repo49.2%estimated ± 6.1 pp, medium confidence
NOVA-6358.4%estimated ± 1.6 pp, low confidence
QwenClawBench65.0%estimated ± 2.6 pp, medium confidence
SuperGPQA77.2%estimated ± 13.9 pp, low confidence
Toolathlon50.4%estimated ± 3.5 pp, high confidence
LiveCodeBench Pro87.8%estimated ± 6.0 pp, low confidence
MedXpertQA (Text)55.3%estimated ± 4.5 pp, medium confidence
MMLU-Pro (Arcee)87.0%estimated ± 3.4 pp, low confidence
ResearchClawBench19.0%estimated ± 1.8 pp, medium confidence
SWE-Rebench61.8%estimated ± 3.6 pp, medium confidence
ARC-AGI-30.0%estimated ± 3.8 pp, high confidence
FrontierCode 1.1 Extended57.7%estimated ± 3.0 pp, medium confidence
HealthBench Professional51.1%estimated ± 4.8 pp, low confidence
AIME 202598.1%estimated ± 1.5 pp, low confidence
LiveCodeBench87.0%estimated ± 10.2 pp, low confidence
OpenHarmony Bench60.9%estimated ± 4.9 pp, low confidence
MMMLU90.0%estimated ± 1.1 pp, medium confidence
MATH-50098.0%estimated ± 3.1 pp, low confidence