benchgap
Other · model

Atria Dawn Preview benchmark scores

As of 2026-10-07, Atria Dawn Preview (Other) has measured scores on 11 benchmarks and estimated scores on 46 more.

BenchmarkScoreSource
AA-SciCode50.1%estimated ± 3.8 pp, high confidence
GDPval-AA24.8%estimated ± 3.7 pp, low confidence
LiveCodeBench v689.4%estimated ± 2.1 pp, high confidence
SciCode51.2%estimated ± 4.8 pp, high confidence
Terminal-Bench 2.178.3%measured
BrowseComp92.5%measured
HLE w/ tools55.6%estimated ± 2.4 pp, medium confidence
VITA-Bench46.6%estimated ± 9.9 pp, low confidence
τ²-bench results94.8%estimated ± 14.1 pp, low confidence
AA Coding Index64.8%estimated ± 7.5 pp, medium confidence
APEX-Agents-AA38.3%estimated ± 3.6 pp, medium confidence
DeepSearchQA96.0%measured
SWE-bench Verified82.2%estimated ± 2.5 pp, high confidence
AutomationBench53.8%measured
BFCL v477.0%measured
CyberGym86.5%measured
JobBench50.3%measured
SkillsBench66.4%measured
SWE-bench Pro59.6%measured
τ³-bench results41.2%measured
WideResearch81.9%measured
AA Agentic Index52.5%estimated ± 5.8 pp, low confidence
Gert Labs70.3%estimated ± 6.0 pp, low confidence
AA-AnalystAgent47.4%estimated ± 12.5 pp, low confidence
AA EnterpriseOps-Gym30.0%estimated ± 5.1 pp, low confidence
AA Harvey LAB93.5%estimated ± 9.8 pp, low confidence
ApprenticeBench29.5%estimated ± 4.0 pp, low confidence
CursorBench 3.154.4%estimated ± 5.6 pp, low confidence
CursorBench 3.257.6%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.012.2%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main39.9%estimated ± 5.5 pp, medium confidence
FrontierSWE v214.4%estimated ± 9.0 pp, medium confidence
OSWorld-Verified80.4%estimated ± 4.7 pp, high confidence
Vals LiveCodeBench84.7%estimated ± 4.7 pp, medium confidence
Vals SWE-bench81.4%estimated ± 7.5 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)66.4%estimated ± 7.4 pp, medium confidence
AA AutomationBench53.6%estimated ± 5.0 pp, medium confidence
GDP.pdf17.8%estimated ± 5.2 pp, low confidence
AA Tau3 Banking41.4%estimated ± 10.2 pp, low confidence
AA Terminal-Bench 2.184.3%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.016.5%estimated ± 8.9 pp, medium confidence
CursorBench 4.035.2%estimated ± 3.6 pp, medium confidence
OSWorld 2.068.6%estimated ± 13.8 pp, low confidence
PostTrainBench v1.125.4%estimated ± 4.7 pp, high confidence
Toolathlon-Verified73.2%estimated ± 1.3 pp, low confidence
ExploitGym39.9%estimated ± 6.1 pp, low confidence
Claw-Eval64.7%estimated ± 2.3 pp, low confidence
DeepPlanning13.0%estimated ± 8.3 pp, low confidence
MCP Atlas78.1%estimated ± 8.2 pp, medium confidence
MCP-Tasks7.3%estimated ± 3.6 pp, low confidence
NL2Repo47.0%estimated ± 6.4 pp, medium confidence
Toolathlon57.3%estimated ± 3.5 pp, medium confidence
LiveCodeBench Pro83.4%estimated ± 6.0 pp, low confidence
ResearchClawBench20.0%estimated ± 2.2 pp, low confidence
FrontierCode 1.1 Extended55.9%estimated ± 3.0 pp, low confidence
Agents' Last Exam30.3%estimated ± 2.1 pp, high confidence
OpenHarmony Bench53.4%estimated ± 4.9 pp, low confidence