benchgap
Other · model

Ornith-1.5-397B benchmark scores

As of 2026-10-07, Ornith-1.5-397B (Other) has measured scores on 17 benchmarks and estimated scores on 71 more.

BenchmarkScoreSource
AA-GPQA Diamond91.2%estimated ± 2.7 pp, high confidence
AA-HLE43.1%estimated ± 4.0 pp, high confidence
AA-SciCode52.8%estimated ± 3.0 pp, high confidence
Artificial Analysis Intelligence Index38.3%estimated ± 5.1 pp, medium confidence
CritPt18.8%estimated ± 4.1 pp, high confidence
GDPval-AA24.6%estimated ± 3.7 pp, low confidence
GPQA Diamond92.8%measured
HLE44.6%measured
LiveCodeBench v690.7%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy45.4%estimated ± 8.1 pp, medium confidence
SciCode51.2%estimated ± 3.8 pp, medium confidence
Terminal-Bench 2.186.1%measured
BrowseComp86.6%measured
HLE w/ tools56.1%measured
VITA-Bench48.2%estimated ± 9.9 pp, low confidence
AA Coding Index71.8%estimated ± 2.5 pp, high confidence
APEX-Agents33.9%estimated ± 0.6 pp, medium confidence
APEX-Agents-AA35.6%estimated ± 4.6 pp, high confidence
DeepSearchQA89.6%estimated ± 8.1 pp, low confidence
SWE-bench Verified86.0%measured
AutomationBench34.8%estimated ± 10.3 pp, low confidence
BFCL v4100.0%estimated ± 8.6 pp, low confidence
CyberGym79.2%estimated ± 2.4 pp, medium confidence
JobBench52.3%estimated ± 3.4 pp, low confidence
SkillsBench61.3%estimated ± 5.6 pp, low confidence
SWE-bench Pro65.1%measured
τ³-bench results69.6%estimated ± 1.8 pp, low confidence
WideResearch80.8%measured
AA Agentic Index46.8%estimated ± 3.3 pp, medium confidence
MMLU-Pro85.7%estimated ± 2.5 pp, low confidence
GPQA92.8%measured
Gert Labs55.5%estimated ± 6.1 pp, low confidence
AA-AnalystAgent49.6%estimated ± 2.5 pp, medium confidence
AA EnterpriseOps-Gym43.2%estimated ± 6.9 pp, low confidence
AA Harvey LAB89.8%estimated ± 1.7 pp, medium confidence
ARC-AGI-193.1%estimated ± 2.7 pp, high confidence
ARC-AGI-270.1%estimated ± 10.1 pp, low confidence
CursorBench 3.157.9%estimated ± 5.6 pp, low confidence
CursorBench 3.263.7%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.013.5%measured
FrontierCode 1.1 Main43.1%estimated ± 2.7 pp, medium confidence
FrontierSWE v224.2%estimated ± 9.0 pp, medium confidence
OSWorld-Verified80.9%estimated ± 5.4 pp, medium confidence
Vals GPQA Diamond92.1%estimated ± 2.0 pp, high confidence
Vals LiveCodeBench86.5%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro88.1%estimated ± 1.2 pp, high confidence
Vals SWE-bench92.6%estimated ± 5.4 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)75.0%estimated ± 3.4 pp, high confidence
AA AutomationBench37.9%estimated ± 5.7 pp, low confidence
GDP.pdf10.6%estimated ± 6.6 pp, low confidence
AA ITBench50.8%estimated ± 3.4 pp, medium confidence
AA Tau3 Banking34.3%estimated ± 10.2 pp, low confidence
AA Terminal-Bench 2.184.9%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.026.1%estimated ± 8.9 pp, medium confidence
CursorBench 4.039.4%estimated ± 3.6 pp, high confidence
DeepSWE56.0%measured
HLE w/o tools44.6%measured
OSWorld 2.019.4%estimated ± 5.5 pp, low confidence
PostTrainBench v1.128.9%estimated ± 4.7 pp, high confidence
SWE Multilingual79.6%measured
Toolathlon-Verified71.2%measured
Vibe Code Bench62.3%estimated ± 8.1 pp, low confidence
ExploitGym20.8%estimated ± 6.1 pp, low confidence
C-Eval93.3%estimated ± 0.8 pp, low confidence
Claw-Eval81.4%measured
DeepPlanning50.1%estimated ± 13.1 pp, low confidence
MCP Atlas80.0%measured
MCP-Tasks77.1%estimated ± 0.6 pp, low confidence
MMLU-ProX86.5%estimated ± 1.7 pp, low confidence
MMLU-Redux96.3%estimated ± 2.1 pp, low confidence
NL2Repo59.5%measured
NOVA-6359.1%estimated ± 1.6 pp, low confidence
QwenClawBench65.9%estimated ± 2.6 pp, medium confidence
SuperGPQA80.5%estimated ± 9.1 pp, medium confidence
Toolathlon55.4%estimated ± 3.5 pp, high confidence
HealthBench Hard34.4%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro87.5%estimated ± 6.0 pp, low confidence
MedXpertQA (Text)63.2%estimated ± 4.5 pp, medium confidence
MMLU-Pro (Arcee)87.8%estimated ± 3.4 pp, low confidence
ResearchClawBench21.6%estimated ± 1.3 pp, medium confidence
SWE-Rebench61.8%estimated ± 3.6 pp, low confidence
ARC-AGI-30.2%estimated ± 2.1 pp, high confidence
FrontierCode 1.1 Extended57.6%estimated ± 3.0 pp, medium confidence
HealthBench Professional55.6%estimated ± 4.8 pp, medium confidence
LiveCodeBench90.1%estimated ± 10.2 pp, low confidence
OpenHarmony Bench59.3%estimated ± 2.7 pp, high confidence
MMMLU90.7%estimated ± 1.1 pp, low confidence
MMLU90.5%estimated ± 3.6 pp, low confidence