benchgap
Other · model

Ornith-1.5-9B benchmark scores

As of 2026-10-07, Ornith-1.5-9B (Other) has measured scores on 15 benchmarks and estimated scores on 73 more.

BenchmarkScoreSource
AA-GPQA Diamond83.9%estimated ± 2.7 pp, high confidence
AA-HLE27.9%estimated ± 4.0 pp, high confidence
AA-SciCode39.7%estimated ± 3.0 pp, high confidence
Artificial Analysis Intelligence Index19.0%estimated ± 5.1 pp, medium confidence
CritPt2.6%estimated ± 4.1 pp, high confidence
GDPval-AA18.7%estimated ± 3.7 pp, low confidence
GPQA Diamond86.4%measured
HLE20.2%measured
LiveCodeBench v683.8%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy26.1%estimated ± 8.1 pp, medium confidence
SciCode44.1%estimated ± 3.8 pp, high confidence
Terminal-Bench 2.146.2%measured
BrowseComp56.4%measured
HLE w/ tools30.5%measured
VITA-Bench36.8%estimated ± 9.9 pp, low confidence
AA Coding Index45.5%estimated ± 4.6 pp, high confidence
APEX-Agents13.2%estimated ± 0.6 pp, low confidence
APEX-Agents-AA8.8%estimated ± 4.6 pp, medium confidence
DeepSearchQA76.3%estimated ± 8.1 pp, low confidence
SWE-bench Verified70.6%measured
AutomationBench15.2%estimated ± 10.3 pp, low confidence
BFCL v474.2%estimated ± 8.6 pp, low confidence
CyberGym76.3%estimated ± 2.4 pp, low confidence
JobBench4.0%estimated ± 3.4 pp, low confidence
SkillsBench17.1%estimated ± 5.6 pp, low confidence
SWE-bench Pro47.5%measured
τ³-bench results69.5%estimated ± 1.8 pp, medium confidence
WideResearch59.5%measured
AA Agentic Index7.7%estimated ± 3.3 pp, low confidence
MMLU-Pro84.4%estimated ± 2.5 pp, medium confidence
GPQA86.4%measured
Gert Labs44.1%estimated ± 6.1 pp, low confidence
AA-AnalystAgent4.5%estimated ± 2.5 pp, low confidence
AA EnterpriseOps-Gym32.7%estimated ± 6.9 pp, low confidence
AA Harvey LAB77.1%estimated ± 1.7 pp, low confidence
ARC-AGI-183.2%estimated ± 2.7 pp, medium confidence
ARC-AGI-243.5%estimated ± 10.1 pp, low confidence
CursorBench 3.145.9%estimated ± 5.6 pp, low confidence
CursorBench 3.234.5%estimated ± 3.6 pp, medium confidence
Terminal-Bench 3.03.1%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main4.2%estimated ± 2.7 pp, low confidence
FrontierSWE v20.0%estimated ± 9.0 pp, low confidence
OSWorld-Verified50.5%estimated ± 5.4 pp, low confidence
Vals GPQA Diamond85.6%estimated ± 2.0 pp, high confidence
Vals LiveCodeBench73.6%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro83.7%estimated ± 1.2 pp, high confidence
Vals SWE-bench63.9%estimated ± 5.4 pp, low confidence
Terminal-Bench 2.1 (Vals AI)41.6%estimated ± 7.4 pp, low confidence
AA AutomationBench2.6%estimated ± 5.7 pp, low confidence
GDP.pdf9.3%estimated ± 6.6 pp, low confidence
AA ITBench73.2%estimated ± 3.4 pp, low confidence
AA Tau3 Banking16.5%estimated ± 10.2 pp, low confidence
AA Terminal-Bench 2.152.1%estimated ± 3.5 pp, medium confidence
AA Terminal-Bench 4.00.0%estimated ± 8.9 pp, low confidence
CursorBench 4.025.9%estimated ± 3.6 pp, medium confidence
DeepSWE27.8%estimated ± 9.6 pp, low confidence
HLE w/o tools20.2%measured
OSWorld 2.019.4%estimated ± 5.5 pp, low confidence
PostTrainBench v1.118.6%estimated ± 4.7 pp, medium confidence
SWE Multilingual54.4%measured
Toolathlon-Verified41.2%measured
Vibe Code Bench0.7%estimated ± 8.1 pp, medium confidence
ExploitGym0.0%estimated ± 6.1 pp, low confidence
C-Eval88.7%estimated ± 0.8 pp, low confidence
Claw-Eval66.5%measured
DeepPlanning36.9%estimated ± 13.1 pp, low confidence
MCP Atlas54.2%measured
MCP-Tasks75.8%estimated ± 0.6 pp, low confidence
MMLU-ProX76.5%estimated ± 1.7 pp, low confidence
MMLU-Redux92.5%estimated ± 2.1 pp, medium confidence
NL2Repo32.4%measured
NOVA-6347.0%estimated ± 1.6 pp, low confidence
QwenClawBench54.6%estimated ± 2.6 pp, high confidence
SuperGPQA61.1%estimated ± 9.1 pp, medium confidence
Toolathlon24.6%estimated ± 3.5 pp, medium confidence
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro73.1%estimated ± 6.0 pp, low confidence
MedXpertQA (Text)44.6%estimated ± 4.5 pp, low confidence
MMLU-Pro (Arcee)85.6%estimated ± 3.4 pp, medium confidence
ResearchClawBench18.0%estimated ± 1.3 pp, high confidence
SWE-Rebench57.0%estimated ± 3.6 pp, medium confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
FrontierCode 1.1 Extended51.2%estimated ± 3.0 pp, low confidence
HealthBench Professional43.2%estimated ± 4.8 pp, low confidence
LiveCodeBench80.4%estimated ± 10.2 pp, low confidence
OpenHarmony Bench44.8%estimated ± 2.7 pp, medium confidence
MMMLU88.4%estimated ± 1.1 pp, medium confidence
MMLU90.3%estimated ± 3.6 pp, low confidence