benchgap
Other · model

Ornith-1.0-9B benchmark scores

As of 2026-10-07, Ornith-1.0-9B (Other) has measured scores on 6 benchmarks and estimated scores on 35 more.

BenchmarkScoreSource
AA-SciCode34.6%estimated ± 3.0 pp, medium confidence
GDPval-AA24.0%estimated ± 9.4 pp, low confidence
LiveCodeBench v681.9%estimated ± 2.1 pp, high confidence
SciCode43.5%estimated ± 3.8 pp, high confidence
Terminal-Bench 2.143.1%measured
BrowseComp72.2%estimated ± 9.4 pp, low confidence
VITA-Bench40.3%estimated ± 10.7 pp, low confidence
AA Coding Index39.6%estimated ± 4.6 pp, medium confidence
SWE-bench Verified69.4%measured
BFCL v464.6%estimated ± 8.6 pp, low confidence
JobBench32.8%estimated ± 8.4 pp, low confidence
SWE-bench Pro42.9%measured
AA Agentic Index20.1%estimated ± 5.1 pp, medium confidence
Gert Labs54.1%estimated ± 9.1 pp, low confidence
CursorBench 3.142.4%estimated ± 5.6 pp, low confidence
CursorBench 3.223.7%estimated ± 3.6 pp, medium confidence
Terminal-Bench 3.02.7%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main3.2%estimated ± 2.7 pp, low confidence
FrontierSWE v20.0%estimated ± 9.0 pp, low confidence
Vals LiveCodeBench64.1%estimated ± 4.7 pp, low confidence
Vals SWE-bench56.5%estimated ± 5.4 pp, low confidence
Terminal-Bench 2.1 (Vals AI)39.3%estimated ± 7.4 pp, low confidence
AA Terminal-Bench 2.152.1%estimated ± 3.5 pp, medium confidence
AA Terminal-Bench 4.00.0%estimated ± 8.9 pp, low confidence
CursorBench 4.022.4%estimated ± 3.6 pp, medium confidence
DeepSWE22.2%estimated ± 9.6 pp, low confidence
PostTrainBench v1.116.3%estimated ± 4.7 pp, medium confidence
SWE Multilingual52.0%measured
Vibe Code Bench0.7%estimated ± 8.1 pp, medium confidence
Claw-Eval63.1%measured
MCP Atlas58.5%estimated ± 11.2 pp, low confidence
MCP-Tasks79.0%estimated ± 6.4 pp, low confidence
MMLU-ProX75.5%estimated ± 1.7 pp, low confidence
NL2Repo27.2%measured
NOVA-6345.7%estimated ± 1.6 pp, low confidence
LiveCodeBench Pro68.6%estimated ± 6.0 pp, low confidence
ResearchClawBench17.2%estimated ± 1.3 pp, high confidence
SWE-Rebench57.0%estimated ± 3.6 pp, medium confidence
FrontierCode 1.1 Extended49.1%estimated ± 3.0 pp, low confidence
LiveCodeBench79.6%estimated ± 10.2 pp, low confidence
OpenHarmony Bench40.7%estimated ± 2.7 pp, medium confidence