benchgap
Other · model

Ornith-1.0-397B benchmark scores

As of 2026-10-07, Ornith-1.0-397B (Other) has measured scores on 6 benchmarks and estimated scores on 35 more.

BenchmarkScoreSource
AA-SciCode50.2%estimated ± 3.0 pp, high confidence
GDPval-AA32.7%estimated ± 9.4 pp, low confidence
LiveCodeBench v689.6%estimated ± 2.1 pp, high confidence
SciCode49.6%estimated ± 3.8 pp, medium confidence
Terminal-Bench 2.177.5%measured
BrowseComp81.6%estimated ± 9.4 pp, low confidence
VITA-Bench54.9%estimated ± 10.7 pp, low confidence
AA Coding Index66.9%estimated ± 4.6 pp, high confidence
SWE-bench Verified82.4%measured
BFCL v4100.0%estimated ± 8.6 pp, low confidence
JobBench45.5%estimated ± 8.4 pp, low confidence
SWE-bench Pro62.2%measured
AA Agentic Index27.4%estimated ± 5.1 pp, low confidence
Gert Labs62.4%estimated ± 9.1 pp, low confidence
CursorBench 3.156.1%estimated ± 5.6 pp, low confidence
CursorBench 3.260.8%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.011.7%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main33.1%estimated ± 2.7 pp, medium confidence
FrontierSWE v219.1%estimated ± 9.0 pp, medium confidence
Vals LiveCodeBench85.7%estimated ± 4.7 pp, medium confidence
Vals SWE-bench82.2%estimated ± 5.4 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)65.8%estimated ± 7.4 pp, medium confidence
AA Terminal-Bench 2.184.1%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.015.6%estimated ± 8.9 pp, medium confidence
CursorBench 4.037.2%estimated ± 3.6 pp, medium confidence
DeepSWE47.2%estimated ± 9.6 pp, medium confidence
PostTrainBench v1.127.0%estimated ± 4.7 pp, high confidence
SWE Multilingual78.9%measured
Vibe Code Bench59.8%estimated ± 8.1 pp, medium confidence
Claw-Eval77.1%measured
MCP Atlas77.0%estimated ± 11.2 pp, low confidence
MCP-Tasks100.0%estimated ± 6.4 pp, low confidence
MMLU-ProX86.3%estimated ± 1.7 pp, low confidence
NL2Repo48.2%measured
NOVA-6358.8%estimated ± 1.6 pp, low confidence
LiveCodeBench Pro85.4%estimated ± 6.0 pp, low confidence
ResearchClawBench20.6%estimated ± 1.3 pp, medium confidence
SWE-Rebench61.8%estimated ± 3.6 pp, medium confidence
FrontierCode 1.1 Extended56.7%estimated ± 3.0 pp, low confidence
LiveCodeBench87.9%estimated ± 10.2 pp, low confidence
OpenHarmony Bench54.4%estimated ± 2.7 pp, high confidence