benchgap
Other · model

Ornith-1.0-35B benchmark scores

As of 2026-10-07, Ornith-1.0-35B (Other) has measured scores on 6 benchmarks and estimated scores on 35 more.

BenchmarkScoreSource
AA-SciCode41.8%estimated ± 3.0 pp, high confidence
GDPval-AA27.9%estimated ± 9.4 pp, low confidence
LiveCodeBench v685.0%estimated ± 2.1 pp, high confidence
SciCode46.5%estimated ± 3.8 pp, high confidence
Terminal-Bench 2.164.2%measured
BrowseComp76.9%estimated ± 9.4 pp, low confidence
VITA-Bench46.9%estimated ± 10.7 pp, low confidence
AA Coding Index48.8%estimated ± 4.6 pp, high confidence
SWE-bench Verified75.6%measured
BFCL v483.5%estimated ± 8.6 pp, low confidence
JobBench40.1%estimated ± 8.4 pp, low confidence
SWE-bench Pro50.4%measured
AA Agentic Index23.4%estimated ± 5.1 pp, medium confidence
Gert Labs58.2%estimated ± 9.1 pp, low confidence
CursorBench 3.148.0%estimated ± 5.6 pp, low confidence
CursorBench 3.241.1%estimated ± 3.6 pp, medium confidence
Terminal-Bench 3.06.3%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main12.1%estimated ± 2.7 pp, low confidence
FrontierSWE v20.0%estimated ± 9.0 pp, low confidence
Vals LiveCodeBench77.7%estimated ± 4.7 pp, medium confidence
Vals SWE-bench66.8%estimated ± 5.4 pp, low confidence
Terminal-Bench 2.1 (Vals AI)55.6%estimated ± 7.4 pp, medium confidence
AA Terminal-Bench 2.158.7%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.01.1%estimated ± 8.9 pp, medium confidence
CursorBench 4.028.1%estimated ± 3.6 pp, medium confidence
DeepSWE30.3%estimated ± 9.6 pp, low confidence
PostTrainBench v1.120.2%estimated ± 4.7 pp, medium confidence
SWE Multilingual69.3%measured
Vibe Code Bench5.1%estimated ± 8.1 pp, medium confidence
Claw-Eval69.8%measured
MCP Atlas69.3%estimated ± 11.2 pp, low confidence
MCP-Tasks92.5%estimated ± 6.4 pp, low confidence
MMLU-ProX82.6%estimated ± 1.7 pp, medium confidence
NL2Repo34.6%measured
NOVA-6354.6%estimated ± 1.6 pp, low confidence
LiveCodeBench Pro75.7%estimated ± 6.0 pp, low confidence
ResearchClawBench18.8%estimated ± 1.3 pp, high confidence
SWE-Rebench57.4%estimated ± 3.6 pp, medium confidence
FrontierCode 1.1 Extended52.5%estimated ± 3.0 pp, low confidence
LiveCodeBench83.7%estimated ± 10.2 pp, low confidence
OpenHarmony Bench46.4%estimated ± 2.7 pp, medium confidence