Other · model
Ornith-1.0-9B benchmark scores
As of 2026-10-07, Ornith-1.0-9B (Other) has measured scores on 6 benchmarks and estimated scores on 35 more.
| Benchmark | Score | Source |
|---|---|---|
| AA-SciCode | 34.6% | estimated ± 3.0 pp, medium confidence |
| GDPval-AA | 24.0% | estimated ± 9.4 pp, low confidence |
| LiveCodeBench v6 | 81.9% | estimated ± 2.1 pp, high confidence |
| SciCode | 43.5% | estimated ± 3.8 pp, high confidence |
| Terminal-Bench 2.1 | 43.1% | measured |
| BrowseComp | 72.2% | estimated ± 9.4 pp, low confidence |
| VITA-Bench | 40.3% | estimated ± 10.7 pp, low confidence |
| AA Coding Index | 39.6% | estimated ± 4.6 pp, medium confidence |
| SWE-bench Verified | 69.4% | measured |
| BFCL v4 | 64.6% | estimated ± 8.6 pp, low confidence |
| JobBench | 32.8% | estimated ± 8.4 pp, low confidence |
| SWE-bench Pro | 42.9% | measured |
| AA Agentic Index | 20.1% | estimated ± 5.1 pp, medium confidence |
| Gert Labs | 54.1% | estimated ± 9.1 pp, low confidence |
| CursorBench 3.1 | 42.4% | estimated ± 5.6 pp, low confidence |
| CursorBench 3.2 | 23.7% | estimated ± 3.6 pp, medium confidence |
| Terminal-Bench 3.0 | 2.7% | estimated ± 9.1 pp, low confidence |
| FrontierCode 1.1 Main | 3.2% | estimated ± 2.7 pp, low confidence |
| FrontierSWE v2 | 0.0% | estimated ± 9.0 pp, low confidence |
| Vals LiveCodeBench | 64.1% | estimated ± 4.7 pp, low confidence |
| Vals SWE-bench | 56.5% | estimated ± 5.4 pp, low confidence |
| Terminal-Bench 2.1 (Vals AI) | 39.3% | estimated ± 7.4 pp, low confidence |
| AA Terminal-Bench 2.1 | 52.1% | estimated ± 3.5 pp, medium confidence |
| AA Terminal-Bench 4.0 | 0.0% | estimated ± 8.9 pp, low confidence |
| CursorBench 4.0 | 22.4% | estimated ± 3.6 pp, medium confidence |
| DeepSWE | 22.2% | estimated ± 9.6 pp, low confidence |
| PostTrainBench v1.1 | 16.3% | estimated ± 4.7 pp, medium confidence |
| SWE Multilingual | 52.0% | measured |
| Vibe Code Bench | 0.7% | estimated ± 8.1 pp, medium confidence |
| Claw-Eval | 63.1% | measured |
| MCP Atlas | 58.5% | estimated ± 11.2 pp, low confidence |
| MCP-Tasks | 79.0% | estimated ± 6.4 pp, low confidence |
| MMLU-ProX | 75.5% | estimated ± 1.7 pp, low confidence |
| NL2Repo | 27.2% | measured |
| NOVA-63 | 45.7% | estimated ± 1.6 pp, low confidence |
| LiveCodeBench Pro | 68.6% | estimated ± 6.0 pp, low confidence |
| ResearchClawBench | 17.2% | estimated ± 1.3 pp, high confidence |
| SWE-Rebench | 57.0% | estimated ± 3.6 pp, medium confidence |
| FrontierCode 1.1 Extended | 49.1% | estimated ± 3.0 pp, low confidence |
| LiveCodeBench | 79.6% | estimated ± 10.2 pp, low confidence |
| OpenHarmony Bench | 40.7% | estimated ± 2.7 pp, medium confidence |