Other · model
Ornith-1.0-397B benchmark scores
As of 2026-10-07, Ornith-1.0-397B (Other) has measured scores on 6 benchmarks and estimated scores on 35 more.
| Benchmark | Score | Source |
|---|---|---|
| AA-SciCode | 50.2% | estimated ± 3.0 pp, high confidence |
| GDPval-AA | 32.7% | estimated ± 9.4 pp, low confidence |
| LiveCodeBench v6 | 89.6% | estimated ± 2.1 pp, high confidence |
| SciCode | 49.6% | estimated ± 3.8 pp, medium confidence |
| Terminal-Bench 2.1 | 77.5% | measured |
| BrowseComp | 81.6% | estimated ± 9.4 pp, low confidence |
| VITA-Bench | 54.9% | estimated ± 10.7 pp, low confidence |
| AA Coding Index | 66.9% | estimated ± 4.6 pp, high confidence |
| SWE-bench Verified | 82.4% | measured |
| BFCL v4 | 100.0% | estimated ± 8.6 pp, low confidence |
| JobBench | 45.5% | estimated ± 8.4 pp, low confidence |
| SWE-bench Pro | 62.2% | measured |
| AA Agentic Index | 27.4% | estimated ± 5.1 pp, low confidence |
| Gert Labs | 62.4% | estimated ± 9.1 pp, low confidence |
| CursorBench 3.1 | 56.1% | estimated ± 5.6 pp, low confidence |
| CursorBench 3.2 | 60.8% | estimated ± 3.6 pp, high confidence |
| Terminal-Bench 3.0 | 11.7% | estimated ± 9.1 pp, low confidence |
| FrontierCode 1.1 Main | 33.1% | estimated ± 2.7 pp, medium confidence |
| FrontierSWE v2 | 19.1% | estimated ± 9.0 pp, medium confidence |
| Vals LiveCodeBench | 85.7% | estimated ± 4.7 pp, medium confidence |
| Vals SWE-bench | 82.2% | estimated ± 5.4 pp, medium confidence |
| Terminal-Bench 2.1 (Vals AI) | 65.8% | estimated ± 7.4 pp, medium confidence |
| AA Terminal-Bench 2.1 | 84.1% | estimated ± 3.5 pp, high confidence |
| AA Terminal-Bench 4.0 | 15.6% | estimated ± 8.9 pp, medium confidence |
| CursorBench 4.0 | 37.2% | estimated ± 3.6 pp, medium confidence |
| DeepSWE | 47.2% | estimated ± 9.6 pp, medium confidence |
| PostTrainBench v1.1 | 27.0% | estimated ± 4.7 pp, high confidence |
| SWE Multilingual | 78.9% | measured |
| Vibe Code Bench | 59.8% | estimated ± 8.1 pp, medium confidence |
| Claw-Eval | 77.1% | measured |
| MCP Atlas | 77.0% | estimated ± 11.2 pp, low confidence |
| MCP-Tasks | 100.0% | estimated ± 6.4 pp, low confidence |
| MMLU-ProX | 86.3% | estimated ± 1.7 pp, low confidence |
| NL2Repo | 48.2% | measured |
| NOVA-63 | 58.8% | estimated ± 1.6 pp, low confidence |
| LiveCodeBench Pro | 85.4% | estimated ± 6.0 pp, low confidence |
| ResearchClawBench | 20.6% | estimated ± 1.3 pp, medium confidence |
| SWE-Rebench | 61.8% | estimated ± 3.6 pp, medium confidence |
| FrontierCode 1.1 Extended | 56.7% | estimated ± 3.0 pp, low confidence |
| LiveCodeBench | 87.9% | estimated ± 10.2 pp, low confidence |
| OpenHarmony Bench | 54.4% | estimated ± 2.7 pp, high confidence |