Other · model
Ornith-1.0-35B benchmark scores
As of 2026-10-07, Ornith-1.0-35B (Other) has measured scores on 6 benchmarks and estimated scores on 35 more.
| Benchmark | Score | Source |
|---|---|---|
| AA-SciCode | 41.8% | estimated ± 3.0 pp, high confidence |
| GDPval-AA | 27.9% | estimated ± 9.4 pp, low confidence |
| LiveCodeBench v6 | 85.0% | estimated ± 2.1 pp, high confidence |
| SciCode | 46.5% | estimated ± 3.8 pp, high confidence |
| Terminal-Bench 2.1 | 64.2% | measured |
| BrowseComp | 76.9% | estimated ± 9.4 pp, low confidence |
| VITA-Bench | 46.9% | estimated ± 10.7 pp, low confidence |
| AA Coding Index | 48.8% | estimated ± 4.6 pp, high confidence |
| SWE-bench Verified | 75.6% | measured |
| BFCL v4 | 83.5% | estimated ± 8.6 pp, low confidence |
| JobBench | 40.1% | estimated ± 8.4 pp, low confidence |
| SWE-bench Pro | 50.4% | measured |
| AA Agentic Index | 23.4% | estimated ± 5.1 pp, medium confidence |
| Gert Labs | 58.2% | estimated ± 9.1 pp, low confidence |
| CursorBench 3.1 | 48.0% | estimated ± 5.6 pp, low confidence |
| CursorBench 3.2 | 41.1% | estimated ± 3.6 pp, medium confidence |
| Terminal-Bench 3.0 | 6.3% | estimated ± 9.1 pp, low confidence |
| FrontierCode 1.1 Main | 12.1% | estimated ± 2.7 pp, low confidence |
| FrontierSWE v2 | 0.0% | estimated ± 9.0 pp, low confidence |
| Vals LiveCodeBench | 77.7% | estimated ± 4.7 pp, medium confidence |
| Vals SWE-bench | 66.8% | estimated ± 5.4 pp, low confidence |
| Terminal-Bench 2.1 (Vals AI) | 55.6% | estimated ± 7.4 pp, medium confidence |
| AA Terminal-Bench 2.1 | 58.7% | estimated ± 3.5 pp, high confidence |
| AA Terminal-Bench 4.0 | 1.1% | estimated ± 8.9 pp, medium confidence |
| CursorBench 4.0 | 28.1% | estimated ± 3.6 pp, medium confidence |
| DeepSWE | 30.3% | estimated ± 9.6 pp, low confidence |
| PostTrainBench v1.1 | 20.2% | estimated ± 4.7 pp, medium confidence |
| SWE Multilingual | 69.3% | measured |
| Vibe Code Bench | 5.1% | estimated ± 8.1 pp, medium confidence |
| Claw-Eval | 69.8% | measured |
| MCP Atlas | 69.3% | estimated ± 11.2 pp, low confidence |
| MCP-Tasks | 92.5% | estimated ± 6.4 pp, low confidence |
| MMLU-ProX | 82.6% | estimated ± 1.7 pp, medium confidence |
| NL2Repo | 34.6% | measured |
| NOVA-63 | 54.6% | estimated ± 1.6 pp, low confidence |
| LiveCodeBench Pro | 75.7% | estimated ± 6.0 pp, low confidence |
| ResearchClawBench | 18.8% | estimated ± 1.3 pp, high confidence |
| SWE-Rebench | 57.4% | estimated ± 3.6 pp, medium confidence |
| FrontierCode 1.1 Extended | 52.5% | estimated ± 3.0 pp, low confidence |
| LiveCodeBench | 83.7% | estimated ± 10.2 pp, low confidence |
| OpenHarmony Bench | 46.4% | estimated ± 2.7 pp, medium confidence |