Other · model
Laguna S 2.1 benchmark scores
As of 2026-10-07, Laguna S 2.1 (Other) has measured scores on 5 benchmarks and estimated scores on 33 more.
| Benchmark | Score | Source |
|---|---|---|
| AA-SciCode | 50.0% | estimated ± 3.8 pp, high confidence |
| GDPval-AA | 25.7% | estimated ± 4.6 pp, medium confidence |
| LiveCodeBench v6 | 89.3% | estimated ± 2.1 pp, high confidence |
| SciCode | 51.1% | estimated ± 4.8 pp, high confidence |
| Terminal-Bench 2.1 | 70.2% | measured |
| BrowseComp | 71.7% | estimated ± 6.4 pp, medium confidence |
| HLE w/ tools | 38.3% | estimated ± 5.9 pp, medium confidence |
| AA Coding Index | 67.0% | estimated ± 2.5 pp, medium confidence |
| SWE-bench Verified | 82.0% | estimated ± 2.5 pp, high confidence |
| CyberGym | 22.5% | estimated ± 6.0 pp, low confidence |
| JobBench | 11.3% | estimated ± 3.4 pp, low confidence |
| SWE-bench Pro | 59.4% | measured |
| WideResearch | 69.1% | estimated ± 3.9 pp, medium confidence |
| AA Agentic Index | 18.8% | estimated ± 3.3 pp, low confidence |
| AA-AnalystAgent | 5.9% | estimated ± 6.7 pp, low confidence |
| CursorBench 3.1 | 54.2% | estimated ± 5.6 pp, low confidence |
| CursorBench 3.2 | 57.3% | estimated ± 3.6 pp, high confidence |
| Terminal-Bench 3.0 | 8.2% | estimated ± 9.1 pp, low confidence |
| FrontierCode 1.1 Main | 23.0% | estimated ± 3.7 pp, low confidence |
| FrontierSWE v2 | 14.1% | estimated ± 9.0 pp, medium confidence |
| Vals LiveCodeBench | 84.6% | estimated ± 4.7 pp, medium confidence |
| Vals SWE-bench | 81.2% | estimated ± 7.5 pp, medium confidence |
| Terminal-Bench 2.1 (Vals AI) | 60.2% | estimated ± 7.4 pp, medium confidence |
| AA Terminal-Bench 2.1 | 75.2% | estimated ± 3.5 pp, high confidence |
| AA Terminal-Bench 4.0 | 7.4% | estimated ± 8.9 pp, medium confidence |
| CursorBench 4.0 | 35.0% | estimated ± 3.6 pp, medium confidence |
| DeepSWE | 40.4% | measured |
| OSWorld 2.0 | 19.4% | estimated ± 5.5 pp, low confidence |
| PostTrainBench v1.1 | 25.3% | estimated ± 4.7 pp, high confidence |
| SWE Multilingual | 78.5% | measured |
| Toolathlon-Verified | 49.7% | measured |
| MCP Atlas | 67.4% | estimated ± 7.8 pp, medium confidence |
| MMLU-ProX | 86.1% | estimated ± 1.7 pp, low confidence |
| NL2Repo | 46.8% | estimated ± 6.4 pp, medium confidence |
| NOVA-63 | 58.6% | estimated ± 1.6 pp, low confidence |
| LiveCodeBench Pro | 83.3% | estimated ± 6.0 pp, low confidence |
| FrontierCode 1.1 Extended | 55.8% | estimated ± 3.0 pp, low confidence |
| OpenHarmony Bench | 53.3% | estimated ± 4.9 pp, low confidence |