benchgap
Other · model

Laguna S 2.1 benchmark scores

As of 2026-10-07, Laguna S 2.1 (Other) has measured scores on 5 benchmarks and estimated scores on 33 more.

BenchmarkScoreSource
AA-SciCode50.0%estimated ± 3.8 pp, high confidence
GDPval-AA25.7%estimated ± 4.6 pp, medium confidence
LiveCodeBench v689.3%estimated ± 2.1 pp, high confidence
SciCode51.1%estimated ± 4.8 pp, high confidence
Terminal-Bench 2.170.2%measured
BrowseComp71.7%estimated ± 6.4 pp, medium confidence
HLE w/ tools38.3%estimated ± 5.9 pp, medium confidence
AA Coding Index67.0%estimated ± 2.5 pp, medium confidence
SWE-bench Verified82.0%estimated ± 2.5 pp, high confidence
CyberGym22.5%estimated ± 6.0 pp, low confidence
JobBench11.3%estimated ± 3.4 pp, low confidence
SWE-bench Pro59.4%measured
WideResearch69.1%estimated ± 3.9 pp, medium confidence
AA Agentic Index18.8%estimated ± 3.3 pp, low confidence
AA-AnalystAgent5.9%estimated ± 6.7 pp, low confidence
CursorBench 3.154.2%estimated ± 5.6 pp, low confidence
CursorBench 3.257.3%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.08.2%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main23.0%estimated ± 3.7 pp, low confidence
FrontierSWE v214.1%estimated ± 9.0 pp, medium confidence
Vals LiveCodeBench84.6%estimated ± 4.7 pp, medium confidence
Vals SWE-bench81.2%estimated ± 7.5 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)60.2%estimated ± 7.4 pp, medium confidence
AA Terminal-Bench 2.175.2%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.07.4%estimated ± 8.9 pp, medium confidence
CursorBench 4.035.0%estimated ± 3.6 pp, medium confidence
DeepSWE40.4%measured
OSWorld 2.019.4%estimated ± 5.5 pp, low confidence
PostTrainBench v1.125.3%estimated ± 4.7 pp, high confidence
SWE Multilingual78.5%measured
Toolathlon-Verified49.7%measured
MCP Atlas67.4%estimated ± 7.8 pp, medium confidence
MMLU-ProX86.1%estimated ± 1.7 pp, low confidence
NL2Repo46.8%estimated ± 6.4 pp, medium confidence
NOVA-6358.6%estimated ± 1.6 pp, low confidence
LiveCodeBench Pro83.3%estimated ± 6.0 pp, low confidence
FrontierCode 1.1 Extended55.8%estimated ± 3.0 pp, low confidence
OpenHarmony Bench53.3%estimated ± 4.9 pp, low confidence