benchgap
Other · model

Laguna M.1 benchmark scores

As of 2026-10-07, Laguna M.1 (Other) has measured scores on 9 benchmarks and estimated scores on 41 more.

BenchmarkScoreSource
AA-GPQA Diamond68.7%estimated ± 4.2 pp, medium confidence
AA-HLE10.7%estimated ± 7.4 pp, low confidence
AA-SciCode44.7%estimated ± 3.8 pp, high confidence
Artificial Analysis Intelligence Index16.8%estimated ± 7.9 pp, low confidence
CritPt2.6%estimated ± 6.6 pp, low confidence
GPQA Diamond39.9%estimated ± 1.7 pp, medium confidence
HLE15.4%estimated ± 9.8 pp, low confidence
LiveCodeBench v684.5%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy21.7%estimated ± 7.7 pp, low confidence
SciCode17.3%estimated ± 2.9 pp, low confidence
Terminal-Bench 2.116.9%estimated ± 6.9 pp, low confidence
AA Coding Index49.6%estimated ± 5.1 pp, medium confidence
SWE-bench Verified74.6%measured
SWE-bench Pro49.2%measured
GPQA73.4%estimated ± 1.6 pp, medium confidence
ARC-AGI-10.2%estimated ± 5.9 pp, low confidence
ARC-AGI-20.0%estimated ± 12.6 pp, low confidence
CursorBench 3.147.1%estimated ± 5.6 pp, low confidence
CursorBench 3.238.4%estimated ± 3.6 pp, medium confidence
Terminal-Bench 3.00.2%estimated ± 6.9 pp, low confidence
FrontierCode 1.1 Main9.9%estimated ± 2.7 pp, low confidence
FrontierSWE v20.0%estimated ± 9.0 pp, low confidence
Vals GPQA Diamond27.0%measured
Vals LiveCodeBench68.1%measured
Vals MMLU-Pro68.8%measured
Vals SWE-bench57.6%measured
Terminal-Bench 2.1 (Vals AI)34.1%measured
VulcanBench v377.6%estimated ± 5.6 pp, low confidence
AA Terminal-Bench 2.114.0%estimated ± 4.4 pp, medium confidence
AA Terminal-Bench 4.00.0%estimated ± 14.0 pp, low confidence
CursorBench 4.027.2%estimated ± 3.6 pp, medium confidence
DeepSWE56.3%estimated ± 7.2 pp, low confidence
HLE w/o tools25.1%estimated ± 5.1 pp, low confidence
PostTrainBench v1.116.5%estimated ± 2.7 pp, high confidence
SWE Multilingual63.1%measured
Vibe Code Bench2.7%estimated ± 8.1 pp, medium confidence
MMLU-ProX80.1%estimated ± 1.7 pp, low confidence
NL2Repo37.5%estimated ± 4.7 pp, medium confidence
NOVA-6351.6%estimated ± 1.6 pp, low confidence
Terminal-Bench 2.045.8%measured
LiveCodeBench Pro74.6%estimated ± 6.0 pp, low confidence
React Native Evals49.3%estimated ± 2.4 pp, low confidence
SWE-Rebench57.1%estimated ± 3.6 pp, medium confidence
ARC-AGI-30.0%estimated ± 9.2 pp, low confidence
FrontierCode 1.1 Extended5.4%estimated ± 2.4 pp, low confidence
HealthBench Professional49.8%estimated ± 1.3 pp, low confidence
HLE-Verified0.0%estimated ± 4.6 pp, low confidence
LABBench270.4%estimated ± 1.8 pp, low confidence
LiveCodeBench83.1%estimated ± 10.2 pp, low confidence
OpenHarmony Bench47.3%estimated ± 3.2 pp, medium confidence