benchgap
Meta · model

Muse Spark 1.1 benchmark scores

As of 2026-10-07, Muse Spark 1.1 (Meta) has measured scores on 31 benchmarks and estimated scores on 83 more.

BenchmarkScoreSource
AA-GPQA Diamond89.8%measured
AA-HLE46.2%measured
AA-SciCode58.8%measured
Artificial Analysis Intelligence Index33.7%measured
CritPt15.1%measured
GDPval-AA35.7%measured
GPQA Diamond93.6%estimated ± 1.0 pp, medium confidence
HLE62.1%measured
AA-LCR77.7%measured
LiveCodeBench v690.2%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy52.1%measured
SciCode48.5%estimated ± 2.9 pp, medium confidence
Terminal-Bench 2.180.0%measured
BrowseComp80.9%estimated ± 1.8 pp, low confidence
HLE w/ tools18.0%estimated ± 5.0 pp, low confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
VITA-Bench51.3%estimated ± 9.9 pp, low confidence
τ²-bench results92.6%estimated ± 13.3 pp, low confidence
AA Coding Index71.3%measured
APEX-Agents40.4%estimated ± 0.6 pp, low confidence
APEX-Agents-AA35.2%estimated ± 2.9 pp, medium confidence
BioMysteryBench (human-difficult)46.5%estimated ± 5.9 pp, low confidence
DeepSearchQA84.9%measured
SWE-bench Verified83.5%estimated ± 2.5 pp, high confidence
AutomationBench28.4%estimated ± 5.8 pp, low confidence
BFCL v476.0%estimated ± 2.6 pp, medium confidence
CyberGym59.0%measured
JobBench54.7%measured
SWE-bench Pro61.5%measured
τ³-bench results69.6%estimated ± 1.8 pp, low confidence
WideResearch82.3%estimated ± 1.2 pp, medium confidence
AA Agentic Index27.5%measured
MMLU-Pro85.9%estimated ± 2.5 pp, low confidence
GPQA93.6%estimated ± 1.0 pp, medium confidence
Gert Labs69.7%estimated ± 3.7 pp, medium confidence
AA-AnalystAgent51.7%estimated ± 5.5 pp, medium confidence
AA EnterpriseOps-Gym45.7%estimated ± 5.9 pp, low confidence
AA Harvey LAB85.8%estimated ± 1.2 pp, medium confidence
ApprenticeBench3.4%estimated ± 4.0 pp, low confidence
ARC-AGI-196.8%estimated ± 3.7 pp, high confidence
ARC-AGI-266.2%estimated ± 12.6 pp, low confidence
CursorBench 3.165.7%estimated ± 3.4 pp, medium confidence
CursorBench 3.260.0%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.012.7%estimated ± 6.9 pp, medium confidence
FrontierCode 1.1 Main33.0%estimated ± 3.7 pp, low confidence
FrontierSWE v217.8%estimated ± 9.0 pp, medium confidence
OSWorld-Verified80.8%measured
Vals GPQA Diamond91.2%measured
Vals LiveCodeBench85.9%measured
Vals MMLU-Pro88.7%measured
Vals SWE-bench82.0%measured
Terminal-Bench 2.1 (Vals AI)69.3%measured
VulcanBench v384.2%estimated ± 5.6 pp, low confidence
AA AutomationBench54.9%estimated ± 5.0 pp, medium confidence
GDP.pdf18.1%estimated ± 4.9 pp, high confidence
AA ITBench56.3%estimated ± 3.6 pp, low confidence
AA Tau3 Banking22.0%estimated ± 5.8 pp, medium confidence
AA Terminal-Bench 2.184.6%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.018.5%estimated ± 8.9 pp, medium confidence
CursorBench 4.034.2%estimated ± 1.6 pp, medium confidence
CWE-bench v145.0%estimated ± 6.9 pp, medium confidence
DeepSWE53.3%measured
HLE w/o tools52.2%measured
OSWorld 2.014.2%measured
PostTrainBench v1.126.0%estimated ± 2.7 pp, high confidence
Toolathlon-Verified69.8%estimated ± 1.0 pp, low confidence
Vibe Code Bench56.1%estimated ± 10.3 pp, low confidence
CharXiv88.4%measured
CharXiv w/o tools80.5%estimated ± 2.3 pp, high confidence
ExploitGym0.8%measured
C-Eval93.3%estimated ± 0.8 pp, low confidence
Claw-Eval66.2%estimated ± 5.6 pp, low confidence
DeepPlanning53.8%estimated ± 13.1 pp, low confidence
MathVision93.3%estimated ± 0.5 pp, medium confidence
MCP Atlas88.1%measured
MCP-Tasks77.1%estimated ± 0.6 pp, low confidence
MMLU-Redux97.5%estimated ± 1.1 pp, low confidence
MMMU-Pro81.1%estimated ± 3.2 pp, high confidence
NL2Repo47.2%estimated ± 4.7 pp, high confidence
QwenClawBench73.0%estimated ± 2.6 pp, medium confidence
ScreenSpot Pro84.3%estimated ± 7.1 pp, medium confidence
SuperGPQA86.5%estimated ± 7.5 pp, low confidence
Terminal-Bench 2.073.1%estimated ± 5.9 pp, medium confidence
Toolathlon75.6%measured
VideoMMMU86.8%estimated ± 1.5 pp, high confidence
ERQA70.4%estimated ± 3.7 pp, high confidence
HealthBench Hard30.5%estimated ± 7.7 pp, low confidence
LiveCodeBench Pro84.9%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)78.4%estimated ± 6.4 pp, low confidence
MedXpertQA (Text)68.9%estimated ± 8.9 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals84.7%estimated ± 2.4 pp, medium confidence
ResearchClawBench19.9%estimated ± 1.8 pp, low confidence
SWE-Rebench58.1%estimated ± 8.9 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, high confidence
BioMysteryBench (human-solvable)87.6%estimated ± 0.6 pp, medium confidence
DRACO76.2%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended33.8%estimated ± 2.4 pp, low confidence
HealthBench (raw)54.6%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)52.6%estimated ± 2.2 pp, low confidence
HealthBench Professional59.3%measured
HealthBench Professional (raw)64.6%estimated ± 5.0 pp, medium confidence
HLE-Verified17.2%estimated ± 2.0 pp, low confidence
LABBench282.4%estimated ± 1.8 pp, low confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
Agents' Last Exam24.0%estimated ± 2.1 pp, medium confidence
OpenHarmony Bench58.8%estimated ± 2.4 pp, high confidence
BabyVision76.3%measured
SimpleVQA71.0%estimated ± 7.8 pp, low confidence
MMLU95.2%estimated ± 1.1 pp, low confidence
MathVision w/ Python83.5%estimated ± 0.9 pp, low confidence
RealWorldQA87.0%estimated ± 1.1 pp, medium confidence
Video-MME (with subtitle)89.1%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.589.9%estimated ± 7.4 pp, low confidence