benchgap
Meta · model

Muse Spark benchmark scores

As of 2026-10-07, Muse Spark (Meta) has measured scores on 37 benchmarks and estimated scores on 82 more.

BenchmarkScoreSource
AA-GPQA Diamond88.4%measured
AA-HLE40.7%measured
AA-SciCode48.7%estimated ± 3.2 pp, high confidence
AIME2696.4%estimated ± 0.6 pp, low confidence
Artificial Analysis Intelligence Index31.3%measured
CritPt11.3%measured
GDPval-AA25.1%measured
GPQA Diamond89.5%measured
HLE50.4%measured
IFBench74.5%estimated ± 6.9 pp, medium confidence
AA-LCR78.0%measured
LiveCodeBench v685.8%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy49.6%measured
SciCode47.3%estimated ± 3.8 pp, high confidence
BrowseComp56.7%estimated ± 5.2 pp, low confidence
HLE w/ tools3.5%estimated ± 5.0 pp, low confidence
IFEval93.7%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
VITA-Bench41.0%estimated ± 10.7 pp, low confidence
τ²-bench results91.5%measured
AA Coding Index58.6%measured
AA-MMMU-Pro80.5%measured
APEX-Agents23.0%estimated ± 10.0 pp, low confidence
APEX-Agents-AA14.0%estimated ± 6.0 pp, low confidence
BioMysteryBench (human-difficult)45.4%estimated ± 5.9 pp, low confidence
DeepSearchQA74.8%measured
SWE-bench Verified77.4%measured
AutomationBench28.4%estimated ± 5.8 pp, low confidence
BFCL v474.1%estimated ± 2.6 pp, high confidence
CyberGym43.5%measured
JobBench33.6%estimated ± 8.4 pp, low confidence
SWE-bench Pro52.4%measured
WideResearch75.7%estimated ± 9.1 pp, low confidence
AA Agentic Index15.8%estimated ± 1.7 pp, high confidence
MMLU-Pro85.7%estimated ± 2.5 pp, low confidence
FrontierMath v2 (Tier 4)14.6%measured
FrontierMath v2 (Tiers 1-3)39.0%measured
GPQA89.5%estimated ± 0.1 pp, high confidence
AA-IFBench75.9%measured
Gert Labs54.8%estimated ± 6.0 pp, low confidence
AA-AnalystAgent17.5%estimated ± 12.0 pp, low confidence
AA EnterpriseOps-Gym40.4%estimated ± 3.6 pp, medium confidence
ApprenticeBench3.4%estimated ± 4.0 pp, low confidence
ARC-AGI-182.1%estimated ± 1.9 pp, high confidence
ARC-AGI-242.5%measured
CursorBench 3.147.9%estimated ± 4.8 pp, low confidence
CursorBench 3.245.4%estimated ± 3.6 pp, medium confidence
FrontierCode 1.1 Main16.8%estimated ± 2.7 pp, low confidence
FrontierSWE v21.6%estimated ± 9.0 pp, low confidence
OfficeQA Pro60.4%estimated ± 7.2 pp, medium confidence
OSWorld-Verified65.8%estimated ± 4.7 pp, medium confidence
Vals GPQA Diamond89.6%measured
Vals LiveCodeBench79.8%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro87.3%measured
Vals SWE-bench74.4%measured
Terminal-Bench 2.1 (Vals AI)51.7%estimated ± 6.3 pp, medium confidence
AA AutomationBench4.2%estimated ± 7.0 pp, medium confidence
GDP.pdf9.4%estimated ± 5.4 pp, medium confidence
AA ITBench47.8%estimated ± 3.6 pp, medium confidence
AA Tau3 Banking22.0%estimated ± 5.8 pp, medium confidence
CursorBench 4.021.7%estimated ± 1.6 pp, medium confidence
CWE-bench v135.0%estimated ± 6.9 pp, low confidence
DeepSWE30.0%estimated ± 6.4 pp, low confidence
HLE w/o tools42.8%measured
PostTrainBench v1.122.6%estimated ± 3.3 pp, medium confidence
Toolathlon-Verified61.4%estimated ± 2.2 pp, low confidence
Vibe Code Bench19.7%measured
CharXiv86.4%measured
CharXiv w/o tools84.9%estimated ± 1.8 pp, medium confidence
ExploitGym1.5%estimated ± 9.3 pp, low confidence
C-Eval93.3%estimated ± 0.8 pp, low confidence
Claw-Eval63.8%measured
HMMT Nov 202593.2%estimated ± 2.6 pp, low confidence
MathVision88.9%estimated ± 1.4 pp, low confidence
MCP Atlas65.4%estimated ± 11.1 pp, low confidence
MCP-Tasks80.4%estimated ± 6.4 pp, low confidence
MMAnswerBench85.2%estimated ± 1.3 pp, low confidence
MMLU-Redux97.3%estimated ± 1.1 pp, low confidence
MMMU-Pro80.4%measured
NL2Repo42.7%estimated ± 4.7 pp, high confidence
QwenClawBench55.9%estimated ± 6.0 pp, low confidence
ScreenSpot Pro84.1%measured
SuperGPQA85.5%estimated ± 7.5 pp, low confidence
Terminal-Bench 2.059.0%measured
Toolathlon43.3%estimated ± 6.3 pp, medium confidence
VideoMMMU87.0%estimated ± 1.0 pp, high confidence
ERQA64.7%measured
HealthBench Hard42.8%measured
LiveCodeBench Pro80.0%measured
MedXpertQA (MM)78.4%measured
MedXpertQA (Text)52.6%measured
MMLU-Pro (Arcee)86.7%estimated ± 3.4 pp, low confidence
React Native Evals72.2%estimated ± 2.4 pp, medium confidence
ResearchClawBench17.4%estimated ± 1.3 pp, high confidence
SWE-Rebench60.6%estimated ± 3.6 pp, medium confidence
ARC-AGI-31.7%estimated ± 3.4 pp, high confidence
FrontierMath (legacy)39.6%estimated ± 0.7 pp, low confidence
BioMysteryBench (human-solvable)84.4%estimated ± 0.6 pp, low confidence
DRACO71.8%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended21.6%estimated ± 2.4 pp, low confidence
HealthBench (raw)49.7%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional58.8%estimated ± 1.3 pp, medium confidence
HealthBench Professional (raw)60.0%estimated ± 5.0 pp, medium confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench281.6%estimated ± 1.8 pp, low confidence
MMMU87.7%estimated ± 0.9 pp, medium confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
Agents' Last Exam24.0%estimated ± 2.1 pp, medium confidence
OpenHarmony Bench51.4%estimated ± 2.8 pp, high confidence
SimpleVQA71.3%measured
ZeroBench33.0%measured
MMMLU89.6%estimated ± 1.1 pp, medium confidence
MMLU94.9%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python81.7%estimated ± 0.6 pp, high confidence
MathVision w/ Python77.4%estimated ± 0.9 pp, low confidence
RealWorldQA85.4%estimated ± 1.1 pp, medium confidence
Video-MME (with subtitle)88.7%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.591.7%estimated ± 5.9 pp, low confidence