benchgap
Meta · model

Muse Glimmer 30B benchmark scores

As of 2026-10-07, Muse Glimmer 30B (Meta) has measured scores on 31 benchmarks and estimated scores on 95 more.

BenchmarkScoreSource
AA-GPQA Diamond83.5%measured
AA-HLE22.0%measured
AA-SciCode44.9%measured
AIME2694.7%measured
Artificial Analysis Intelligence Index17.5%measured
CritPt2.6%measured
GDPval-AA14.5%measured
GPQA Diamond82.5%estimated ± 2.2 pp, high confidence
HLE32.5%estimated ± 9.5 pp, medium confidence
IFBench77.0%measured
AA-LCR83.3%measured
LiveCodeBench v685.3%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy27.0%measured
SciCode43.6%measured
Terminal-Bench 2.151.7%measured
BrowseComp78.3%estimated ± 3.1 pp, high confidence
HLE w/ tools39.0%estimated ± 6.1 pp, low confidence
IFEval92.8%estimated ± 3.5 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
VITA-Bench46.4%estimated ± 9.9 pp, low confidence
τ²-bench results89.0%estimated ± 2.4 pp, medium confidence
AA Coding Index49.0%measured
AA-MMMU-Pro74.3%measured
APEX-Agents30.3%estimated ± 0.6 pp, medium confidence
APEX-Agents-AA34.3%estimated ± 5.7 pp, medium confidence
BioMysteryBench (human-difficult)32.8%estimated ± 5.9 pp, low confidence
DeepSearchQA74.6%measured
SWE-bench Verified76.0%measured
AutomationBench48.1%estimated ± 7.0 pp, low confidence
BFCL v470.6%estimated ± 2.6 pp, high confidence
CyberGym53.6%estimated ± 5.3 pp, low confidence
JobBench45.9%estimated ± 6.8 pp, medium confidence
SkillsBench44.3%measured
SWE-bench Pro51.2%measured
τ³-bench results69.6%estimated ± 1.8 pp, low confidence
WideResearch73.0%estimated ± 1.5 pp, low confidence
AA Agentic Index10.5%measured
MMLU-Pro85.6%estimated ± 3.5 pp, high confidence
GPQA82.5%estimated ± 3.2 pp, high confidence
AA-IFBench78.8%estimated ± 8.5 pp, medium confidence
Gert Labs52.9%estimated ± 7.2 pp, medium confidence
AA-AnalystAgent16.4%estimated ± 3.6 pp, medium confidence
AA EnterpriseOps-Gym34.7%measured
AA Harvey LAB79.4%estimated ± 1.2 pp, low confidence
ApprenticeBench0.0%estimated ± 5.2 pp, low confidence
ARC-AGI-132.9%estimated ± 5.4 pp, low confidence
ARC-AGI-210.0%estimated ± 14.2 pp, low confidence
CursorBench 3.131.4%estimated ± 3.4 pp, low confidence
CursorBench 3.242.8%estimated ± 3.6 pp, medium confidence
Terminal-Bench 3.03.8%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main13.1%estimated ± 2.7 pp, low confidence
FrontierSWE v20.0%estimated ± 9.0 pp, low confidence
OfficeQA Pro50.3%estimated ± 7.2 pp, low confidence
OSWorld-Verified65.9%measured
Vals GPQA Diamond77.7%estimated ± 7.4 pp, medium confidence
Vals LiveCodeBench84.7%estimated ± 0.8 pp, medium confidence
Vals MMLU-Pro84.4%estimated ± 2.2 pp, high confidence
Vals SWE-bench69.3%estimated ± 5.3 pp, low confidence
Terminal-Bench 2.1 (Vals AI)51.1%estimated ± 6.0 pp, low confidence
AA AutomationBench6.8%measured
GDP.pdf10.0%measured
AA ITBench43.3%estimated ± 3.6 pp, medium confidence
AA Tau3 Banking23.5%measured
AA Terminal-Bench 2.151.7%measured
AA Terminal-Bench 4.00.5%measured
CursorBench 4.015.5%estimated ± 1.6 pp, medium confidence
CWE-bench v137.0%estimated ± 6.0 pp, medium confidence
DeepSWE12.7%estimated ± 6.4 pp, low confidence
HLE w/o tools19.0%estimated ± 5.7 pp, medium confidence
OSWorld 2.00.0%estimated ± 3.6 pp, medium confidence
PostTrainBench v1.120.6%estimated ± 4.7 pp, medium confidence
Terminal-Bench 4.04.5%estimated ± 5.1 pp, low confidence
Terminal-Bench-Science 0.10.0%estimated ± 4.1 pp, low confidence
Toolathlon-Verified73.0%estimated ± 2.2 pp, low confidence
Vibe Code Bench6.7%estimated ± 8.1 pp, medium confidence
CharXiv78.8%measured
CharXiv w/o tools77.9%estimated ± 1.8 pp, medium confidence
ExploitGym13.6%estimated ± 5.0 pp, low confidence
C-Eval90.0%estimated ± 1.0 pp, low confidence
Claw-Eval69.7%estimated ± 5.7 pp, medium confidence
DeepPlanning47.9%estimated ± 13.1 pp, low confidence
HMMT Feb 202594.7%estimated ± 1.9 pp, medium confidence
HMMT Feb 202685.6%estimated ± 4.0 pp, high confidence
MathVision90.4%estimated ± 2.4 pp, medium confidence
MCP Atlas75.5%measured
MCP-Tasks77.0%estimated ± 0.6 pp, low confidence
MMAnswerBench82.3%estimated ± 1.3 pp, high confidence
MMLU-Redux95.0%estimated ± 1.1 pp, medium confidence
MMMU-Pro74.0%measured
NL2Repo35.3%estimated ± 5.1 pp, medium confidence
QwenClawBench62.7%estimated ± 2.6 pp, high confidence
ScreenSpot Pro75.4%measured
SuperGPQA72.0%estimated ± 7.5 pp, medium confidence
Toolathlon27.0%estimated ± 6.3 pp, low confidence
VideoMMMU84.1%estimated ± 1.0 pp, high confidence
ERQA65.3%estimated ± 3.7 pp, high confidence
HealthBench Hard3.8%estimated ± 7.7 pp, low confidence
LiveCodeBench Pro76.4%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)60.8%estimated ± 2.4 pp, high confidence
MMLU-Pro (Arcee)86.9%estimated ± 4.0 pp, medium confidence
React Native Evals76.4%estimated ± 3.5 pp, high confidence
ResearchClawBench18.7%estimated ± 1.8 pp, medium confidence
SWE-Rebench57.8%estimated ± 3.6 pp, medium confidence
ARC-AGI-31.7%estimated ± 3.4 pp, medium confidence
BioMysteryBench (human-solvable)81.7%estimated ± 0.6 pp, low confidence
DRACO67.5%estimated ± 3.0 pp, low confidence
FrontierCode 1.1 Extended52.8%estimated ± 3.0 pp, low confidence
HealthBench (raw)30.3%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional57.7%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)42.2%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench267.5%estimated ± 1.8 pp, low confidence
AIME 202595.4%estimated ± 1.5 pp, medium confidence
MMMU82.4%estimated ± 0.9 pp, high confidence
LiveCodeBench96.2%estimated ± 8.9 pp, low confidence
Agents' Last Exam23.0%estimated ± 6.3 pp, low confidence
OpenHarmony Bench51.0%estimated ± 2.4 pp, high confidence
SimpleVQA82.7%estimated ± 5.7 pp, low confidence
MMLU89.8%estimated ± 1.1 pp, medium confidence
MMMU-Pro w/ Python75.6%estimated ± 0.6 pp, high confidence
MATH-50097.6%estimated ± 3.1 pp, medium confidence
MathVision w/ Python76.8%estimated ± 0.9 pp, low confidence
RealWorldQA84.9%estimated ± 1.1 pp, medium confidence
Video-MME (with subtitle)87.2%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.575.8%measured