benchgap
Xiaomi · model

MiMo-V2.5 benchmark scores

As of 2026-10-07, MiMo-V2.5 (Xiaomi) has measured scores on 13 benchmarks and estimated scores on 67 more.

BenchmarkScoreSource
AA-GPQA Diamond82.5%estimated ± 4.2 pp, high confidence
AA-HLE24.0%estimated ± 7.4 pp, medium confidence
AA-SciCode48.3%estimated ± 3.8 pp, high confidence
Artificial Analysis Intelligence Index20.9%estimated ± 7.9 pp, medium confidence
CritPt3.6%estimated ± 6.6 pp, medium confidence
GDPval-AA32.6%estimated ± 8.3 pp, low confidence
GPQA Diamond85.8%estimated ± 1.7 pp, medium confidence
HLE37.3%estimated ± 9.8 pp, low confidence
LiveCodeBench v687.9%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy24.9%estimated ± 7.7 pp, medium confidence
SciCode39.7%estimated ± 2.9 pp, low confidence
Terminal-Bench 2.178.3%estimated ± 6.9 pp, medium confidence
BrowseComp63.6%estimated ± 5.7 pp, medium confidence
VITA-Bench39.6%estimated ± 10.7 pp, low confidence
AA Coding Index54.3%estimated ± 5.8 pp, medium confidence
AA-MMMU-Pro75.9%estimated ± 4.9 pp, medium confidence
APEX-Agents-AA31.8%estimated ± 10.0 pp, low confidence
DeepSearchQA80.5%estimated ± 10.2 pp, low confidence
SWE-bench Verified79.4%estimated ± 2.5 pp, high confidence
BFCL v462.3%estimated ± 8.6 pp, low confidence
CyberGym41.1%estimated ± 10.0 pp, low confidence
JobBench31.8%estimated ± 8.4 pp, low confidence
SWE-bench Pro56.1%measured
τ³-bench results68.9%estimated ± 8.9 pp, medium confidence
WideResearch70.4%estimated ± 5.9 pp, low confidence
AA Agentic Index19.7%estimated ± 5.1 pp, medium confidence
GPQA86.9%estimated ± 1.6 pp, high confidence
Gert Labs46.9%measured
ApprenticeBench0.0%estimated ± 3.1 pp, low confidence
ARC-AGI-172.6%estimated ± 5.9 pp, medium confidence
ARC-AGI-224.6%estimated ± 12.6 pp, low confidence
CursorBench 3.152.0%estimated ± 5.6 pp, low confidence
CursorBench 3.252.3%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.06.1%estimated ± 6.9 pp, low confidence
FrontierCode 1.1 Main25.1%estimated ± 5.2 pp, low confidence
FrontierSWE v28.2%estimated ± 9.0 pp, medium confidence
OSWorld-Verified58.0%estimated ± 5.9 pp, medium confidence
Vals GPQA Diamond81.6%measured
Vals LiveCodeBench81.5%measured
Vals MMLU-Pro82.9%measured
Vals SWE-bench71.0%measured
Terminal-Bench 2.1 (Vals AI)60.7%measured
VulcanBench v381.6%estimated ± 5.6 pp, low confidence
AA Terminal-Bench 2.177.5%estimated ± 4.4 pp, high confidence
AA Terminal-Bench 4.014.5%estimated ± 14.0 pp, low confidence
CursorBench 4.032.5%estimated ± 3.6 pp, medium confidence
DeepSWE56.3%estimated ± 7.2 pp, low confidence
HLE w/o tools26.5%estimated ± 5.1 pp, medium confidence
OSWorld 2.00.0%estimated ± 4.5 pp, low confidence
PostTrainBench v1.120.5%estimated ± 2.7 pp, high confidence
CharXiv81.0%measured
CharXiv w/o tools82.1%estimated ± 1.8 pp, medium confidence
Claw-Eval62.3%measured
MathVision86.8%estimated ± 2.6 pp, high confidence
MCP Atlas57.0%estimated ± 11.2 pp, low confidence
MCP-Tasks77.5%estimated ± 6.4 pp, low confidence
MMMU-Pro77.9%measured
NL2Repo41.1%estimated ± 4.7 pp, high confidence
ScreenSpot Pro73.6%estimated ± 7.1 pp, medium confidence
Terminal-Bench 2.065.8%measured
Toolathlon52.3%estimated ± 6.9 pp, medium confidence
VideoMMMU84.8%estimated ± 1.0 pp, high confidence
ERQA66.5%estimated ± 3.7 pp, high confidence
LiveCodeBench Pro80.6%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)70.0%estimated ± 2.4 pp, high confidence
React Native Evals67.1%estimated ± 2.4 pp, low confidence
ResearchClawBench16.9%measured
ARC-AGI-30.0%estimated ± 9.2 pp, low confidence
FrontierCode 1.1 Extended17.0%estimated ± 2.4 pp, low confidence
HealthBench Professional50.1%estimated ± 1.3 pp, low confidence
HLE-Verified0.0%estimated ± 4.6 pp, low confidence
LABBench279.1%estimated ± 1.8 pp, low confidence
MMMU84.4%estimated ± 1.1 pp, medium confidence
OpenHarmony Bench51.5%estimated ± 3.2 pp, high confidence
SimpleVQA67.1%estimated ± 7.8 pp, low confidence
MMMU-Pro w/ Python79.3%estimated ± 0.6 pp, high confidence
MathVision w/ Python76.8%estimated ± 0.9 pp, low confidence
RealWorldQA85.4%estimated ± 1.1 pp, medium confidence
Video-MME (with subtitle)87.7%measured
OmniDocBench 1.591.7%estimated ± 5.9 pp, low confidence