benchgap
Xiaomi · model

MiMo-V2-Pro benchmark scores

As of 2026-10-07, MiMo-V2-Pro (Xiaomi) has measured scores on 12 benchmarks and estimated scores on 58 more.

BenchmarkScoreSource
AA-GPQA Diamond87.0%measured
AA-HLE30.4%measured
AA-SciCode47.3%estimated ± 3.9 pp, high confidence
Artificial Analysis Intelligence Index28.6%measured
CritPt0.3%measured
GDPval-AA29.9%estimated ± 8.3 pp, low confidence
GPQA Diamond87.1%estimated ± 2.2 pp, high confidence
HLE23.5%estimated ± 9.5 pp, medium confidence
IFBench67.6%estimated ± 6.9 pp, medium confidence
AA-LCR68.3%measured
LiveCodeBench v686.0%estimated ± 8.1 pp, low confidence
AA-Omniscience Accuracy26.6%measured
SciCode47.6%estimated ± 3.8 pp, high confidence
BrowseComp58.7%estimated ± 5.7 pp, medium confidence
IFEval93.3%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
VITA-Bench35.6%estimated ± 10.7 pp, low confidence
τ²-bench results95.0%measured
AA Coding Index56.4%estimated ± 5.1 pp, medium confidence
APEX-Agents-AA29.2%estimated ± 10.0 pp, low confidence
BioMysteryBench (human-difficult)32.5%estimated ± 5.9 pp, low confidence
DeepSearchQA77.0%estimated ± 10.2 pp, low confidence
SWE-bench Verified78.0%measured
BFCL v449.6%estimated ± 8.6 pp, low confidence
CyberGym28.0%estimated ± 10.0 pp, low confidence
JobBench25.7%estimated ± 8.4 pp, low confidence
SWE-bench Pro54.4%estimated ± 3.2 pp, high confidence
τ³-bench results99.9%estimated ± 8.9 pp, low confidence
WideResearch48.9%estimated ± 5.9 pp, low confidence
AA Agentic Index17.7%estimated ± 5.1 pp, medium confidence
MMLU-Pro85.6%estimated ± 3.5 pp, high confidence
GPQA89.9%estimated ± 3.2 pp, high confidence
AA-IFBench68.8%measured
Gert Labs36.7%measured
AA EnterpriseOps-Gym46.6%estimated ± 3.6 pp, medium confidence
ApprenticeBench0.0%estimated ± 3.1 pp, low confidence
ARC-AGI-183.9%estimated ± 5.4 pp, medium confidence
ARC-AGI-247.1%estimated ± 14.2 pp, low confidence
FrontierCode 1.1 Main18.6%estimated ± 2.7 pp, low confidence
OSWorld-Verified53.3%estimated ± 5.9 pp, medium confidence
Vals GPQA Diamond83.5%estimated ± 7.4 pp, medium confidence
Vals LiveCodeBench80.1%estimated ± 9.9 pp, low confidence
Vals MMLU-Pro84.3%estimated ± 2.2 pp, high confidence
Vals SWE-bench73.6%estimated ± 6.8 pp, medium confidence
AA ITBench45.6%estimated ± 3.6 pp, medium confidence
HLE w/o tools27.5%estimated ± 5.7 pp, medium confidence
OSWorld 2.00.0%estimated ± 4.5 pp, low confidence
Vibe Code Bench23.1%estimated ± 8.1 pp, medium confidence
C-Eval93.8%estimated ± 1.0 pp, low confidence
Claw-Eval57.8%measured
MCP Atlas47.3%estimated ± 11.2 pp, low confidence
MCP-Tasks69.3%estimated ± 6.4 pp, low confidence
MMLU-Redux94.9%estimated ± 1.1 pp, medium confidence
NL2Repo43.2%estimated ± 6.1 pp, medium confidence
SuperGPQA71.6%estimated ± 7.5 pp, medium confidence
Toolathlon49.7%estimated ± 6.9 pp, medium confidence
HealthBench Hard24.8%estimated ± 7.7 pp, low confidence
MMLU-Pro (Arcee)78.3%estimated ± 4.0 pp, low confidence
ResearchClawBench15.3%measured
SWE-Rebench61.3%estimated ± 3.6 pp, medium confidence
ARC-AGI-31.7%estimated ± 3.4 pp, medium confidence
BioMysteryBench (human-solvable)81.7%estimated ± 0.6 pp, low confidence
HealthBench (raw)39.6%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional57.7%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)50.7%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench272.5%estimated ± 1.8 pp, low confidence
LiveCodeBench85.2%estimated ± 10.2 pp, low confidence
MMLU89.6%estimated ± 1.1 pp, medium confidence