benchgap
Moonshot · model

Kimi K2 benchmark scores

As of 2026-10-07, Kimi K2 (Moonshot) has measured scores on 10 benchmarks and estimated scores on 38 more.

BenchmarkScoreSource
AA-GPQA Diamond76.6%measured
AA-HLE7.4%measured
AIME2695.5%estimated ± 0.6 pp, low confidence
Artificial Analysis Intelligence Index12.7%measured
CritPt0.0%measured
GDPval-AA3.2%estimated ± 12.5 pp, low confidence
GPQA Diamond71.9%estimated ± 2.2 pp, high confidence
HLE19.0%estimated ± 9.5 pp, medium confidence
IFBench56.0%estimated ± 6.9 pp, low confidence
AA-LCR53.0%measured
AA-Omniscience Accuracy27.4%measured
IFEval89.1%estimated ± 1.3 pp, high confidence
LongBench v20.0%estimated ± 5.9 pp, low confidence
τ²-bench results61.1%measured
APEX-Agents-AA1.6%estimated ± 11.4 pp, low confidence
BioMysteryBench (human-difficult)33.1%estimated ± 5.9 pp, low confidence
BFCL v447.6%estimated ± 9.8 pp, low confidence
AA Agentic Index12.0%estimated ± 12.3 pp, low confidence
MMLU-Pro85.6%estimated ± 3.5 pp, high confidence
FrontierMath v2 (Tier 4)0.0%measured
FrontierMath v2 (Tiers 1-3)21.4%measured
GPQA73.0%estimated ± 3.2 pp, high confidence
AA-IFBench41.5%measured
AA EnterpriseOps-Gym6.1%estimated ± 3.6 pp, low confidence
ARC-AGI-18.4%estimated ± 5.4 pp, low confidence
ARC-AGI-22.7%estimated ± 14.2 pp, low confidence
Vals GPQA Diamond67.3%estimated ± 7.4 pp, medium confidence
Vals MMLU-Pro84.5%estimated ± 2.2 pp, high confidence
AA ITBench83.2%estimated ± 3.6 pp, low confidence
HLE w/o tools5.9%estimated ± 5.7 pp, low confidence
C-Eval86.6%estimated ± 1.0 pp, low confidence
HMMT Nov 202599.9%estimated ± 2.6 pp, low confidence
MCP-Tasks87.9%estimated ± 8.2 pp, low confidence
MMAnswerBench81.9%estimated ± 1.3 pp, low confidence
MMLU-Redux95.0%estimated ± 1.1 pp, medium confidence
SuperGPQA72.3%estimated ± 7.5 pp, medium confidence
HealthBench Hard0.6%estimated ± 7.7 pp, low confidence
MMLU-Pro (Arcee)77.1%estimated ± 4.0 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, medium confidence
FrontierMath (legacy)22.3%estimated ± 0.7 pp, low confidence
BioMysteryBench (human-solvable)81.7%estimated ± 0.6 pp, low confidence
HealthBench (raw)11.4%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional57.7%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)25.1%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench258.5%estimated ± 1.8 pp, low confidence
MMLU89.9%estimated ± 1.1 pp, medium confidence