benchgap
Moonshot · model

Kimi K2.7 Code benchmark scores

As of 2026-10-07, Kimi K2.7 Code (Moonshot) has measured scores on 19 benchmarks and estimated scores on 80 more.

BenchmarkScoreSource
AA-GPQA Diamond89.6%measured
AA-HLE35.0%measured
AA-SciCode47.8%measured
Artificial Analysis Intelligence Index25.8%measured
CritPt10.0%measured
GDPval-AA27.0%measured
GPQA Diamond89.2%estimated ± 2.2 pp, high confidence
HLE43.8%estimated ± 9.5 pp, medium confidence
IFBench63.2%estimated ± 6.9 pp, medium confidence
AA-LCR79.3%measured
LiveCodeBench v689.2%estimated ± 3.1 pp, high confidence
AA-Omniscience Accuracy39.6%measured
SciCode40.9%estimated ± 2.9 pp, low confidence
Terminal-Bench 2.182.6%estimated ± 6.9 pp, medium confidence
BrowseComp79.8%estimated ± 8.4 pp, medium confidence
HLE w/ tools45.4%estimated ± 6.1 pp, medium confidence
IFEval93.0%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
VITA-Bench46.6%estimated ± 9.9 pp, low confidence
τ²-bench results90.1%measured
AA Coding Index60.8%measured
APEX-Agents30.7%estimated ± 0.6 pp, medium confidence
APEX-Agents-AA34.6%estimated ± 5.7 pp, medium confidence
BioMysteryBench (human-difficult)40.7%estimated ± 5.9 pp, low confidence
DeepSearchQA85.8%estimated ± 4.5 pp, high confidence
SWE-bench Verified76.3%estimated ± 4.4 pp, high confidence
AutomationBench28.4%estimated ± 9.7 pp, low confidence
BFCL v474.5%estimated ± 2.6 pp, high confidence
CyberGym62.8%estimated ± 5.3 pp, low confidence
JobBench46.4%estimated ± 6.8 pp, medium confidence
SWE-bench Pro54.8%estimated ± 4.5 pp, high confidence
τ³-bench results69.6%estimated ± 1.8 pp, low confidence
WideResearch77.8%estimated ± 9.1 pp, low confidence
AA Agentic Index22.5%measured
MMLU-Pro85.7%estimated ± 3.5 pp, high confidence
GPQA88.8%estimated ± 3.2 pp, high confidence
AA-IFBench63.1%measured
Gert Labs63.9%estimated ± 8.2 pp, low confidence
AA-AnalystAgent19.1%estimated ± 12.0 pp, low confidence
AA EnterpriseOps-Gym38.0%estimated ± 3.6 pp, medium confidence
AA Harvey LAB82.2%estimated ± 1.2 pp, medium confidence
ARC-AGI-177.1%estimated ± 5.4 pp, medium confidence
ARC-AGI-237.0%estimated ± 14.2 pp, low confidence
CursorBench 3.138.6%estimated ± 3.4 pp, low confidence
CursorBench 3.249.7%measured
Terminal-Bench 3.010.6%estimated ± 6.9 pp, low confidence
FrontierCode 1.1 Main33.0%estimated ± 1.6 pp, low confidence
FrontierSWE v29.4%estimated ± 9.5 pp, medium confidence
OSWorld-Verified69.7%estimated ± 8.7 pp, low confidence
Vals GPQA Diamond88.0%estimated ± 7.4 pp, medium confidence
Vals LiveCodeBench82.1%measured
Vals MMLU-Pro86.8%estimated ± 2.2 pp, high confidence
Vals SWE-bench78.2%measured
Terminal-Bench 2.1 (Vals AI)67.0%measured
VulcanBench v372.0%estimated ± 5.1 pp, low confidence
AA AutomationBench4.5%estimated ± 7.0 pp, medium confidence
GDP.pdf9.4%estimated ± 5.4 pp, medium confidence
AA ITBench43.8%estimated ± 3.6 pp, medium confidence
AA Tau3 Banking22.0%estimated ± 5.8 pp, medium confidence
AA Terminal-Bench 2.182.9%estimated ± 4.4 pp, high confidence
AA Terminal-Bench 4.021.6%estimated ± 14.0 pp, low confidence
CursorBench 4.023.4%estimated ± 1.6 pp, medium confidence
CWE-bench v137.0%estimated ± 6.9 pp, low confidence
DeepSWE34.7%estimated ± 6.4 pp, low confidence
HLE w/o tools32.6%estimated ± 5.7 pp, medium confidence
PostTrainBench v1.124.9%estimated ± 2.7 pp, high confidence
ProgramBench53.6%measured
Toolathlon-Verified51.5%estimated ± 2.3 pp, low confidence
Vibe Code Bench34.7%estimated ± 10.3 pp, low confidence
ExploitGym0.0%estimated ± 12.2 pp, low confidence
C-Eval93.1%estimated ± 1.0 pp, medium confidence
Claw-Eval69.9%estimated ± 5.7 pp, medium confidence
DeepPlanning48.2%estimated ± 13.1 pp, low confidence
MCP Atlas76.0%measured
MCP-Tasks77.0%estimated ± 0.6 pp, low confidence
MMLU-Redux96.6%estimated ± 1.1 pp, medium confidence
NL2Repo44.7%estimated ± 4.7 pp, high confidence
QwenClawBench63.0%estimated ± 2.6 pp, high confidence
SuperGPQA80.9%estimated ± 7.5 pp, medium confidence
Terminal-Bench 2.071.5%estimated ± 5.9 pp, medium confidence
Toolathlon50.1%estimated ± 6.3 pp, medium confidence
HealthBench Hard19.6%estimated ± 7.7 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals78.3%estimated ± 2.4 pp, medium confidence
ResearchClawBench18.7%estimated ± 1.8 pp, medium confidence
SWE-Rebench58.1%estimated ± 8.9 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, medium confidence
BioMysteryBench (human-solvable)83.3%estimated ± 0.6 pp, low confidence
DRACO72.6%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended55.1%estimated ± 1.6 pp, low confidence
HealthBench (raw)44.3%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional57.7%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)55.0%estimated ± 5.0 pp, low confidence
HLE-Verified16.6%estimated ± 2.0 pp, low confidence
LABBench276.4%estimated ± 1.8 pp, low confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
OpenHarmony Bench52.1%measured
MMLU93.3%estimated ± 1.1 pp, low confidence