benchgap
Mistral · model

Mistral Large 4 benchmark scores

As of 2026-10-07, Mistral Large 4 (Mistral) has measured scores on 12 benchmarks and estimated scores on 85 more.

BenchmarkScoreSource
AA-GPQA Diamond90.8%estimated ± 6.8 pp, medium confidence
AA-HLE35.0%measured
AA-SciCode54.2%measured
Artificial Analysis Intelligence Index38.4%measured
CritPt10.6%measured
GDPval-AA46.2%measured
GPQA Diamond89.2%estimated ± 2.2 pp, high confidence
HLE44.4%estimated ± 9.5 pp, medium confidence
AA-LCR81.3%measured
LiveCodeBench v695.2%estimated ± 3.3 pp, medium confidence
AA-Omniscience Accuracy25.8%measured
SciCode53.8%estimated ± 3.2 pp, high confidence
Terminal-Bench 2.187.7%estimated ± 3.9 pp, high confidence
BrowseComp82.7%estimated ± 10.1 pp, low confidence
HLE w/ tools55.3%estimated ± 6.1 pp, medium confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
AA Coding Index70.6%estimated ± 6.4 pp, medium confidence
AA-MMMU-Pro76.4%measured
APEX-Agents38.6%estimated ± 10.0 pp, low confidence
APEX-Agents-AA34.6%estimated ± 9.1 pp, medium confidence
BioMysteryBench (human-difficult)31.9%estimated ± 5.9 pp, low confidence
DeepSearchQA90.1%estimated ± 8.4 pp, medium confidence
SWE-bench Verified86.4%estimated ± 4.4 pp, medium confidence
AutomationBench45.5%estimated ± 7.6 pp, low confidence
BFCL v477.1%estimated ± 2.6 pp, medium confidence
CyberGym78.8%estimated ± 6.2 pp, medium confidence
JobBench47.3%estimated ± 8.6 pp, medium confidence
SWE-bench Pro65.7%estimated ± 6.8 pp, medium confidence
WideResearch96.5%estimated ± 9.1 pp, low confidence
AA Agentic Index41.4%estimated ± 2.8 pp, high confidence
MMLU-Pro85.6%estimated ± 3.5 pp, high confidence
GPQA91.8%estimated ± 3.2 pp, high confidence
Gert Labs69.3%estimated ± 8.9 pp, medium confidence
AA-AnalystAgent44.0%estimated ± 6.5 pp, medium confidence
AA EnterpriseOps-Gym43.6%estimated ± 4.4 pp, high confidence
ARC-AGI-192.6%estimated ± 5.4 pp, medium confidence
ARC-AGI-271.1%estimated ± 14.2 pp, low confidence
CursorBench 3.154.4%estimated ± 3.4 pp, medium confidence
CursorBench 3.260.4%estimated ± 5.3 pp, medium confidence
FrontierCode 1.1 Main47.7%estimated ± 4.3 pp, low confidence
FrontierSWE v224.6%estimated ± 14.6 pp, low confidence
OfficeQA Pro50.3%estimated ± 7.2 pp, medium confidence
OSWorld-Verified79.8%estimated ± 8.7 pp, low confidence
Vals GPQA Diamond91.2%estimated ± 7.5 pp, medium confidence
Vals MMLU-Pro84.2%estimated ± 2.2 pp, high confidence
Vals SWE-bench82.1%estimated ± 6.3 pp, low confidence
Terminal-Bench 2.1 (Vals AI)75.6%estimated ± 8.5 pp, medium confidence
AA AutomationBench59.9%measured
GDP.pdf18.6%measured
AA Tau3 Banking46.9%estimated ± 5.8 pp, medium confidence
AA Terminal-Bench 2.186.5%estimated ± 2.6 pp, high confidence
AA Terminal-Bench 4.026.8%measured
CursorBench 4.037.3%estimated ± 3.5 pp, high confidence
CWE-bench v156.2%estimated ± 6.0 pp, medium confidence
HLE w/o tools32.6%estimated ± 5.7 pp, medium confidence
OSWorld 2.055.2%estimated ± 13.3 pp, low confidence
PostTrainBench v1.130.4%estimated ± 7.2 pp, low confidence
Terminal-Bench 4.030.5%estimated ± 5.1 pp, medium confidence
Terminal-Bench-Science 0.14.7%estimated ± 4.1 pp, low confidence
Toolathlon-Verified76.1%estimated ± 2.2 pp, medium confidence
Vibe Code Bench78.4%measured
CharXiv83.8%estimated ± 7.5 pp, low confidence
CharXiv w/o tools80.5%estimated ± 2.4 pp, high confidence
ExploitGym15.0%estimated ± 5.0 pp, low confidence
C-Eval95.5%estimated ± 1.0 pp, low confidence
MathVision89.1%estimated ± 4.9 pp, high confidence
MCP Atlas85.5%estimated ± 2.7 pp, low confidence
MMLU-Redux94.7%estimated ± 1.1 pp, medium confidence
MMMU-Pro77.6%estimated ± 4.4 pp, medium confidence
NL2Repo56.5%estimated ± 6.5 pp, medium confidence
QwenClawBench100.0%estimated ± 6.0 pp, low confidence
ScreenSpot Pro75.2%estimated ± 10.7 pp, low confidence
SuperGPQA70.8%estimated ± 7.5 pp, medium confidence
Toolathlon55.9%estimated ± 6.3 pp, medium confidence
ERQA65.9%estimated ± 3.9 pp, high confidence
HealthBench Hard32.9%estimated ± 7.7 pp, low confidence
MedXpertQA (MM)74.3%estimated ± 6.5 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals85.4%estimated ± 3.0 pp, low confidence
ResearchClawBench19.2%estimated ± 2.2 pp, medium confidence
SWE-Rebench62.2%estimated ± 6.5 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, medium confidence
BioMysteryBench (human-solvable)83.8%estimated ± 0.6 pp, low confidence
DRACO80.5%estimated ± 3.0 pp, medium confidence
HealthBench (raw)44.3%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, medium confidence
HealthBench Professional57.7%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)55.0%estimated ± 5.0 pp, low confidence
HLE-Verified20.3%estimated ± 6.0 pp, low confidence
LABBench277.7%estimated ± 2.4 pp, low confidence
MMMU84.2%estimated ± 0.9 pp, high confidence
Agents' Last Exam39.0%estimated ± 6.3 pp, medium confidence
OpenHarmony Bench58.8%estimated ± 2.4 pp, high confidence
MMLU89.3%estimated ± 1.1 pp, medium confidence
MMMU-Pro w/ Python79.5%estimated ± 1.6 pp, high confidence
RealWorldQA85.7%estimated ± 1.2 pp, medium confidence
OmniDocBench 1.591.7%estimated ± 6.4 pp, low confidence