benchgap
Google · model

Gemini 3.1 Flash-Lite benchmark scores

As of 2026-10-07, Gemini 3.1 Flash-Lite (Google) has measured scores on 9 benchmarks and estimated scores on 64 more.

BenchmarkScoreSource
AA-GPQA Diamond81.9%estimated ± 4.2 pp, high confidence
AA-HLE23.4%estimated ± 7.4 pp, medium confidence
AA-SciCode45.7%estimated ± 3.8 pp, medium confidence
Artificial Analysis Intelligence Index20.5%estimated ± 7.9 pp, medium confidence
CritPt3.4%estimated ± 6.6 pp, medium confidence
GDPval-AA15.6%estimated ± 8.8 pp, medium confidence
GPQA Diamond85.5%estimated ± 1.7 pp, medium confidence
HLE43.5%estimated ± 9.8 pp, low confidence
AA-Omniscience Accuracy33.7%estimated ± 7.7 pp, medium confidence
SciCode37.0%estimated ± 2.9 pp, low confidence
Terminal-Bench 2.116.9%estimated ± 6.9 pp, low confidence
BrowseComp58.7%estimated ± 5.7 pp, medium confidence
AA Coding Index43.1%estimated ± 5.8 pp, medium confidence
APEX-Agents-AA10.6%estimated ± 11.4 pp, low confidence
SWE-bench Verified71.7%estimated ± 4.4 pp, medium confidence
CyberGym30.2%estimated ± 10.0 pp, low confidence
SWE-bench Pro0.0%estimated ± 2.1 pp, high confidence
τ³-bench results97.0%estimated ± 8.9 pp, low confidence
WideResearch54.2%estimated ± 5.9 pp, low confidence
JevBench 1.414.3%measured
AA Agentic Index10.2%estimated ± 6.4 pp, medium confidence
GPQA90.0%estimated ± 1.6 pp, high confidence
Gert Labs38.5%measured
ApprenticeBench0.0%estimated ± 3.1 pp, low confidence
ARC-AGI-171.5%estimated ± 5.9 pp, medium confidence
ARC-AGI-222.4%estimated ± 12.6 pp, low confidence
CursorBench 3.136.5%estimated ± 5.8 pp, low confidence
CursorBench 3.248.6%estimated ± 5.4 pp, low confidence
Terminal-Bench 3.00.2%estimated ± 6.9 pp, low confidence
FrontierCode 1.1 Main0.0%estimated ± 4.3 pp, low confidence
FrontierSWE v24.7%estimated ± 9.5 pp, low confidence
OSWorld-Verified53.8%estimated ± 5.9 pp, medium confidence
Vals GPQA Diamond81.1%measured
Vals LiveCodeBench80.1%measured
Vals MMLU-Pro86.2%measured
Vals SWE-bench62.8%measured
Terminal-Bench 2.1 (Vals AI)34.1%measured
VulcanBench v380.9%estimated ± 5.6 pp, low confidence
AA Terminal-Bench 2.114.0%estimated ± 4.4 pp, medium confidence
AA Terminal-Bench 4.00.0%estimated ± 14.0 pp, low confidence
CursorBench 4.021.4%estimated ± 4.4 pp, low confidence
DeepSWE56.3%estimated ± 7.2 pp, low confidence
HLE w/o tools34.2%estimated ± 5.1 pp, medium confidence
OSWorld 2.00.0%estimated ± 4.5 pp, low confidence
PostTrainBench v1.115.8%estimated ± 2.7 pp, high confidence
Vibe Code Bench0.0%measured
CharXiv73.2%measured
CharXiv w/o tools77.2%estimated ± 2.3 pp, medium confidence
Claw-Eval54.5%estimated ± 6.8 pp, low confidence
MathVision81.9%estimated ± 2.6 pp, high confidence
MCP Atlas43.5%estimated ± 14.8 pp, low confidence
MMMU-Pro75.4%estimated ± 3.2 pp, high confidence
NL2Repo38.5%estimated ± 4.7 pp, medium confidence
ScreenSpot Pro58.3%estimated ± 7.1 pp, medium confidence
Terminal-Bench 2.044.2%estimated ± 5.9 pp, medium confidence
Toolathlon33.8%estimated ± 7.2 pp, low confidence
VideoMMMU84.0%estimated ± 1.5 pp, high confidence
ERQA62.1%estimated ± 3.7 pp, high confidence
MedXpertQA (MM)72.3%estimated ± 6.4 pp, low confidence
React Native Evals55.8%estimated ± 2.4 pp, low confidence
ResearchClawBench14.0%estimated ± 2.2 pp, medium confidence
SWE-Rebench0.0%estimated ± 6.5 pp, low confidence
ARC-AGI-30.0%estimated ± 9.2 pp, low confidence
FrontierCode 1.1 Extended8.8%estimated ± 2.4 pp, low confidence
HealthBench Professional56.3%estimated ± 1.3 pp, medium confidence
HLE-Verified0.0%estimated ± 4.6 pp, low confidence
LABBench281.0%estimated ± 1.8 pp, low confidence
OpenHarmony Bench48.9%estimated ± 3.2 pp, medium confidence
SimpleVQA62.6%estimated ± 7.8 pp, low confidence
MathVision w/ Python76.8%estimated ± 0.9 pp, low confidence
RealWorldQA83.7%estimated ± 1.1 pp, low confidence
Video-MME (with subtitle)86.0%estimated ± 0.9 pp, low confidence
OmniDocBench 1.580.7%estimated ± 7.4 pp, low confidence