benchgap
NVIDIA · model

Nemotron 3.5 Lightning 30B A3B NVFP4 benchmark scores

As of 2026-10-07, Nemotron 3.5 Lightning 30B A3B NVFP4 (NVIDIA) has measured scores on 21 benchmarks and estimated scores on 80 more.

BenchmarkScoreSource
AA-GPQA Diamond74.3%measured
AA-HLE10.6%measured
AA-SciCode32.1%measured
Artificial Analysis Intelligence Index12.9%measured
CritPt0.0%measured
GDPval-AA7.0%measured
GPQA Diamond75.6%measured
HLE15.8%estimated ± 5.7 pp, medium confidence
IFBench72.9%measured
AA-LCR49.2%measured
LiveCodeBench v674.3%estimated ± 3.1 pp, high confidence
AA-Omniscience Accuracy14.4%measured
SciCode31.4%measured
Terminal-Bench 2.123.5%measured
BrowseComp36.8%measured
HLE w/ tools33.2%estimated ± 4.6 pp, medium confidence
IFEval92.8%estimated ± 3.5 pp, high confidence
LongBench v20.0%estimated ± 5.9 pp, low confidence
VITA-Bench5.2%estimated ± 9.9 pp, low confidence
τ²-bench results79.9%estimated ± 13.3 pp, low confidence
AA Coding Index26.8%measured
APEX-Agents7.0%estimated ± 10.0 pp, low confidence
APEX-Agents-AA6.0%estimated ± 4.6 pp, medium confidence
BioMysteryBench (human-difficult)21.4%estimated ± 5.9 pp, low confidence
DeepSearchQA61.7%estimated ± 4.5 pp, medium confidence
SWE-bench Verified52.8%measured
AutomationBench28.4%estimated ± 9.7 pp, low confidence
BFCL v465.5%estimated ± 2.6 pp, high confidence
CyberGym44.8%estimated ± 4.6 pp, medium confidence
JobBench0.0%estimated ± 8.6 pp, low confidence
SWE-bench Pro27.2%estimated ± 3.2 pp, high confidence
τ³-bench results9.5%measured
WideResearch50.8%estimated ± 4.6 pp, medium confidence
AA Agentic Index6.1%measured
MMLU-Pro81.6%measured
GPQA75.6%measured
AA-IFBench58.4%estimated ± 8.5 pp, medium confidence
Gert Labs20.9%estimated ± 7.4 pp, low confidence
AA-AnalystAgent4.4%estimated ± 12.0 pp, low confidence
AA EnterpriseOps-Gym32.5%estimated ± 6.7 pp, low confidence
AA Harvey LAB79.4%estimated ± 1.2 pp, low confidence
ARC-AGI-160.3%estimated ± 2.7 pp, medium confidence
ARC-AGI-224.3%estimated ± 10.1 pp, low confidence
CursorBench 3.10.0%estimated ± 3.4 pp, low confidence
CursorBench 3.215.9%estimated ± 4.4 pp, medium confidence
Terminal-Bench 3.01.1%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main0.1%estimated ± 2.7 pp, low confidence
FrontierSWE v20.0%estimated ± 14.6 pp, low confidence
OSWorld-Verified46.2%estimated ± 5.4 pp, low confidence
Vals GPQA Diamond74.4%estimated ± 2.0 pp, medium confidence
Vals LiveCodeBench81.7%estimated ± 0.8 pp, low confidence
Vals MMLU-Pro50.2%estimated ± 1.2 pp, medium confidence
Vals SWE-bench38.4%estimated ± 5.3 pp, low confidence
Terminal-Bench 2.1 (Vals AI)24.1%estimated ± 7.4 pp, low confidence
AA AutomationBench2.6%estimated ± 5.7 pp, low confidence
GDP.pdf9.4%estimated ± 5.4 pp, low confidence
AA Tau3 Banking22.0%estimated ± 5.8 pp, low confidence
AA Terminal-Bench 2.152.1%estimated ± 3.5 pp, medium confidence
AA Terminal-Bench 4.00.0%estimated ± 8.9 pp, low confidence
CursorBench 4.06.3%estimated ± 1.6 pp, medium confidence
CWE-bench v111.9%estimated ± 6.9 pp, low confidence
DeepSWE0.4%estimated ± 6.4 pp, low confidence
HLE w/o tools10.5%measured
OSWorld 2.017.7%estimated ± 13.8 pp, low confidence
PostTrainBench v1.18.4%estimated ± 6.1 pp, low confidence
SWE Multilingual36.5%measured
Toolathlon-Verified20.8%estimated ± 2.3 pp, low confidence
Vibe Code Bench0.6%estimated ± 8.1 pp, low confidence
ExploitGym0.0%estimated ± 6.1 pp, low confidence
C-Eval84.5%estimated ± 0.9 pp, low confidence
Claw-Eval62.4%estimated ± 7.9 pp, low confidence
DeepPlanning13.0%estimated ± 8.3 pp, low confidence
MCP Atlas64.6%estimated ± 7.8 pp, low confidence
MCP-Tasks0.0%estimated ± 3.6 pp, low confidence
MMLU-ProX68.6%estimated ± 1.7 pp, low confidence
MMLU-Redux91.2%estimated ± 0.7 pp, high confidence
NL2Repo18.0%estimated ± 5.1 pp, low confidence
NOVA-6335.7%estimated ± 1.6 pp, low confidence
QwenClawBench55.9%estimated ± 6.0 pp, low confidence
SuperGPQA55.2%estimated ± 7.5 pp, medium confidence
Toolathlon3.0%estimated ± 3.5 pp, medium confidence
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
MedXpertQA (Text)27.7%estimated ± 4.5 pp, low confidence
MMLU-Pro (Arcee)81.4%estimated ± 3.4 pp, medium confidence
React Native Evals70.8%estimated ± 3.5 pp, high confidence
ResearchClawBench9.0%estimated ± 2.0 pp, medium confidence
SWE-Rebench57.0%estimated ± 3.6 pp, low confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
BioMysteryBench (human-solvable)81.7%estimated ± 0.6 pp, low confidence
DRACO64.4%estimated ± 3.0 pp, low confidence
FrontierCode 1.1 Extended26.3%estimated ± 3.4 pp, low confidence
HealthBench (raw)15.9%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional24.1%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)29.1%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench255.7%estimated ± 1.8 pp, low confidence
LiveCodeBench45.2%estimated ± 8.9 pp, low confidence
OpenHarmony Bench51.0%estimated ± 2.4 pp, medium confidence
MMMLU80.9%estimated ± 1.1 pp, low confidence
MMLU81.3%estimated ± 1.1 pp, medium confidence