benchgap
Other · model

dots3-note Preview benchmark scores

As of 2026-10-07, dots3-note Preview (Other) has measured scores on 27 benchmarks and estimated scores on 82 more.

BenchmarkScoreSource
AA-GPQA Diamond93.2%estimated ± 2.7 pp, high confidence
AA-HLE47.3%estimated ± 6.3 pp, medium confidence
AA-SciCode50.8%estimated ± 3.0 pp, high confidence
Artificial Analysis Intelligence Index45.8%estimated ± 7.0 pp, medium confidence
CritPt23.0%estimated ± 5.7 pp, medium confidence
GDPval-AA24.0%estimated ± 3.7 pp, low confidence
GPQA Diamond93.7%estimated ± 1.5 pp, high confidence
HLE52.6%measured
IFBench80.4%measured
LiveCodeBench v691.5%measured
AA-Omniscience Accuracy55.2%estimated ± 6.2 pp, medium confidence
SciCode47.8%estimated ± 3.8 pp, high confidence
Terminal-Bench 2.175.1%measured
BrowseComp83.3%measured
HLE w/ tools52.6%measured
IFEval93.9%measured
VITA-Bench45.3%estimated ± 9.9 pp, low confidence
AA Coding Index68.1%estimated ± 4.6 pp, high confidence
AA-MMMU-Pro79.7%estimated ± 3.0 pp, medium confidence
APEX-Agents30.8%measured
APEX-Agents-AA30.8%estimated ± 4.6 pp, high confidence
DeepSearchQA92.1%measured
SWE-bench Verified78.4%measured
AutomationBench31.8%estimated ± 10.3 pp, low confidence
BFCL v493.7%estimated ± 8.6 pp, low confidence
CyberGym77.8%estimated ± 2.4 pp, medium confidence
JobBench19.9%estimated ± 3.4 pp, low confidence
SkillsBench52.8%measured
SWE-bench Pro61.0%measured
τ³-bench results67.9%estimated ± 11.4 pp, low confidence
WideResearch78.9%measured
AA Agentic Index26.5%estimated ± 3.3 pp, medium confidence
MMLU-Pro86.3%estimated ± 3.6 pp, high confidence
GPQA93.6%estimated ± 1.6 pp, high confidence
AA-IFBench72.8%estimated ± 8.2 pp, medium confidence
Gert Labs54.5%estimated ± 6.1 pp, low confidence
AA-AnalystAgent41.9%estimated ± 2.5 pp, medium confidence
AA EnterpriseOps-Gym42.1%estimated ± 6.9 pp, low confidence
AA Harvey LAB88.7%estimated ± 1.7 pp, medium confidence
ApprenticeBench23.1%estimated ± 10.1 pp, low confidence
ARC-AGI-195.0%estimated ± 1.9 pp, high confidence
ARC-AGI-281.4%measured
CursorBench 3.155.3%estimated ± 5.6 pp, low confidence
CursorBench 3.259.4%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.010.3%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main19.9%estimated ± 2.7 pp, low confidence
FrontierSWE v216.9%estimated ± 9.0 pp, medium confidence
OSWorld-Verified76.6%estimated ± 5.4 pp, medium confidence
Vals GPQA Diamond93.4%estimated ± 2.2 pp, high confidence
Vals LiveCodeBench85.2%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro89.9%estimated ± 2.0 pp, high confidence
Vals SWE-bench83.8%estimated ± 5.4 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)64.0%estimated ± 7.4 pp, medium confidence
AA AutomationBench20.2%estimated ± 5.7 pp, low confidence
GDP.pdf9.5%estimated ± 6.6 pp, low confidence
AA ITBench53.8%estimated ± 3.4 pp, low confidence
AA Tau3 Banking30.9%estimated ± 10.2 pp, low confidence
AA Terminal-Bench 2.182.9%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.012.8%estimated ± 8.9 pp, medium confidence
CursorBench 4.036.2%estimated ± 3.6 pp, medium confidence
DeepSWE49.3%estimated ± 9.6 pp, medium confidence
HLE w/o tools42.9%estimated ± 5.3 pp, medium confidence
OSWorld 2.019.4%estimated ± 5.5 pp, low confidence
PostTrainBench v1.126.3%estimated ± 4.7 pp, high confidence
SWE Multilingual75.7%measured
Toolathlon-Verified55.6%measured
Vibe Code Bench28.1%estimated ± 8.1 pp, medium confidence
CharXiv89.5%estimated ± 2.5 pp, high confidence
CharXiv w/o tools83.1%measured
ExploitGym10.1%estimated ± 6.1 pp, low confidence
C-Eval93.3%estimated ± 0.8 pp, low confidence
Claw-Eval73.4%measured
HMMT Feb 202697.3%estimated ± 9.7 pp, low confidence
MathVision87.7%measured
MCP Atlas77.0%estimated ± 0.7 pp, medium confidence
MCP-Tasks79.1%estimated ± 6.0 pp, low confidence
MMLU-ProX85.0%estimated ± 1.7 pp, medium confidence
MMLU-Redux96.8%estimated ± 2.1 pp, low confidence
MMMU-Pro79.1%measured
NL2Repo49.8%measured
NOVA-6357.4%estimated ± 1.6 pp, low confidence
ScreenSpot Pro84.6%estimated ± 2.5 pp, medium confidence
SuperGPQA83.8%estimated ± 9.1 pp, medium confidence
Toolathlon53.9%estimated ± 3.5 pp, high confidence
V*93.6%estimated ± 5.1 pp, medium confidence
VideoMMMU86.8%measured
ERQA65.6%estimated ± 2.4 pp, low confidence
LiveCodeBench Pro84.5%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)72.6%estimated ± 2.4 pp, high confidence
ResearchClawBench19.7%estimated ± 1.3 pp, high confidence
SWE-Rebench61.6%estimated ± 3.6 pp, medium confidence
ARC-AGI-30.3%estimated ± 11.5 pp, low confidence
FrontierCode 1.1 Extended56.3%estimated ± 3.0 pp, low confidence
HealthBench Professional55.8%estimated ± 4.6 pp, high confidence
HLE-Verified35.9%estimated ± 0.7 pp, low confidence
LABBench280.7%estimated ± 2.5 pp, low confidence
MMMU85.1%estimated ± 1.1 pp, low confidence
LiveCodeBench85.5%estimated ± 10.2 pp, low confidence
OpenHarmony Bench55.2%estimated ± 2.7 pp, high confidence
IMOAnswerBench90.9%measured
BabyVision50.0%measured
MMVU79.9%measured
SimpleVQA72.5%measured
ZeroBench19.0%measured
MMMU-Pro w/ Python80.4%estimated ± 0.6 pp, high confidence
MathVision w/ Python93.8%estimated ± 0.5 pp, low confidence
RealWorldQA26.0%estimated ± 0.8 pp, low confidence
Video-MME (with subtitle)88.8%estimated ± 1.2 pp, medium confidence
OmniDocBench 1.591.7%estimated ± 5.9 pp, low confidence