benchgap
Other · model

Sakana Fugu-Ultra benchmark scores

As of 2026-10-07, Sakana Fugu-Ultra (Other) has measured scores on 9 benchmarks and estimated scores on 52 more.

BenchmarkScoreSource
AA-GPQA Diamond92.2%estimated ± 2.7 pp, high confidence
AA-HLE52.9%estimated ± 4.0 pp, high confidence
AA-SciCode57.5%estimated ± 3.1 pp, high confidence
Artificial Analysis Intelligence Index43.3%estimated ± 5.1 pp, medium confidence
CritPt30.3%estimated ± 4.1 pp, high confidence
GPQA Diamond95.5%measured
HLE58.3%estimated ± 5.7 pp, medium confidence
AA-LCR75.9%estimated ± 4.5 pp, low confidence
LiveCodeBench v693.2%measured
AA-Omniscience Accuracy58.1%estimated ± 8.1 pp, medium confidence
SciCode58.7%measured
Terminal-Bench 2.182.1%measured
AA Coding Index67.5%estimated ± 6.4 pp, low confidence
SWE-bench Verified91.1%estimated ± 2.5 pp, high confidence
SWE-bench Pro73.7%measured
MMLU-Pro85.8%estimated ± 2.5 pp, low confidence
GPQA95.5%measured
ARC-AGI-196.5%estimated ± 2.7 pp, high confidence
ARC-AGI-291.5%estimated ± 10.1 pp, low confidence
CursorBench 3.163.2%estimated ± 5.6 pp, low confidence
CursorBench 3.269.1%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.015.2%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main47.1%estimated ± 5.5 pp, medium confidence
FrontierSWE v239.5%estimated ± 9.0 pp, medium confidence
Vals GPQA Diamond94.9%estimated ± 2.0 pp, medium confidence
Vals LiveCodeBench88.5%estimated ± 0.8 pp, low confidence
Vals MMLU-Pro88.8%estimated ± 1.2 pp, medium confidence
Vals SWE-bench77.7%estimated ± 5.3 pp, low confidence
Terminal-Bench 2.1 (Vals AI)69.4%estimated ± 7.4 pp, medium confidence
AA Terminal-Bench 2.184.8%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.021.1%estimated ± 8.9 pp, medium confidence
CursorBench 4.046.0%estimated ± 3.6 pp, high confidence
HLE w/o tools50.0%measured
PostTrainBench v1.134.9%estimated ± 4.7 pp, high confidence
CharXiv86.6%measured
CharXiv w/o tools79.2%estimated ± 2.3 pp, high confidence
C-Eval96.1%estimated ± 0.9 pp, low confidence
MathVision90.1%estimated ± 2.6 pp, high confidence
MMLU-Redux95.2%estimated ± 2.1 pp, medium confidence
MMMU-Pro80.5%estimated ± 3.2 pp, high confidence
NL2Repo47.7%estimated ± 3.1 pp, low confidence
ScreenSpot Pro82.0%estimated ± 7.1 pp, medium confidence
SuperGPQA83.5%estimated ± 9.2 pp, low confidence
VideoMMMU86.3%estimated ± 1.5 pp, high confidence
ERQA69.5%estimated ± 3.7 pp, high confidence
HealthBench Hard36.6%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro90.8%measured
MedXpertQA (MM)77.7%estimated ± 6.4 pp, low confidence
MedXpertQA (Text)75.1%estimated ± 4.5 pp, low confidence
MMLU-Pro (Arcee)88.7%estimated ± 3.4 pp, low confidence
ARC-AGI-334.2%estimated ± 2.1 pp, high confidence
FrontierCode 1.1 Extended59.9%estimated ± 3.0 pp, medium confidence
HealthBench Professional60.7%estimated ± 4.8 pp, medium confidence
OpenHarmony Bench60.9%estimated ± 4.9 pp, low confidence
SimpleVQA70.1%estimated ± 7.8 pp, low confidence
MMMLU91.4%estimated ± 1.1 pp, low confidence
MMLU90.5%estimated ± 3.6 pp, low confidence
MathVision w/ Python77.6%estimated ± 0.9 pp, low confidence
RealWorldQA86.6%estimated ± 1.1 pp, medium confidence
Video-MME (with subtitle)88.7%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.588.9%estimated ± 7.4 pp, low confidence