benchgap
Other · model

Sakana Fugu benchmark scores

As of 2026-10-07, Sakana Fugu (Other) has measured scores on 9 benchmarks and estimated scores on 52 more.

BenchmarkScoreSource
AA-GPQA Diamond91.7%estimated ± 2.7 pp, high confidence
AA-HLE52.9%estimated ± 4.0 pp, high confidence
AA-SciCode58.8%estimated ± 3.1 pp, medium confidence
Artificial Analysis Intelligence Index40.7%estimated ± 5.1 pp, medium confidence
CritPt30.3%estimated ± 4.1 pp, high confidence
GPQA Diamond95.5%measured
HLE55.9%estimated ± 5.7 pp, medium confidence
AA-LCR74.9%estimated ± 4.5 pp, medium confidence
LiveCodeBench v692.9%measured
AA-Omniscience Accuracy58.1%estimated ± 8.1 pp, medium confidence
SciCode60.1%measured
Terminal-Bench 2.180.2%measured
AA Coding Index66.9%estimated ± 6.4 pp, low confidence
SWE-bench Verified81.7%estimated ± 2.5 pp, high confidence
SWE-bench Pro59.0%measured
MMLU-Pro85.8%estimated ± 2.5 pp, low confidence
GPQA95.5%measured
ARC-AGI-196.5%estimated ± 2.7 pp, high confidence
ARC-AGI-291.5%estimated ± 10.1 pp, low confidence
CursorBench 3.154.0%estimated ± 5.6 pp, low confidence
CursorBench 3.256.8%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.013.6%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main39.6%estimated ± 5.5 pp, medium confidence
FrontierSWE v213.4%estimated ± 9.0 pp, medium confidence
Vals GPQA Diamond94.9%estimated ± 2.0 pp, medium confidence
Vals LiveCodeBench88.8%estimated ± 0.8 pp, low confidence
Vals MMLU-Pro88.8%estimated ± 1.2 pp, medium confidence
Vals SWE-bench78.0%estimated ± 5.3 pp, low confidence
Terminal-Bench 2.1 (Vals AI)67.9%estimated ± 7.4 pp, medium confidence
AA Terminal-Bench 2.184.6%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.018.8%estimated ± 8.9 pp, medium confidence
CursorBench 4.034.7%estimated ± 3.6 pp, medium confidence
HLE w/o tools47.2%measured
PostTrainBench v1.125.0%estimated ± 4.7 pp, high confidence
CharXiv85.1%measured
CharXiv w/o tools78.5%estimated ± 2.3 pp, high confidence
C-Eval96.1%estimated ± 0.9 pp, low confidence
MathVision89.3%estimated ± 2.6 pp, high confidence
MMLU-Redux95.2%estimated ± 2.1 pp, medium confidence
MMMU-Pro79.9%estimated ± 3.2 pp, high confidence
NL2Repo47.5%estimated ± 3.1 pp, low confidence
ScreenSpot Pro80.0%estimated ± 7.1 pp, medium confidence
SuperGPQA83.5%estimated ± 9.2 pp, low confidence
VideoMMMU86.0%estimated ± 1.5 pp, high confidence
ERQA68.7%estimated ± 3.7 pp, high confidence
HealthBench Hard36.6%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro87.8%measured
MedXpertQA (MM)77.1%estimated ± 6.4 pp, low confidence
MedXpertQA (Text)75.1%estimated ± 4.5 pp, low confidence
MMLU-Pro (Arcee)88.7%estimated ± 3.4 pp, low confidence
ARC-AGI-334.2%estimated ± 2.1 pp, high confidence
FrontierCode 1.1 Extended55.7%estimated ± 3.0 pp, low confidence
HealthBench Professional60.7%estimated ± 4.8 pp, medium confidence
OpenHarmony Bench53.3%estimated ± 4.9 pp, low confidence
SimpleVQA69.3%estimated ± 7.8 pp, low confidence
MMMLU91.4%estimated ± 1.1 pp, low confidence
MMLU90.5%estimated ± 3.6 pp, low confidence
MathVision w/ Python76.9%estimated ± 0.9 pp, low confidence
RealWorldQA86.3%estimated ± 1.1 pp, medium confidence
Video-MME (with subtitle)88.4%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.588.1%estimated ± 7.4 pp, low confidence