benchgap
Other · model

Pareto 26.9 benchmark scores

As of 2026-10-07, Pareto 26.9 (Other) has measured scores on 4 benchmarks and estimated scores on 39 more.

BenchmarkScoreSource
AA-GPQA Diamond92.0%estimated ± 2.7 pp, high confidence
AA-HLE47.1%estimated ± 5.0 pp, high confidence
Artificial Analysis Intelligence Index42.4%estimated ± 5.1 pp, medium confidence
CritPt21.6%estimated ± 6.2 pp, medium confidence
GPQA Diamond95.0%estimated ± 1.9 pp, high confidence
HLE57.4%estimated ± 5.7 pp, medium confidence
AA-Omniscience Accuracy50.3%estimated ± 8.6 pp, medium confidence
AA Coding Index77.4%estimated ± 2.5 pp, high confidence
AA-MMMU-Pro76.0%estimated ± 4.9 pp, medium confidence
SWE-bench Verified93.4%estimated ± 5.2 pp, low confidence
SWE-bench Pro77.1%estimated ± 7.8 pp, low confidence
MMLU-Pro85.8%estimated ± 2.5 pp, low confidence
GPQA94.8%estimated ± 2.1 pp, high confidence
ARC-AGI-196.3%estimated ± 3.7 pp, high confidence
ARC-AGI-280.9%estimated ± 13.0 pp, low confidence
CursorBench 3.268.1%estimated ± 6.4 pp, low confidence
FrontierCode 1.1 Main53.3%estimated ± 3.7 pp, low confidence
Vals GPQA Diamond92.3%estimated ± 3.0 pp, high confidence
Vals MMLU-Pro89.7%estimated ± 1.9 pp, high confidence
Terminal-Bench 2.1 (Vals AI)81.8%estimated ± 4.6 pp, medium confidence
AA Terminal-Bench 4.047.1%estimated ± 4.7 pp, high confidence
DeepSWE74.0%measured
HLE w/o tools49.0%measured
PostTrainBench v1.142.6%estimated ± 4.8 pp, high confidence
Terminal-Bench 4.051.0%measured
CharXiv81.0%estimated ± 5.4 pp, medium confidence
CharXiv w/o tools82.3%estimated ± 1.8 pp, medium confidence
MathVision86.2%estimated ± 3.6 pp, high confidence
MMMU-Pro78.0%measured
NL2Repo61.3%estimated ± 6.6 pp, medium confidence
ScreenSpot Pro73.0%estimated ± 10.0 pp, low confidence
VideoMMMU84.8%estimated ± 1.0 pp, high confidence
ERQA62.4%estimated ± 6.1 pp, medium confidence
MedXpertQA (MM)70.2%estimated ± 2.4 pp, high confidence
MedXpertQA (Text)66.1%estimated ± 8.9 pp, low confidence
ARC-AGI-30.5%estimated ± 11.6 pp, low confidence
HealthBench Professional57.0%estimated ± 4.8 pp, medium confidence
MMMU84.4%estimated ± 1.1 pp, medium confidence
SimpleVQA67.0%estimated ± 8.2 pp, medium confidence
MMMU-Pro w/ Python79.4%estimated ± 0.6 pp, high confidence
RealWorldQA86.1%estimated ± 4.8 pp, medium confidence
Video-MME (with subtitle)87.6%estimated ± 1.8 pp, medium confidence
OmniDocBench 1.591.7%estimated ± 5.9 pp, low confidence