benchgap
Other · model

Hy4 preview benchmark scores

As of 2026-10-07, Hy4 preview (Other) has measured scores on 25 benchmarks and estimated scores on 69 more.

BenchmarkScoreSource
AA-GPQA Diamond90.9%estimated ± 2.7 pp, high confidence
AA-HLE41.5%estimated ± 4.0 pp, high confidence
AA-SciCode52.7%estimated ± 3.0 pp, high confidence
Artificial Analysis Intelligence Index40.4%estimated ± 5.1 pp, medium confidence
CritPt16.9%measured
GDPval-AA25.4%estimated ± 3.7 pp, low confidence
GPQA Diamond92.3%measured
HLE55.4%measured
LiveCodeBench v692.1%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy45.0%estimated ± 7.9 pp, medium confidence
SciCode54.1%estimated ± 4.8 pp, high confidence
Terminal-Bench 2.185.4%measured
BrowseComp83.3%estimated ± 5.2 pp, low confidence
HLE w/ tools55.4%measured
VITA-Bench49.6%estimated ± 9.9 pp, low confidence
AA Coding Index74.4%estimated ± 2.5 pp, high confidence
AA-MMMU-Pro83.7%estimated ± 2.9 pp, high confidence
APEX-Agents37.1%measured
APEX-Agents-AA41.7%estimated ± 3.6 pp, low confidence
BioMysteryBench (human-difficult)47.2%estimated ± 6.8 pp, low confidence
DeepSearchQA90.2%estimated ± 10.9 pp, low confidence
SWE-bench Verified82.9%estimated ± 2.4 pp, high confidence
AutomationBench32.1%measured
CyberGym78.4%measured
JobBench61.7%measured
SkillsBench62.9%measured
SWE-bench Pro65.7%measured
τ³-bench results69.6%estimated ± 1.8 pp, low confidence
WideResearch83.9%measured
AA Agentic Index50.6%estimated ± 3.3 pp, medium confidence
MMLU-Pro85.7%estimated ± 2.5 pp, low confidence
GPQA92.3%measured
Gert Labs67.4%estimated ± 6.0 pp, low confidence
AA-AnalystAgent48.4%estimated ± 2.5 pp, medium confidence
AA EnterpriseOps-Gym48.3%estimated ± 5.1 pp, low confidence
AA Harvey LAB89.6%estimated ± 1.7 pp, medium confidence
ApprenticeBench8.8%estimated ± 4.0 pp, medium confidence
ARC-AGI-192.4%estimated ± 2.7 pp, high confidence
ARC-AGI-267.1%estimated ± 10.1 pp, low confidence
CursorBench 3.158.3%estimated ± 5.6 pp, low confidence
CursorBench 3.264.2%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.03.5%estimated ± 6.9 pp, low confidence
FrontierCode 1.1 Main43.0%estimated ± 3.7 pp, low confidence
FrontierSWE v225.3%estimated ± 9.0 pp, medium confidence
OfficeQA Pro66.2%measured
OSWorld-Verified78.6%estimated ± 4.7 pp, high confidence
Vals GPQA Diamond91.6%estimated ± 2.0 pp, high confidence
Vals LiveCodeBench86.7%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro87.9%estimated ± 1.2 pp, high confidence
Vals SWE-bench92.1%estimated ± 5.4 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)55.1%measured
AA AutomationBench56.8%estimated ± 5.0 pp, medium confidence
GDP.pdf11.6%estimated ± 3.4 pp, medium confidence
AA ITBench51.4%estimated ± 3.4 pp, medium confidence
AA Tau3 Banking47.4%estimated ± 10.7 pp, low confidence
AA Terminal-Bench 2.184.9%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.025.2%estimated ± 8.9 pp, medium confidence
CursorBench 4.039.8%estimated ± 3.6 pp, high confidence
CWE-bench v153.0%measured
DeepSWE64.3%measured
HLE w/o tools43.4%measured
OSWorld 2.019.4%estimated ± 5.5 pp, low confidence
PostTrainBench v1.129.3%estimated ± 4.7 pp, high confidence
ProgramBench17.5%measured
SWE Multilingual82.9%measured
Toolathlon-Verified74.1%measured
ExploitGym9.7%estimated ± 9.3 pp, low confidence
C-Eval93.3%estimated ± 0.8 pp, low confidence
Claw-Eval66.2%estimated ± 5.6 pp, low confidence
DeepPlanning51.8%estimated ± 13.1 pp, low confidence
MCP Atlas83.7%measured
MCP-Tasks77.1%estimated ± 0.6 pp, low confidence
MMLU-ProX87.7%estimated ± 1.7 pp, low confidence
MMLU-Redux96.9%estimated ± 2.1 pp, low confidence
NL2Repo58.9%measured
NOVA-6360.4%estimated ± 1.6 pp, low confidence
QwenClawBench69.0%estimated ± 2.6 pp, medium confidence
SuperGPQA84.8%estimated ± 9.1 pp, low confidence
Terminal-Bench 2.062.8%estimated ± 5.9 pp, medium confidence
Toolathlon52.5%estimated ± 5.2 pp, low confidence
HealthBench Hard31.2%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro88.0%estimated ± 6.0 pp, low confidence
MedXpertQA (Text)61.3%estimated ± 4.5 pp, medium confidence
MMLU-Pro (Arcee)87.6%estimated ± 3.4 pp, low confidence
ResearchClawBench19.5%estimated ± 1.8 pp, low confidence
ARC-AGI-30.1%estimated ± 2.1 pp, high confidence
BioMysteryBench (human-solvable)88.4%estimated ± 0.6 pp, medium confidence
DRACO77.2%measured
FrontierCode 1.1 Extended57.8%estimated ± 3.0 pp, medium confidence
HealthBench Professional54.6%estimated ± 4.8 pp, medium confidence
Agents' Last Exam22.8%measured
OpenHarmony Bench59.1%estimated ± 2.7 pp, high confidence
MMMLU90.6%estimated ± 1.1 pp, medium confidence
MMLU90.4%estimated ± 3.6 pp, low confidence