benchgap
Calibration

AA-HLE → LABBench2

LABBench2 is estimated from AA-HLE with a linear curve fitted on 7 models measured on both: y = 0.4257·x + 0.6276, R² = 0.64, cross-validated error 2.4 pp. It is used for 11 estimates.

Estimated modelAA-HLELABBench2Source
Claude Haiku 5.544.4%81.7%estimated ± 2.4 pp, medium confidence
Claude Sonnet 5.555.0%86.2%estimated ± 2.4 pp, medium confidence
DeepSeek V4.1 Flash39.2%79.5%estimated ± 2.4 pp, low confidence
GPT-4 Turbo3.1%64.1%estimated ± 2.4 pp, low confidence
Grok 4.743.1%81.1%estimated ± 2.4 pp, medium confidence
Ling 3.1 Flash39.4%79.5%estimated ± 2.4 pp, low confidence
Mercury 2.511.8%67.8%estimated ± 2.4 pp, low confidence
MiMo-V2.6-Flash35.1%77.7%estimated ± 2.4 pp, low confidence
MiMo-V2.6-Pro49.4%83.8%estimated ± 2.4 pp, medium confidence
Mistral Large 435.0%77.7%estimated ± 2.4 pp, low confidence
Step 5 Preview46.5%82.6%estimated ± 2.4 pp, medium confidence