benchgap
Calibration

AutomationBench → Harvey LAB

Harvey LAB is estimated from AutomationBench with a linear curve fitted on 8 models measured on both: y = 0.1546·x + 0.8358, R² = 0.77, cross-validated error 3.2 pp. It is used for 8 estimates.

Estimated modelAutomationBenchHarvey LABSource
GPT-6 Astra (high)66.6%93.9%estimated ± 3.2 pp, high confidence
GPT-6 Astra (xhigh)67.2%94.0%estimated ± 3.2 pp, high confidence
Gemini 4 Argon (high)77.5%95.6%estimated ± 3.2 pp, medium confidence
MiMo-V2.6-Pro58.6%92.6%estimated ± 3.2 pp, high confidence
Mistral Large 4 Preview59.9%92.8%estimated ± 3.2 pp, high confidence
GPT-6 Luna (max)53.2%91.8%estimated ± 3.2 pp, high confidence
GPT-6.1 Sol (xhigh)66.6%93.9%estimated ± 3.2 pp, high confidence
Grok 4.6 (xhigh)67.0%93.9%estimated ± 3.2 pp, high confidence