benchgap
Calibration

ARC-AGI-1 → HealthBench (length-adjusted)

HealthBench (length-adjusted) is estimated from ARC-AGI-1 with a Michaelis–Menten curve fitted on 6 models measured on both: y = 1.5144·x / (1.57246 + x), R² = 0.36, cross-validated error 2.5 pp. It is used for 20 estimates.

Estimated modelARC-AGI-1HealthBench (length-adjusted)Source
Claude Fable 598.5%58.3%estimated ± 2.5 pp, low confidence
Claude Fable 5.197.5%58.0%estimated ± 2.5 pp, low confidence
Claude Opus 4.5 Thinking80.0%51.1%estimated ± 2.5 pp, low confidence
Claude Opus 4.6 (Adaptive)93.0%56.3%estimated ± 2.5 pp, low confidence
Claude Sonnet 4.5 Thinking63.7%43.6%estimated ± 2.5 pp, low confidence
Claude Sonnet 4.686.0%53.5%estimated ± 2.5 pp, low confidence
DeepSeek V4 Flash 073189.0%54.7%estimated ± 2.5 pp, low confidence
DeepSeek V4 Pro 081390.0%55.1%estimated ± 2.5 pp, low confidence
Gemini 3.5 Flash92.5%56.1%estimated ± 2.5 pp, low confidence
Gemini 3.6 Flash91.2%55.6%estimated ± 2.5 pp, low confidence
Gemini 3.7 Flash95.5%57.2%estimated ± 2.5 pp, low confidence
Gemini 3 Pro75.0%48.9%estimated ± 2.5 pp, low confidence
GPT-5.172.8%47.9%estimated ± 2.5 pp, low confidence
GPT-5.286.2%53.6%estimated ± 2.5 pp, low confidence
GPT-5.4 mini63.7%43.7%estimated ± 2.5 pp, low confidence
GPT-5.4 nano51.5%37.4%estimated ± 2.5 pp, low confidence
GPT-5.4 Pro94.5%56.8%estimated ± 2.5 pp, low confidence
GPT-5.5 Pro95.0%57.0%estimated ± 2.5 pp, low confidence
Inkling-Small84.0%52.7%estimated ± 2.5 pp, low confidence
Kimi K394.5%56.8%estimated ± 2.5 pp, low confidence