benchgap
Calibration

ARC-AGI-2 → HLE-Verified

HLE-Verified is estimated from ARC-AGI-2 with a offset logistic curve fitted on 5 models measured on both: y = 0.3534 + (0.5460 − 0.3534) / (1 + exp(−200.00·(x − 0.8315))), R² = 0.99, cross-validated error 0.7 pp. It is used for 36 estimates.

Estimated modelARC-AGI-2HLE-VerifiedSource
Claude Fable 589.2%54.6%estimated ± 0.7 pp, medium confidence
Claude Fable 5.190.0%54.6%estimated ± 0.7 pp, medium confidence
Claude Opus 4.5 Thinking37.6%35.3%estimated ± 0.7 pp, low confidence
Claude Opus 4.6 (Adaptive)68.8%35.3%estimated ± 0.7 pp, low confidence
Claude Opus 4.7 (Adaptive)75.8%35.3%estimated ± 0.7 pp, low confidence
Claude Opus 4.872.1%35.3%estimated ± 0.7 pp, low confidence
Claude Opus 5.591.7%54.6%estimated ± 0.7 pp, medium confidence
Claude Sonnet 4.513.6%35.3%estimated ± 0.7 pp, low confidence
Claude Sonnet 4.5 Thinking13.6%35.3%estimated ± 0.7 pp, low confidence
Claude Sonnet 4.658.3%35.3%estimated ± 0.7 pp, low confidence
DeepSeek V4 Flash 073161.4%35.3%estimated ± 0.7 pp, low confidence
DeepSeek V4 Pro 081361.3%35.3%estimated ± 0.7 pp, low confidence
dots3-note Preview81.4%35.9%estimated ± 0.7 pp, low confidence
Gemini 3.1 Pro77.1%35.3%estimated ± 0.7 pp, low confidence
Gemini 3.5 Flash72.1%35.3%estimated ± 0.7 pp, low confidence
Gemini 3.6 Flash60.4%35.3%estimated ± 0.7 pp, low confidence
Gemini 3 Pro31.1%35.3%estimated ± 0.7 pp, low confidence
Gemini 3 Pro Deep Think45.1%35.3%estimated ± 0.7 pp, low confidence
GPT-5.252.9%35.3%estimated ± 0.7 pp, low confidence
GPT-5.474.0%35.3%estimated ± 0.7 pp, low confidence
GPT-5.4 mini18.9%35.3%estimated ± 0.7 pp, low confidence
GPT-5.4 nano5.7%35.3%estimated ± 0.7 pp, low confidence
GPT-5.4 Pro83.3%46.4%estimated ± 0.7 pp, low confidence
GPT-5.585.0%54.1%estimated ± 0.7 pp, medium confidence
GPT-5.5 Pro84.2%52.5%estimated ± 0.7 pp, medium confidence
GPT-5.6 Luna59.5%35.3%estimated ± 0.7 pp, low confidence
GPT-6.1 Sol94.2%54.6%estimated ± 0.7 pp, low confidence
GPT-6 Astra95.0%54.6%estimated ± 0.7 pp, low confidence
GPT-6 Luna59.3%35.3%estimated ± 0.7 pp, low confidence
GPT-6 Sol89.6%54.6%estimated ± 0.7 pp, medium confidence
Grok 4.2053.3%35.3%estimated ± 0.7 pp, low confidence
Grok 4.552.6%35.3%estimated ± 0.7 pp, low confidence
Grok 4.667.1%35.3%estimated ± 0.7 pp, low confidence
Inkling-Small40.1%35.3%estimated ± 0.7 pp, low confidence
Kimi K360.4%35.3%estimated ± 0.7 pp, low confidence
Muse Spark42.5%35.3%estimated ± 0.7 pp, low confidence