benchgap
Calibration

ARC-AGI-1 → LABBench2

LABBench2 is estimated from ARC-AGI-1 with a Hill curve fitted on 5 models measured on both: y = 0.0639 + (1.2000 − 0.0639)·x^6.00 / (0.85725^6.00 + x^6.00), R² = 0.77, cross-validated error 1.3 pp. It is used for 30 estimates.

Estimated modelARC-AGI-1LABBench2Source
Claude Fable 598.5%85.6%estimated ± 1.3 pp, medium confidence
Claude Fable 5.197.5%84.1%estimated ± 1.3 pp, medium confidence
Claude Opus 4.5 Thinking80.0%51.6%estimated ± 1.3 pp, low confidence
Claude Opus 4.6 (Adaptive)93.0%76.8%estimated ± 1.3 pp, low confidence
Claude Opus 4.892.5%75.9%estimated ± 1.3 pp, low confidence
Claude Opus 5.597.5%84.1%estimated ± 1.3 pp, medium confidence
Claude Sonnet 4.5 Thinking63.7%22.7%estimated ± 1.3 pp, low confidence
Claude Sonnet 4.686.0%63.7%estimated ± 1.3 pp, low confidence
DeepSeek V4 Flash 073189.0%69.6%estimated ± 1.3 pp, low confidence
DeepSeek V4 Pro 081390.0%71.4%estimated ± 1.3 pp, low confidence
Gemini 3.5 Flash92.5%75.9%estimated ± 1.3 pp, low confidence
Gemini 3.6 Flash91.2%73.6%estimated ± 1.3 pp, low confidence
Gemini 3 Pro75.0%41.6%estimated ± 1.3 pp, low confidence
GPT-5.172.8%37.4%estimated ± 1.3 pp, low confidence
GPT-5.286.2%64.1%estimated ± 1.3 pp, low confidence
GPT-5.493.7%78.0%estimated ± 1.3 pp, low confidence
GPT-5.4 mini63.7%22.8%estimated ± 1.3 pp, low confidence
GPT-5.4 nano51.5%11.5%estimated ± 1.3 pp, low confidence
GPT-5.4 Pro94.5%79.3%estimated ± 1.3 pp, low confidence
GPT-5.595.0%80.2%estimated ± 1.3 pp, low confidence
GPT-5.5 Pro95.0%80.2%estimated ± 1.3 pp, low confidence
GPT-5.6 Luna88.0%67.6%estimated ± 1.3 pp, low confidence
GPT-6.1 Sol96.5%82.6%estimated ± 1.3 pp, medium confidence
GPT-6 Astra98.5%85.6%estimated ± 1.3 pp, medium confidence
GPT-6 Luna86.7%65.1%estimated ± 1.3 pp, low confidence
GPT-6 Sol95.5%81.0%estimated ± 1.3 pp, medium confidence
Grok 4.585.7%63.1%estimated ± 1.3 pp, low confidence
Grok 4.687.0%65.7%estimated ± 1.3 pp, low confidence
Inkling-Small84.0%59.7%estimated ± 1.3 pp, low confidence
Kimi K394.5%79.3%estimated ± 1.3 pp, low confidence