benchgap
Calibration

AA-HLE → CritPt

CritPt is estimated from AA-HLE with a offset logistic curve fitted on 186 models measured on both: y = 0.0019 + (0.3319 − 0.0019) / (1 + exp(−15.66·(x − 0.4217))), R² = 0.91, cross-validated error 2.3 pp. It is used for 9 estimates.

Estimated modelAA-HLECritPtSource
Claude 3 Opus2.8%0.3%estimated ± 2.3 pp, high confidence
DeepSeek R1 Distill Qwen 32B4.6%0.3%estimated ± 2.3 pp, high confidence
Gemini 1.0 Pro4.2%0.3%estimated ± 2.3 pp, high confidence
Gemini 1.5 Pro4.6%0.3%estimated ± 2.3 pp, high confidence
GPT-4 Turbo3.1%0.3%estimated ± 2.3 pp, high confidence
GPT-4o mini4.2%0.3%estimated ± 2.3 pp, high confidence
o3-mini7.9%0.3%estimated ± 2.3 pp, high confidence
Phi-4 Multimodal Instruct5.0%0.3%estimated ± 2.3 pp, high confidence
Qwen2.5 Coder 32B Instruct3.5%0.3%estimated ± 2.3 pp, high confidence