benchgap
Calibration

IFEval → AA-IFBench

AA-IFBench is estimated from IFEval with a inverse Michaelis–Menten curve fitted on 16 models measured on both: y = 0.30789·(x − 0.5994) / (0.5994 + 0.4832 − x), R² = 0.83, cross-validated error 8.2 pp. It is used for 17 estimates.

Estimated modelIFEvalAA-IFBenchSource
Agents-A194.8%79.9%estimated ± 8.2 pp, medium confidence
Agents-A1-4B94.8%79.8%estimated ± 8.2 pp, medium confidence
Celeris-180.8%23.4%estimated ± 8.2 pp, low confidence
dots3-note Preview93.9%72.8%estimated ± 8.2 pp, medium confidence
K-EXAONE 2.092.4%63.0%estimated ± 8.2 pp, medium confidence
Kanana-2 1.3B Instruct77.6%17.8%estimated ± 8.2 pp, low confidence
Kanana-2 3B Instruct81.0%23.7%estimated ± 8.2 pp, low confidence
LFM2.5-230M71.7%9.9%estimated ± 8.2 pp, low confidence
LFM2.5-VL-3B82.3%26.5%estimated ± 8.2 pp, low confidence
LFM2.5-VL-450M61.2%0.8%estimated ± 8.2 pp, low confidence
Mellum2-12B-A2.5B-Instruct75.8%15.0%estimated ± 8.2 pp, low confidence
Mellum2-12B-A2.5B-Thinking76.5%16.1%estimated ± 8.2 pp, low confidence
MiniCPM5-1B80.4%22.6%estimated ± 8.2 pp, low confidence
MiniCPM5-2B86.7%38.2%estimated ± 8.2 pp, medium confidence
o3-mini93.9%72.8%estimated ± 8.2 pp, medium confidence
Ternary Bonsai 2 27B91.3%57.0%estimated ± 8.2 pp, medium confidence
ZAYA1-8B85.6%34.8%estimated ± 8.2 pp, medium confidence