benchgap
Calibration

IFBench → AA-IFBench

AA-IFBench is estimated from IFBench with a offset logistic curve fitted on 11 models measured on both: y = 0.5678 + (0.8034 − 0.5678) / (1 + exp(−128.63·(x − 0.7493))), R² = 0.75, cross-validated error 8.5 pp. It is used for 21 estimates.

Estimated modelIFBenchAA-IFBenchSource
A.X K275.9%75.1%estimated ± 8.5 pp, medium confidence
Granite 4.2 30B77.2%79.1%estimated ± 8.5 pp, medium confidence
Granite 4.2 3B74.3%64.2%estimated ± 8.5 pp, medium confidence
Granite 4.2 8B79.3%80.3%estimated ± 8.5 pp, medium confidence
Hy3 Preview63.1%56.8%estimated ± 8.5 pp, medium confidence
Inkling79.8%80.3%estimated ± 8.5 pp, medium confidence
Inkling-Small82.2%80.3%estimated ± 8.5 pp, low confidence
LFM2.5-2.6B59.2%56.8%estimated ± 8.5 pp, medium confidence
Ling 3.0 Flash74.5%65.4%estimated ± 8.5 pp, medium confidence
Ling 3.0 Flash FP873.4%59.7%estimated ± 8.5 pp, medium confidence
LLaDA2.2-mini24.9%56.8%estimated ± 8.5 pp, low confidence
MAI-Thinking-185.0%80.3%estimated ± 8.5 pp, low confidence
Mercury 2.577.0%78.8%estimated ± 8.5 pp, medium confidence
Muse Glimmer 30B77.0%78.8%estimated ± 8.5 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP472.9%58.4%estimated ± 8.5 pp, medium confidence
Qwen3.8-27B79.5%80.3%estimated ± 8.5 pp, medium confidence
Qwen3.8-Flash-Next81.3%80.3%estimated ± 8.5 pp, medium confidence
Qwen3.8 Max82.8%80.3%estimated ± 8.5 pp, low confidence
Qwen3.8-Omni-Flash81.5%80.3%estimated ± 8.5 pp, medium confidence
Beam79.7%80.3%estimated ± 8.5 pp, medium confidence
Solar Open 280.0%80.3%estimated ± 8.5 pp, medium confidence