Calibration
IFBench → AA-IFBench
AA-IFBench is estimated from IFBench with a offset logistic curve fitted on 11 models measured on both: y = 0.5678 + (0.8034 − 0.5678) / (1 + exp(−128.63·(x − 0.7493))), R² = 0.75, cross-validated error 8.5 pp. It is used for 21 estimates.
| Estimated model | IFBench | AA-IFBench | Source |
|---|---|---|---|
| A.X K2 | 75.9% | 75.1% | estimated ± 8.5 pp, medium confidence |
| Granite 4.2 30B | 77.2% | 79.1% | estimated ± 8.5 pp, medium confidence |
| Granite 4.2 3B | 74.3% | 64.2% | estimated ± 8.5 pp, medium confidence |
| Granite 4.2 8B | 79.3% | 80.3% | estimated ± 8.5 pp, medium confidence |
| Hy3 Preview | 63.1% | 56.8% | estimated ± 8.5 pp, medium confidence |
| Inkling | 79.8% | 80.3% | estimated ± 8.5 pp, medium confidence |
| Inkling-Small | 82.2% | 80.3% | estimated ± 8.5 pp, low confidence |
| LFM2.5-2.6B | 59.2% | 56.8% | estimated ± 8.5 pp, medium confidence |
| Ling 3.0 Flash | 74.5% | 65.4% | estimated ± 8.5 pp, medium confidence |
| Ling 3.0 Flash FP8 | 73.4% | 59.7% | estimated ± 8.5 pp, medium confidence |
| LLaDA2.2-mini | 24.9% | 56.8% | estimated ± 8.5 pp, low confidence |
| MAI-Thinking-1 | 85.0% | 80.3% | estimated ± 8.5 pp, low confidence |
| Mercury 2.5 | 77.0% | 78.8% | estimated ± 8.5 pp, medium confidence |
| Muse Glimmer 30B | 77.0% | 78.8% | estimated ± 8.5 pp, medium confidence |
| Nemotron 3.5 Lightning 30B A3B NVFP4 | 72.9% | 58.4% | estimated ± 8.5 pp, medium confidence |
| Qwen3.8-27B | 79.5% | 80.3% | estimated ± 8.5 pp, medium confidence |
| Qwen3.8-Flash-Next | 81.3% | 80.3% | estimated ± 8.5 pp, medium confidence |
| Qwen3.8 Max | 82.8% | 80.3% | estimated ± 8.5 pp, low confidence |
| Qwen3.8-Omni-Flash | 81.5% | 80.3% | estimated ± 8.5 pp, medium confidence |
| Beam | 79.7% | 80.3% | estimated ± 8.5 pp, medium confidence |
| Solar Open 2 | 80.0% | 80.3% | estimated ± 8.5 pp, medium confidence |