benchgap
Calibration

AA-IFBench → IFBench

IFBench is estimated from AA-IFBench with a Hill curve fitted on 11 models measured on both: y = 0.5529 + (1.2000 − 0.5529)·x^6.00 / (0.87644^6.00 + x^6.00), R² = 0.71, cross-validated error 6.9 pp. It is used for 119 estimates.

Estimated modelAA-IFBenchIFBenchSource
Claude 3 Haiku36.1%55.6%estimated ± 6.9 pp, low confidence
Claude 4.1 Opus Thinking55.4%59.2%estimated ± 6.9 pp, medium confidence
Claude 4 Sonnet45.4%56.5%estimated ± 6.9 pp, medium confidence
Claude Fable 563.5%63.5%estimated ± 6.9 pp, medium confidence
Claude Opus 4.5 Thinking58.0%60.3%estimated ± 6.9 pp, medium confidence
Claude Opus 4.644.6%56.4%estimated ± 6.9 pp, medium confidence
Claude Opus 4.6 (Adaptive)53.1%58.3%estimated ± 6.9 pp, medium confidence
Claude Opus 4.743.6%56.3%estimated ± 6.9 pp, medium confidence
Claude Opus 4.7 (Adaptive)58.6%60.6%estimated ± 6.9 pp, medium confidence
Claude Opus 4.862.2%62.6%estimated ± 6.9 pp, medium confidence
Claude Sonnet 4.641.2%56.0%estimated ± 6.9 pp, low confidence
Command A+73.9%72.4%estimated ± 6.9 pp, medium confidence
DeepSeek-R139.6%55.8%estimated ± 6.9 pp, low confidence
DeepSeek R1 Distill Qwen 32B22.9%55.3%estimated ± 6.9 pp, low confidence
DeepSeek V334.8%55.5%estimated ± 6.9 pp, low confidence
DeepSeek V3 032441.0%56.0%estimated ± 6.9 pp, low confidence
DeepSeek V3.137.8%55.7%estimated ± 6.9 pp, low confidence
DeepSeek V3.1 (Reasoning)41.5%56.0%estimated ± 6.9 pp, low confidence
DeepSeek V3.249.0%57.2%estimated ± 6.9 pp, medium confidence
DeepSeek V4 Pro 081376.5%75.1%estimated ± 6.9 pp, medium confidence
Exaone 4.0 1.2B25.3%55.3%estimated ± 6.9 pp, low confidence
Exaone 4.0 32B33.5%55.5%estimated ± 6.9 pp, low confidence
Gemini 2.5 Flash39.0%55.8%estimated ± 6.9 pp, low confidence
Gemini 2.5 Pro48.7%57.1%estimated ± 6.9 pp, medium confidence
Gemini 3.1 Pro77.1%75.8%estimated ± 6.9 pp, medium confidence
Gemini 3 Flash55.1%59.1%estimated ± 6.9 pp, medium confidence
Gemini 3 Pro70.4%69.0%estimated ± 6.9 pp, medium confidence
Gemma 3 27B31.8%55.4%estimated ± 6.9 pp, low confidence
Gemma 4 12B73.5%72.0%estimated ± 6.9 pp, medium confidence
Gemma 4 26B A4B72.4%70.9%estimated ± 6.9 pp, medium confidence
Gemma 4 31B75.6%74.2%estimated ± 6.9 pp, medium confidence
Gemma 4 E2B38.0%55.7%estimated ± 6.9 pp, low confidence
Gemma 4 E4B44.2%56.3%estimated ± 6.9 pp, medium confidence
GLM-4.5-Air37.6%55.7%estimated ± 6.9 pp, low confidence
GLM-4.636.7%55.6%estimated ± 6.9 pp, low confidence
GLM-4.767.9%66.8%estimated ± 6.9 pp, medium confidence
GLM-572.3%70.8%estimated ± 6.9 pp, medium confidence
GLM-5.176.3%74.9%estimated ± 6.9 pp, medium confidence
GLM-5.273.3%71.8%estimated ± 6.9 pp, medium confidence
GLM-5-Turbo73.2%71.7%estimated ± 6.9 pp, medium confidence
GLM-5V-Turbo61.1%62.0%estimated ± 6.9 pp, medium confidence
GPT-4.143.0%56.2%estimated ± 6.9 pp, medium confidence
GPT-4.1 mini38.3%55.7%estimated ± 6.9 pp, low confidence
GPT-4.1 nano32.0%55.4%estimated ± 6.9 pp, low confidence
GPT-4o34.3%55.5%estimated ± 6.9 pp, low confidence
GPT-4o mini31.0%55.4%estimated ± 6.9 pp, low confidence
GPT-5.172.9%71.4%estimated ± 6.9 pp, medium confidence
GPT-5.1-Codex70.0%68.6%estimated ± 6.9 pp, medium confidence
GPT-5.1-Codex-Max70.0%68.6%estimated ± 6.9 pp, medium confidence
GPT-5.275.4%74.0%estimated ± 6.9 pp, medium confidence
GPT-5.2-Codex77.6%76.3%estimated ± 6.9 pp, medium confidence
GPT-5.3 Codex75.4%74.0%estimated ± 6.9 pp, medium confidence
GPT-5.473.9%72.4%estimated ± 6.9 pp, medium confidence
GPT-5.4 mini73.3%71.8%estimated ± 6.9 pp, medium confidence
GPT-5.4 nano75.9%74.5%estimated ± 6.9 pp, medium confidence
GPT-5.575.9%74.5%estimated ± 6.9 pp, medium confidence
GPT-5.6 Sol72.7%71.2%estimated ± 6.9 pp, medium confidence
GPT-5.6 Terra71.2%69.7%estimated ± 6.9 pp, medium confidence
GPT-5 (high)73.1%71.6%estimated ± 6.9 pp, medium confidence
GPT-5 (medium)70.6%69.2%estimated ± 6.9 pp, medium confidence
GPT-OSS 120B69.0%67.7%estimated ± 6.9 pp, medium confidence
GPT-OSS 20B65.1%64.6%estimated ± 6.9 pp, medium confidence
Granite-4.0-350M15.9%55.3%estimated ± 6.9 pp, low confidence
Granite-4.0-H-1B26.2%55.3%estimated ± 6.9 pp, low confidence
Granite-4.0-H-350M17.6%55.3%estimated ± 6.9 pp, low confidence
Grok 453.7%58.5%estimated ± 6.9 pp, medium confidence
Grok 4.1 Fast36.5%55.6%estimated ± 6.9 pp, low confidence
Grok 4.1 Fast (Reasoning)52.7%58.2%estimated ± 6.9 pp, medium confidence
Grok 4 Fast (Reasoning)50.5%57.6%estimated ± 6.9 pp, medium confidence
Grok Code Fast 141.4%56.0%estimated ± 6.9 pp, low confidence
K-Exaone64.7%64.3%estimated ± 6.9 pp, medium confidence
Kimi K2.676.0%74.6%estimated ± 6.9 pp, medium confidence
Kimi K241.5%56.0%estimated ± 6.9 pp, low confidence
Kimi K2.570.2%68.8%estimated ± 6.9 pp, medium confidence
Kimi K2.5 (Reasoning)70.2%68.8%estimated ± 6.9 pp, medium confidence
Kimi K2.7 Code63.1%63.2%estimated ± 6.9 pp, medium confidence
LFM2.5-VL-1.6B-Extract33.1%55.5%estimated ± 6.9 pp, low confidence
Llama 3.1 405B39.0%55.8%estimated ± 6.9 pp, low confidence
Llama 4 Maverick43.0%56.2%estimated ± 6.9 pp, medium confidence
Llama 4 Scout39.5%55.8%estimated ± 6.9 pp, low confidence
MiMo-V2.5-Pro79.9%78.9%estimated ± 6.9 pp, medium confidence
MiMo-V2-Flash39.9%55.9%estimated ± 6.9 pp, low confidence
MiMo-V2-Omni53.5%58.5%estimated ± 6.9 pp, medium confidence
MiMo-V2-Pro68.8%67.6%estimated ± 6.9 pp, medium confidence
MiniMax M2.775.7%74.3%estimated ± 6.9 pp, medium confidence
MiniMax M382.9%82.3%estimated ± 6.9 pp, low confidence
Mistral Large 231.2%55.4%estimated ± 6.9 pp, low confidence
Mistral Large 336.2%55.6%estimated ± 6.9 pp, low confidence
Mistral Medium 339.3%55.8%estimated ± 6.9 pp, low confidence
Mistral Medium 3.5 128B68.8%67.6%estimated ± 6.9 pp, medium confidence
Mistral Small 448.2%57.0%estimated ± 6.9 pp, medium confidence
Mistral Small 4 (Reasoning)48.2%57.0%estimated ± 6.9 pp, medium confidence
Muse Spark75.9%74.5%estimated ± 6.9 pp, medium confidence
Nemotron 3 Nano 30B71.1%69.6%estimated ± 6.9 pp, medium confidence
Nemotron 3 Super 100B71.5%70.0%estimated ± 6.9 pp, medium confidence
Nemotron Ultra 253B38.2%55.7%estimated ± 6.9 pp, low confidence
North Mini Code57.6%60.1%estimated ± 6.9 pp, medium confidence
Nova Pro38.1%55.7%estimated ± 6.9 pp, low confidence
o170.3%68.9%estimated ± 6.9 pp, medium confidence
o371.4%69.9%estimated ± 6.9 pp, medium confidence
Phi-423.5%55.3%estimated ± 6.9 pp, low confidence
Qwen3.5-122B-A10B75.7%74.3%estimated ± 6.9 pp, medium confidence
Qwen3.5-27B75.6%74.2%estimated ± 6.9 pp, medium confidence
Qwen3.5-35B-A3B72.5%71.0%estimated ± 6.9 pp, medium confidence
Qwen3.5 397B51.6%57.9%estimated ± 6.9 pp, medium confidence
Qwen3.5 397B (Reasoning)51.6%57.9%estimated ± 6.9 pp, medium confidence
Qwen3.6-27B67.6%66.5%estimated ± 6.9 pp, medium confidence
Qwen3.6-35B-A3B64.4%64.1%estimated ± 6.9 pp, medium confidence
Qwen 3.6 Max (preview)76.6%75.2%estimated ± 6.9 pp, medium confidence
Qwen3 Max44.1%56.3%estimated ± 6.9 pp, medium confidence
Qwen3-Omni-30B-A3B-Instruct31.2%55.4%estimated ± 6.9 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking43.4%56.2%estimated ± 6.9 pp, medium confidence
Sarvam 105B34.4%55.5%estimated ± 6.9 pp, low confidence
Sarvam 30B26.5%55.3%estimated ± 6.9 pp, low confidence
Solar Pro 233.7%55.5%estimated ± 6.9 pp, low confidence
Step 3.7 Flash67.3%66.3%estimated ± 6.9 pp, medium confidence
Trinity-Large-Preview56.3%59.5%estimated ± 6.9 pp, medium confidence
Trinity-Large-Thinking56.3%59.5%estimated ± 6.9 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B47.1%56.8%estimated ± 6.9 pp, medium confidence