benchgap
Calibration

AA-IFBench → IFEval

IFEval is estimated from AA-IFBench with a Hill curve fitted on 16 models measured on both: y = 0.0000 + (0.9472 − 0.0000)·x^2.88 / (0.15867^2.88 + x^2.88), R² = 0.90, cross-validated error 1.3 pp. It is used for 114 estimates.

Estimated modelAA-IFBenchIFEvalSource
Claude 3 Haiku36.1%86.6%estimated ± 1.3 pp, high confidence
Claude 4.1 Opus Thinking55.4%92.2%estimated ± 1.3 pp, high confidence
Claude 4 Sonnet45.4%90.3%estimated ± 1.3 pp, high confidence
Claude Fable 563.5%93.0%estimated ± 1.3 pp, high confidence
Claude Opus 4.5 Thinking58.0%92.5%estimated ± 1.3 pp, high confidence
Claude Opus 4.644.6%90.1%estimated ± 1.3 pp, high confidence
Claude Opus 4.6 (Adaptive)53.1%91.9%estimated ± 1.3 pp, high confidence
Claude Opus 4.743.6%89.8%estimated ± 1.3 pp, high confidence
Claude Opus 4.7 (Adaptive)58.6%92.6%estimated ± 1.3 pp, high confidence
Claude Opus 4.862.2%92.9%estimated ± 1.3 pp, high confidence
Claude Sonnet 4.641.2%89.0%estimated ± 1.3 pp, high confidence
Command A+73.9%93.6%estimated ± 1.3 pp, high confidence
DeepSeek-R139.6%88.4%estimated ± 1.3 pp, high confidence
DeepSeek R1 Distill Qwen 32B22.9%70.3%estimated ± 1.3 pp, medium confidence
DeepSeek V3 032441.0%88.9%estimated ± 1.3 pp, high confidence
DeepSeek V3.137.8%87.5%estimated ± 1.3 pp, high confidence
DeepSeek V3.1 (Reasoning)41.5%89.1%estimated ± 1.3 pp, high confidence
DeepSeek V3.249.0%91.2%estimated ± 1.3 pp, high confidence
DeepSeek V4 Pro 081376.5%93.7%estimated ± 1.3 pp, high confidence
Exaone 4.0 1.2B25.3%75.1%estimated ± 1.3 pp, medium confidence
Exaone 4.0 32B33.5%84.8%estimated ± 1.3 pp, high confidence
Gemini 2.5 Flash39.0%88.1%estimated ± 1.3 pp, high confidence
Gemini 2.5 Pro48.7%91.1%estimated ± 1.3 pp, high confidence
Gemini 3.1 Pro77.1%93.7%estimated ± 1.3 pp, high confidence
Gemini 3.5 Flash76.3%93.7%estimated ± 1.3 pp, high confidence
Gemini 3 Flash55.1%92.2%estimated ± 1.3 pp, high confidence
Gemini 3 Pro70.4%93.4%estimated ± 1.3 pp, high confidence
Gemma 3 27B31.8%83.4%estimated ± 1.3 pp, medium confidence
Gemma 4 12B73.5%93.6%estimated ± 1.3 pp, high confidence
Gemma 4 26B A4B72.4%93.5%estimated ± 1.3 pp, high confidence
Gemma 4 31B75.6%93.7%estimated ± 1.3 pp, high confidence
Gemma 4 E2B38.0%87.6%estimated ± 1.3 pp, high confidence
Gemma 4 E4B44.2%90.0%estimated ± 1.3 pp, high confidence
GLM-4.5-Air37.6%87.4%estimated ± 1.3 pp, high confidence
GLM-4.636.7%86.9%estimated ± 1.3 pp, high confidence
GLM-4.767.9%93.3%estimated ± 1.3 pp, high confidence
GLM-5.176.3%93.7%estimated ± 1.3 pp, high confidence
GLM-5.273.3%93.6%estimated ± 1.3 pp, high confidence
GLM-5-Turbo73.2%93.6%estimated ± 1.3 pp, high confidence
GLM-5V-Turbo61.1%92.8%estimated ± 1.3 pp, high confidence
GPT-4o34.3%85.4%estimated ± 1.3 pp, high confidence
GPT-4o mini31.0%82.7%estimated ± 1.3 pp, medium confidence
GPT-5.172.9%93.6%estimated ± 1.3 pp, high confidence
GPT-5.1-Codex70.0%93.4%estimated ± 1.3 pp, high confidence
GPT-5.1-Codex-Max70.0%93.4%estimated ± 1.3 pp, high confidence
GPT-5.275.4%93.7%estimated ± 1.3 pp, high confidence
GPT-5.2-Codex77.6%93.7%estimated ± 1.3 pp, high confidence
GPT-5.3 Codex75.4%93.7%estimated ± 1.3 pp, high confidence
GPT-5.473.9%93.6%estimated ± 1.3 pp, high confidence
GPT-5.4 mini73.3%93.6%estimated ± 1.3 pp, high confidence
GPT-5.4 nano75.9%93.7%estimated ± 1.3 pp, high confidence
GPT-5.575.9%93.7%estimated ± 1.3 pp, high confidence
GPT-5.6 Sol72.7%93.5%estimated ± 1.3 pp, high confidence
GPT-5.6 Terra71.2%93.5%estimated ± 1.3 pp, high confidence
GPT-5 (high)73.1%93.6%estimated ± 1.3 pp, high confidence
GPT-5 (medium)70.6%93.4%estimated ± 1.3 pp, high confidence
GPT-OSS 120B69.0%93.4%estimated ± 1.3 pp, high confidence
GPT-OSS 20B65.1%93.1%estimated ± 1.3 pp, high confidence
Granite-4.0-350M15.9%47.5%estimated ± 1.3 pp, medium confidence
Granite-4.0-H-1B26.2%76.6%estimated ± 1.3 pp, medium confidence
Granite-4.0-H-350M17.6%54.4%estimated ± 1.3 pp, medium confidence
Grok 453.7%92.0%estimated ± 1.3 pp, high confidence
Grok 4.1 Fast36.5%86.8%estimated ± 1.3 pp, high confidence
Grok 4.1 Fast (Reasoning)52.7%91.8%estimated ± 1.3 pp, high confidence
Grok 4.381.3%93.9%estimated ± 1.3 pp, medium confidence
Grok 4 Fast (Reasoning)50.5%91.4%estimated ± 1.3 pp, high confidence
Grok Code Fast 141.4%89.1%estimated ± 1.3 pp, high confidence
K-Exaone64.7%93.1%estimated ± 1.3 pp, high confidence
Kimi K2.676.0%93.7%estimated ± 1.3 pp, high confidence
Kimi K241.5%89.1%estimated ± 1.3 pp, high confidence
Kimi K2.5 (Reasoning)70.2%93.4%estimated ± 1.3 pp, high confidence
Kimi K2.7 Code63.1%93.0%estimated ± 1.3 pp, high confidence
LFM2.5-VL-1.6B-Extract33.1%84.5%estimated ± 1.3 pp, high confidence
Ling 2.6 Flash57.4%92.4%estimated ± 1.3 pp, high confidence
Llama 3.1 405B39.0%88.1%estimated ± 1.3 pp, high confidence
Llama 4 Maverick43.0%89.6%estimated ± 1.3 pp, high confidence
Llama 4 Scout39.5%88.3%estimated ± 1.3 pp, high confidence
MiMo-V2.5-Pro79.9%93.8%estimated ± 1.3 pp, high confidence
MiMo-V2-Flash39.9%88.5%estimated ± 1.3 pp, high confidence
MiMo-V2-Omni53.5%91.9%estimated ± 1.3 pp, high confidence
MiMo-V2-Pro68.8%93.3%estimated ± 1.3 pp, high confidence
MiniMax M2.775.7%93.7%estimated ± 1.3 pp, high confidence
MiniMax M382.9%93.9%estimated ± 1.3 pp, medium confidence
Mistral Large 231.2%82.9%estimated ± 1.3 pp, medium confidence
Mistral Large 336.2%86.6%estimated ± 1.3 pp, high confidence
Mistral Medium 339.3%88.2%estimated ± 1.3 pp, high confidence
Mistral Medium 3.5 128B68.8%93.3%estimated ± 1.3 pp, high confidence
Mistral Small 448.2%91.0%estimated ± 1.3 pp, high confidence
Mistral Small 4 (Reasoning)48.2%91.0%estimated ± 1.3 pp, high confidence
Muse Spark75.9%93.7%estimated ± 1.3 pp, high confidence
Nemotron 3 Nano 30B71.1%93.5%estimated ± 1.3 pp, high confidence
Nemotron 3 Nano Omni 30B A3B63.2%93.0%estimated ± 1.3 pp, high confidence
Nemotron 3 Super 100B71.5%93.5%estimated ± 1.3 pp, high confidence
Nemotron 3 Ultra81.4%93.9%estimated ± 1.3 pp, medium confidence
Nemotron Ultra 253B38.2%87.7%estimated ± 1.3 pp, high confidence
North Mini Code57.6%92.5%estimated ± 1.3 pp, high confidence
Nova Pro38.1%87.7%estimated ± 1.3 pp, high confidence
o371.4%93.5%estimated ± 1.3 pp, high confidence
Phi-423.5%71.6%estimated ± 1.3 pp, medium confidence
Qwen3.5 397B (Reasoning)51.6%91.6%estimated ± 1.3 pp, high confidence
Qwen3.6-27B67.6%93.3%estimated ± 1.3 pp, high confidence
Qwen3.6-35B-A3B64.4%93.1%estimated ± 1.3 pp, high confidence
Qwen 3.6 Max (preview)76.6%93.7%estimated ± 1.3 pp, high confidence
Qwen3 Max44.1%90.0%estimated ± 1.3 pp, high confidence
Qwen3-Omni-30B-A3B-Instruct31.2%82.9%estimated ± 1.3 pp, medium confidence
Qwen3-Omni-30B-A3B-Thinking43.4%89.8%estimated ± 1.3 pp, high confidence
Sarvam 105B34.4%85.5%estimated ± 1.3 pp, high confidence
Sarvam 30B26.5%77.1%estimated ± 1.3 pp, medium confidence
Solar Pro 233.7%85.0%estimated ± 1.3 pp, high confidence
Solar Pro 371.2%93.5%estimated ± 1.3 pp, high confidence
Step 3.7 Flash67.3%93.3%estimated ± 1.3 pp, high confidence
Trinity-Large-Preview56.3%92.3%estimated ± 1.3 pp, high confidence
Trinity-Large-Thinking56.3%92.3%estimated ± 1.3 pp, high confidence
Ultravox v0.6 Llama 3.3 70B47.1%90.7%estimated ± 1.3 pp, high confidence