Calibration
τ²-bench results → BFCL v4
BFCL v4 is estimated from τ²-bench results with a Michaelis–Menten curve fitted on 6 models measured on both: y = 2.0000·x / (1.85549 + x), R² = 0.55, cross-validated error 8.7 pp. It is used for 61 estimates.
| Estimated model | τ²-bench results | BFCL v4 | Source |
|---|---|---|---|
| GLM-5-Turbo | 98.5% | 69.4% | estimated ± 8.7 pp, low confidence |
| GLM-5V-Turbo | 98.5% | 69.4% | estimated ± 8.7 pp, low confidence |
| Qwen 3.6 Max (preview) | 95.9% | 68.1% | estimated ± 8.7 pp, low confidence |
| MiMo-V2-Pro | 95.0% | 67.7% | estimated ± 8.7 pp, low confidence |
| Qwen3.5-27B | 93.9% | 67.2% | estimated ± 8.7 pp, low confidence |
| Grok 4.1 Fast (Reasoning) | 93.3% | 66.9% | estimated ± 8.7 pp, low confidence |
| Claude Opus 4.6 (Adaptive) | 92.1% | 66.3% | estimated ± 8.7 pp, low confidence |
| GPT-5.2-Codex | 92.1% | 66.3% | estimated ± 8.7 pp, low confidence |
| MiMo-V2-Omni | 91.2% | 65.9% | estimated ± 8.7 pp, low confidence |
| Claude Opus 4.5 Thinking | 89.5% | 65.1% | estimated ± 8.7 pp, low confidence |
| Qwen3.5-35B-A3B | 89.2% | 64.9% | estimated ± 8.7 pp, low confidence |
| Gemini 3 Pro | 87.1% | 63.9% | estimated ± 8.7 pp, low confidence |
| GPT-5 (medium) | 86.5% | 63.6% | estimated ± 8.7 pp, low confidence |
| GPT-5.3 Codex | 86.0% | 63.3% | estimated ± 8.7 pp, low confidence |
| Claude Opus 4.6 | 84.8% | 62.7% | estimated ± 8.7 pp, low confidence |
| GPT-5.2 | 84.8% | 62.7% | estimated ± 8.7 pp, low confidence |
| Qwen3.5 397B (Reasoning) | 83.9% | 62.3% | estimated ± 8.7 pp, low confidence |
| GPT-5.1-Codex | 83.0% | 61.8% | estimated ± 8.7 pp, low confidence |
| GPT-5.1-Codex-Max | 83.0% | 61.8% | estimated ± 8.7 pp, low confidence |
| o3 | 80.7% | 60.6% | estimated ± 8.7 pp, low confidence |
| Ternary Bonsai 2 27B | 80.2% | 60.4% | estimated ± 8.7 pp, low confidence |
| Claude Sonnet 4.6 | 79.5% | 60.0% | estimated ± 8.7 pp, low confidence |
| DeepSeek V3.2 | 78.9% | 59.7% | estimated ± 8.7 pp, low confidence |
| Agents-A1-4B | 78.2% | 59.3% | estimated ± 8.7 pp, low confidence |
| GLM-4.6 | 76.9% | 58.6% | estimated ± 8.7 pp, low confidence |
| Grok Code Fast 1 | 75.7% | 58.0% | estimated ± 8.7 pp, low confidence |
| Grok 4 | 74.9% | 57.5% | estimated ± 8.7 pp, low confidence |
| Qwen3 Max | 74.3% | 57.2% | estimated ± 8.7 pp, low confidence |
| Claude Opus 4.7 | 74.0% | 57.0% | estimated ± 8.7 pp, low confidence |
| Claude 4.1 Opus Thinking | 71.4% | 55.6% | estimated ± 8.7 pp, low confidence |
| Grok 4 Fast (Reasoning) | 65.8% | 52.4% | estimated ± 8.7 pp, low confidence |
| Grok 4.1 Fast | 63.7% | 51.1% | estimated ± 8.7 pp, low confidence |
| o1 | 62.6% | 50.5% | estimated ± 8.7 pp, low confidence |
| Kimi K2 | 61.1% | 49.5% | estimated ± 8.7 pp, low confidence |
| Claude 4 Sonnet | 52.3% | 44.0% | estimated ± 8.7 pp, low confidence |
| GPT-4.1 | 47.1% | 40.5% | estimated ± 8.7 pp, low confidence |
| Sarvam 105B | 46.8% | 40.3% | estimated ± 8.7 pp, low confidence |
| GLM-4.5-Air | 46.5% | 40.1% | estimated ± 8.7 pp, low confidence |
| Gemini 3 Flash | 43.3% | 37.8% | estimated ± 8.7 pp, low confidence |
| DeepSeek V3.1 (Reasoning) | 37.4% | 33.6% | estimated ± 8.7 pp, low confidence |
| DeepSeek-R1 | 36.5% | 32.9% | estimated ± 8.7 pp, low confidence |
| DeepSeek V3.1 | 34.8% | 31.6% | estimated ± 8.7 pp, low confidence |
| Sarvam 30B | 34.5% | 31.4% | estimated ± 8.7 pp, low confidence |
| Solar Pro 2 | 31.9% | 29.3% | estimated ± 8.7 pp, low confidence |
| Mistral Large 2 | 30.7% | 28.4% | estimated ± 8.7 pp, low confidence |
| o3-mini | 28.7% | 26.8% | estimated ± 8.7 pp, low confidence |
| Mistral Medium 3 | 24.3% | 23.2% | estimated ± 8.7 pp, low confidence |
| Qwen3-Omni-30B-A3B-Thinking | 21.3% | 20.6% | estimated ± 8.7 pp, low confidence |
| Claude 3 Haiku | 21.1% | 20.4% | estimated ± 8.7 pp, low confidence |
| Exaone 4.0 1.2B | 20.5% | 19.9% | estimated ± 8.7 pp, low confidence |
| Granite-4.0-H-1B | 19.6% | 19.1% | estimated ± 8.7 pp, low confidence |
| Llama 3.1 405B | 19.0% | 18.6% | estimated ± 8.7 pp, low confidence |
| Qwen3-Omni-30B-A3B-Instruct | 16.4% | 16.2% | estimated ± 8.7 pp, low confidence |
| Gemini 2.5 Flash | 14.9% | 14.9% | estimated ± 8.7 pp, low confidence |
| Granite-4.0-H-350M | 14.6% | 14.6% | estimated ± 8.7 pp, low confidence |
| Nova Pro | 14.0% | 14.0% | estimated ± 8.7 pp, low confidence |
| Granite-4.0-350M | 13.2% | 13.3% | estimated ± 8.7 pp, low confidence |
| Nemotron Ultra 253B | 11.4% | 11.6% | estimated ± 8.7 pp, low confidence |
| LFM2.5-VL-1.6B-Extract | 8.5% | 8.8% | estimated ± 8.7 pp, low confidence |
| Exaone 4.0 32B | 4.1% | 4.3% | estimated ± 8.7 pp, low confidence |
| Phi-4 | 0.0% | 0.0% | estimated ± 8.7 pp, low confidence |