benchgap
Calibration

τ²-bench results → BFCL v4

BFCL v4 is estimated from τ²-bench results with a Michaelis–Menten curve fitted on 6 models measured on both: y = 2.0000·x / (1.85549 + x), R² = 0.55, cross-validated error 8.7 pp. It is used for 61 estimates.

Estimated modelτ²-bench resultsBFCL v4Source
GLM-5-Turbo98.5%69.4%estimated ± 8.7 pp, low confidence
GLM-5V-Turbo98.5%69.4%estimated ± 8.7 pp, low confidence
Qwen 3.6 Max (preview)95.9%68.1%estimated ± 8.7 pp, low confidence
MiMo-V2-Pro95.0%67.7%estimated ± 8.7 pp, low confidence
Qwen3.5-27B93.9%67.2%estimated ± 8.7 pp, low confidence
Grok 4.1 Fast (Reasoning)93.3%66.9%estimated ± 8.7 pp, low confidence
Claude Opus 4.6 (Adaptive)92.1%66.3%estimated ± 8.7 pp, low confidence
GPT-5.2-Codex92.1%66.3%estimated ± 8.7 pp, low confidence
MiMo-V2-Omni91.2%65.9%estimated ± 8.7 pp, low confidence
Claude Opus 4.5 Thinking89.5%65.1%estimated ± 8.7 pp, low confidence
Qwen3.5-35B-A3B89.2%64.9%estimated ± 8.7 pp, low confidence
Gemini 3 Pro87.1%63.9%estimated ± 8.7 pp, low confidence
GPT-5 (medium)86.5%63.6%estimated ± 8.7 pp, low confidence
GPT-5.3 Codex86.0%63.3%estimated ± 8.7 pp, low confidence
Claude Opus 4.684.8%62.7%estimated ± 8.7 pp, low confidence
GPT-5.284.8%62.7%estimated ± 8.7 pp, low confidence
Qwen3.5 397B (Reasoning)83.9%62.3%estimated ± 8.7 pp, low confidence
GPT-5.1-Codex83.0%61.8%estimated ± 8.7 pp, low confidence
GPT-5.1-Codex-Max83.0%61.8%estimated ± 8.7 pp, low confidence
o380.7%60.6%estimated ± 8.7 pp, low confidence
Ternary Bonsai 2 27B80.2%60.4%estimated ± 8.7 pp, low confidence
Claude Sonnet 4.679.5%60.0%estimated ± 8.7 pp, low confidence
DeepSeek V3.278.9%59.7%estimated ± 8.7 pp, low confidence
Agents-A1-4B78.2%59.3%estimated ± 8.7 pp, low confidence
GLM-4.676.9%58.6%estimated ± 8.7 pp, low confidence
Grok Code Fast 175.7%58.0%estimated ± 8.7 pp, low confidence
Grok 474.9%57.5%estimated ± 8.7 pp, low confidence
Qwen3 Max74.3%57.2%estimated ± 8.7 pp, low confidence
Claude Opus 4.774.0%57.0%estimated ± 8.7 pp, low confidence
Claude 4.1 Opus Thinking71.4%55.6%estimated ± 8.7 pp, low confidence
Grok 4 Fast (Reasoning)65.8%52.4%estimated ± 8.7 pp, low confidence
Grok 4.1 Fast63.7%51.1%estimated ± 8.7 pp, low confidence
o162.6%50.5%estimated ± 8.7 pp, low confidence
Kimi K261.1%49.5%estimated ± 8.7 pp, low confidence
Claude 4 Sonnet52.3%44.0%estimated ± 8.7 pp, low confidence
GPT-4.147.1%40.5%estimated ± 8.7 pp, low confidence
Sarvam 105B46.8%40.3%estimated ± 8.7 pp, low confidence
GLM-4.5-Air46.5%40.1%estimated ± 8.7 pp, low confidence
Gemini 3 Flash43.3%37.8%estimated ± 8.7 pp, low confidence
DeepSeek V3.1 (Reasoning)37.4%33.6%estimated ± 8.7 pp, low confidence
DeepSeek-R136.5%32.9%estimated ± 8.7 pp, low confidence
DeepSeek V3.134.8%31.6%estimated ± 8.7 pp, low confidence
Sarvam 30B34.5%31.4%estimated ± 8.7 pp, low confidence
Solar Pro 231.9%29.3%estimated ± 8.7 pp, low confidence
Mistral Large 230.7%28.4%estimated ± 8.7 pp, low confidence
o3-mini28.7%26.8%estimated ± 8.7 pp, low confidence
Mistral Medium 324.3%23.2%estimated ± 8.7 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking21.3%20.6%estimated ± 8.7 pp, low confidence
Claude 3 Haiku21.1%20.4%estimated ± 8.7 pp, low confidence
Exaone 4.0 1.2B20.5%19.9%estimated ± 8.7 pp, low confidence
Granite-4.0-H-1B19.6%19.1%estimated ± 8.7 pp, low confidence
Llama 3.1 405B19.0%18.6%estimated ± 8.7 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct16.4%16.2%estimated ± 8.7 pp, low confidence
Gemini 2.5 Flash14.9%14.9%estimated ± 8.7 pp, low confidence
Granite-4.0-H-350M14.6%14.6%estimated ± 8.7 pp, low confidence
Nova Pro14.0%14.0%estimated ± 8.7 pp, low confidence
Granite-4.0-350M13.2%13.3%estimated ± 8.7 pp, low confidence
Nemotron Ultra 253B11.4%11.6%estimated ± 8.7 pp, low confidence
LFM2.5-VL-1.6B-Extract8.5%8.8%estimated ± 8.7 pp, low confidence
Exaone 4.0 32B4.1%4.3%estimated ± 8.7 pp, low confidence
Phi-40.0%0.0%estimated ± 8.7 pp, low confidence