benchgap
Calibration

Gert Labs → τ³-bench results

τ³-bench results is estimated from Gert Labs with a offset logistic curve fitted on 9 models measured on both: y = 1.0000 + (0.6891 − 1.0000) / (1 + exp(−200.00·(x − 0.3958))), R² = 0.90, cross-validated error 8.9 pp. It is used for 35 estimates.

Estimated modelGert Labsτ³-bench resultsSource
Claude 4 Sonnet39.7%83.2%estimated ± 8.9 pp, medium confidence
Claude Opus 4.661.9%68.9%estimated ± 8.9 pp, medium confidence
Claude Opus 4.765.6%68.9%estimated ± 8.9 pp, low confidence
Claude Sonnet 4.548.5%68.9%estimated ± 8.9 pp, medium confidence
Claude Sonnet 4.662.9%68.9%estimated ± 8.9 pp, medium confidence
DeepSeek V3.229.6%100.0%estimated ± 8.9 pp, low confidence
Gemini 2.5 Pro42.0%69.1%estimated ± 8.9 pp, medium confidence
Gemini 3.1 Flash-Lite38.5%97.0%estimated ± 8.9 pp, low confidence
Gemini 3.1 Pro56.9%68.9%estimated ± 8.9 pp, medium confidence
Gemini 3 Flash56.6%68.9%estimated ± 8.9 pp, medium confidence
Gemini 3 Pro63.2%68.9%estimated ± 8.9 pp, medium confidence
Gemma 4 31B35.3%100.0%estimated ± 8.9 pp, low confidence
GLM-4.740.0%78.9%estimated ± 8.9 pp, medium confidence
GLM-5V-Turbo30.8%100.0%estimated ± 8.9 pp, low confidence
GPT-4.125.7%100.0%estimated ± 8.9 pp, low confidence
GPT-5.141.2%70.0%estimated ± 8.9 pp, medium confidence
GPT-5.1-Codex49.7%68.9%estimated ± 8.9 pp, medium confidence
GPT-5.246.5%68.9%estimated ± 8.9 pp, medium confidence
GPT-5.2-Codex51.8%68.9%estimated ± 8.9 pp, medium confidence
GPT-5.3 Codex57.5%68.9%estimated ± 8.9 pp, medium confidence
GPT-OSS 120B29.6%100.0%estimated ± 8.9 pp, low confidence
Grok 442.3%69.0%estimated ± 8.9 pp, medium confidence
Grok 4.1 Fast47.3%68.9%estimated ± 8.9 pp, medium confidence
Grok 4.2038.4%97.5%estimated ± 8.9 pp, low confidence
Grok 4.343.9%68.9%estimated ± 8.9 pp, medium confidence
Grok Build 0.149.2%68.9%estimated ± 8.9 pp, medium confidence
Hy3 Preview36.9%99.9%estimated ± 8.9 pp, low confidence
Kimi K2.5 (Reasoning)32.6%100.0%estimated ± 8.9 pp, low confidence
MiMo-V2.546.9%68.9%estimated ± 8.9 pp, medium confidence
MiMo-V2-Pro36.7%99.9%estimated ± 8.9 pp, low confidence
Qwen3.5-27B39.4%87.1%estimated ± 8.9 pp, medium confidence
Qwen3.5-35B-A3B29.0%100.0%estimated ± 8.9 pp, low confidence
Qwen3.6-27B54.8%68.9%estimated ± 8.9 pp, medium confidence
Qwen3 Max43.7%68.9%estimated ± 8.9 pp, medium confidence
Trinity-Large-Thinking32.6%100.0%estimated ± 8.9 pp, low confidence