Calibration
Gert Labs → τ³-bench results
τ³-bench results is estimated from Gert Labs with a offset logistic curve fitted on 9 models measured on both: y = 1.0000 + (0.6891 − 1.0000) / (1 + exp(−200.00·(x − 0.3958))), R² = 0.90, cross-validated error 8.9 pp. It is used for 35 estimates.
| Estimated model | Gert Labs | τ³-bench results | Source |
|---|---|---|---|
| Claude 4 Sonnet | 39.7% | 83.2% | estimated ± 8.9 pp, medium confidence |
| Claude Opus 4.6 | 61.9% | 68.9% | estimated ± 8.9 pp, medium confidence |
| Claude Opus 4.7 | 65.6% | 68.9% | estimated ± 8.9 pp, low confidence |
| Claude Sonnet 4.5 | 48.5% | 68.9% | estimated ± 8.9 pp, medium confidence |
| Claude Sonnet 4.6 | 62.9% | 68.9% | estimated ± 8.9 pp, medium confidence |
| DeepSeek V3.2 | 29.6% | 100.0% | estimated ± 8.9 pp, low confidence |
| Gemini 2.5 Pro | 42.0% | 69.1% | estimated ± 8.9 pp, medium confidence |
| Gemini 3.1 Flash-Lite | 38.5% | 97.0% | estimated ± 8.9 pp, low confidence |
| Gemini 3.1 Pro | 56.9% | 68.9% | estimated ± 8.9 pp, medium confidence |
| Gemini 3 Flash | 56.6% | 68.9% | estimated ± 8.9 pp, medium confidence |
| Gemini 3 Pro | 63.2% | 68.9% | estimated ± 8.9 pp, medium confidence |
| Gemma 4 31B | 35.3% | 100.0% | estimated ± 8.9 pp, low confidence |
| GLM-4.7 | 40.0% | 78.9% | estimated ± 8.9 pp, medium confidence |
| GLM-5V-Turbo | 30.8% | 100.0% | estimated ± 8.9 pp, low confidence |
| GPT-4.1 | 25.7% | 100.0% | estimated ± 8.9 pp, low confidence |
| GPT-5.1 | 41.2% | 70.0% | estimated ± 8.9 pp, medium confidence |
| GPT-5.1-Codex | 49.7% | 68.9% | estimated ± 8.9 pp, medium confidence |
| GPT-5.2 | 46.5% | 68.9% | estimated ± 8.9 pp, medium confidence |
| GPT-5.2-Codex | 51.8% | 68.9% | estimated ± 8.9 pp, medium confidence |
| GPT-5.3 Codex | 57.5% | 68.9% | estimated ± 8.9 pp, medium confidence |
| GPT-OSS 120B | 29.6% | 100.0% | estimated ± 8.9 pp, low confidence |
| Grok 4 | 42.3% | 69.0% | estimated ± 8.9 pp, medium confidence |
| Grok 4.1 Fast | 47.3% | 68.9% | estimated ± 8.9 pp, medium confidence |
| Grok 4.20 | 38.4% | 97.5% | estimated ± 8.9 pp, low confidence |
| Grok 4.3 | 43.9% | 68.9% | estimated ± 8.9 pp, medium confidence |
| Grok Build 0.1 | 49.2% | 68.9% | estimated ± 8.9 pp, medium confidence |
| Hy3 Preview | 36.9% | 99.9% | estimated ± 8.9 pp, low confidence |
| Kimi K2.5 (Reasoning) | 32.6% | 100.0% | estimated ± 8.9 pp, low confidence |
| MiMo-V2.5 | 46.9% | 68.9% | estimated ± 8.9 pp, medium confidence |
| MiMo-V2-Pro | 36.7% | 99.9% | estimated ± 8.9 pp, low confidence |
| Qwen3.5-27B | 39.4% | 87.1% | estimated ± 8.9 pp, medium confidence |
| Qwen3.5-35B-A3B | 29.0% | 100.0% | estimated ± 8.9 pp, low confidence |
| Qwen3.6-27B | 54.8% | 68.9% | estimated ± 8.9 pp, medium confidence |
| Qwen3 Max | 43.7% | 68.9% | estimated ± 8.9 pp, medium confidence |
| Trinity-Large-Thinking | 32.6% | 100.0% | estimated ± 8.9 pp, low confidence |