Calibration
AA ITBench → τ²-bench results
τ²-bench results is estimated from AA ITBench with a offset logistic curve fitted on 5 models measured on both: y = 0.9993 + (0.8541 − 0.9993) / (1 + exp(−200.00·(x − 0.4634))), R² = 0.96, cross-validated error 4.8 pp. It is used for 8 estimates.
| Estimated model | AA ITBench | τ²-bench results | Source |
|---|---|---|---|
| Claude Fable 5.1 | 49.5% | 85.4% | estimated ± 4.8 pp, medium confidence |
| Claude Opus 5.5 | 38.2% | 99.9% | estimated ± 4.8 pp, low confidence |
| DeepSeek V4.1 Flash | 46.9% | 89.0% | estimated ± 4.8 pp, medium confidence |
| Gemini 3.8 Flash | 52.5% | 85.4% | estimated ± 4.8 pp, medium confidence |
| GPT-6 Astra | 48.6% | 85.6% | estimated ± 4.8 pp, medium confidence |
| GPT-6 Sol | 49.4% | 85.4% | estimated ± 4.8 pp, medium confidence |
| Grok 4.7 | 42.1% | 99.9% | estimated ± 4.8 pp, low confidence |
| Muse Spark 1.3 | 33.2% | 99.9% | estimated ± 4.8 pp, low confidence |