benchgap
Calibration

AA ITBench → τ²-bench results

τ²-bench results is estimated from AA ITBench with a offset logistic curve fitted on 5 models measured on both: y = 0.9993 + (0.8541 − 0.9993) / (1 + exp(−200.00·(x − 0.4634))), R² = 0.96, cross-validated error 4.8 pp. It is used for 8 estimates.

Estimated modelAA ITBenchτ²-bench resultsSource
Claude Fable 5.149.5%85.4%estimated ± 4.8 pp, medium confidence
Claude Opus 5.538.2%99.9%estimated ± 4.8 pp, low confidence
DeepSeek V4.1 Flash46.9%89.0%estimated ± 4.8 pp, medium confidence
Gemini 3.8 Flash52.5%85.4%estimated ± 4.8 pp, medium confidence
GPT-6 Astra48.6%85.6%estimated ± 4.8 pp, medium confidence
GPT-6 Sol49.4%85.4%estimated ± 4.8 pp, medium confidence
Grok 4.742.1%99.9%estimated ± 4.8 pp, low confidence
Muse Spark 1.333.2%99.9%estimated ± 4.8 pp, low confidence