Calibration
τ³-bench results → DeepPlanning
DeepPlanning is estimated from τ³-bench results with a offset logistic curve fitted on 6 models measured on both: y = 0.1297 + (0.3539 − 0.1297) / (1 + exp(−200.00·(x − 0.6698))), R² = 0.80, cross-validated error 8.3 pp. It is used for 12 estimates.
| Estimated model | τ³-bench results | DeepPlanning | Source |
|---|---|---|---|
| Atria Dawn Preview | 41.2% | 13.0% | estimated ± 8.3 pp, low confidence |
| GLM-5.1 | 70.6% | 35.4% | estimated ± 8.3 pp, low confidence |
| Granite 4.2 30B | 62.0% | 13.0% | estimated ± 8.3 pp, low confidence |
| Granite 4.2 3B | 45.8% | 13.0% | estimated ± 8.3 pp, low confidence |
| Granite 4.2 8B | 58.1% | 13.0% | estimated ± 8.3 pp, low confidence |
| LFM2.5-2.6B | 5.7% | 13.0% | estimated ± 8.3 pp, low confidence |
| Mercury 2.5 | 96.0% | 35.4% | estimated ± 8.3 pp, low confidence |
| MiMo-V2.5-Pro | 72.9% | 35.4% | estimated ± 8.3 pp, low confidence |
| Mistral Medium 3.5 128B | 91.4% | 35.4% | estimated ± 8.3 pp, low confidence |
| Nemotron 3.5 Lightning 30B A3B NVFP4 | 9.5% | 13.0% | estimated ± 8.3 pp, low confidence |
| Nemotron 3 Ultra | 70.9% | 35.4% | estimated ± 8.3 pp, low confidence |
| Pokee-Isaac 28B | 66.2% | 16.9% | estimated ± 8.3 pp, low confidence |