benchgap
Calibration

τ³-bench results → DeepPlanning

DeepPlanning is estimated from τ³-bench results with a offset logistic curve fitted on 6 models measured on both: y = 0.1297 + (0.3539 − 0.1297) / (1 + exp(−200.00·(x − 0.6698))), R² = 0.80, cross-validated error 8.3 pp. It is used for 12 estimates.

Estimated modelτ³-bench resultsDeepPlanningSource
Atria Dawn Preview41.2%13.0%estimated ± 8.3 pp, low confidence
GLM-5.170.6%35.4%estimated ± 8.3 pp, low confidence
Granite 4.2 30B62.0%13.0%estimated ± 8.3 pp, low confidence
Granite 4.2 3B45.8%13.0%estimated ± 8.3 pp, low confidence
Granite 4.2 8B58.1%13.0%estimated ± 8.3 pp, low confidence
LFM2.5-2.6B5.7%13.0%estimated ± 8.3 pp, low confidence
Mercury 2.596.0%35.4%estimated ± 8.3 pp, low confidence
MiMo-V2.5-Pro72.9%35.4%estimated ± 8.3 pp, low confidence
Mistral Medium 3.5 128B91.4%35.4%estimated ± 8.3 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP49.5%13.0%estimated ± 8.3 pp, low confidence
Nemotron 3 Ultra70.9%35.4%estimated ± 8.3 pp, low confidence
Pokee-Isaac 28B66.2%16.9%estimated ± 8.3 pp, low confidence