benchgap
Calibration

OSWorld 2.0 → Toolathlon-Verified

Toolathlon-Verified is estimated from OSWorld 2.0 with a Michaelis–Menten curve fitted on 5 models measured on both: y = 0.8288·x / (0.02655 + x), R² = 0.96, cross-validated error 1.0 pp. It is used for 18 estimates.

Estimated modelOSWorld 2.0Toolathlon-VerifiedSource
Claude Opus 4.713.9%69.6%estimated ± 1.0 pp, low confidence
Claude Opus 4.7 (Adaptive)18.2%72.3%estimated ± 1.0 pp, low confidence
Claude Opus 4.820.6%73.4%estimated ± 1.0 pp, medium confidence
Claude Sonnet 4.68.3%62.8%estimated ± 1.0 pp, low confidence
Gemini 3.7 Flash47.9%78.5%estimated ± 1.0 pp, medium confidence
Gemini 3.8 Flash59.0%79.3%estimated ± 1.0 pp, medium confidence
Gemini 4 Argon69.2%79.8%estimated ± 1.0 pp, medium confidence
GPT-5.513.0%68.8%estimated ± 1.0 pp, low confidence
GPT-5.6 Luna45.6%78.3%estimated ± 1.0 pp, medium confidence
GPT-5.6 Sol62.6%79.5%estimated ± 1.0 pp, medium confidence
GPT-5.6 Terra50.2%78.7%estimated ± 1.0 pp, medium confidence
GPT-6 Astra72.6%80.0%estimated ± 1.0 pp, low confidence
GPT-6 Sol60.5%79.4%estimated ± 1.0 pp, medium confidence
Kimi K2.64.6%52.5%estimated ± 1.0 pp, low confidence
MiniMax M34.6%52.5%estimated ± 1.0 pp, low confidence
Muse Spark 1.114.2%69.8%estimated ± 1.0 pp, low confidence
Muse Spark 1.366.9%79.7%estimated ± 1.0 pp, medium confidence
Qwen3.7 Plus2.8%42.5%estimated ± 1.0 pp, low confidence