benchgap
Calibration

OSWorld 2.0 → JobBench

JobBench is estimated from OSWorld 2.0 with a Michaelis–Menten curve fitted on 7 models measured on both: y = 0.7273·x / (0.07457 + x), R² = 0.80, cross-validated error 4.9 pp. It is used for 13 estimates.

Estimated modelOSWorld 2.0JobBenchSource
Claude Opus 4.713.9%47.3%estimated ± 4.9 pp, medium confidence
Claude Opus 4.820.6%53.4%estimated ± 4.9 pp, medium confidence
Gemini 3.7 Flash47.9%62.9%estimated ± 4.9 pp, medium confidence
Gemini 3.8 Flash59.0%64.6%estimated ± 4.9 pp, medium confidence
Gemini 4 Argon69.2%65.7%estimated ± 4.9 pp, low confidence
GPT-5.6 Luna45.6%62.5%estimated ± 4.9 pp, medium confidence
GPT-5.6 Sol62.6%65.0%estimated ± 4.9 pp, medium confidence
GPT-5.6 Terra50.2%63.3%estimated ± 4.9 pp, medium confidence
GPT-6 Astra72.6%66.0%estimated ± 4.9 pp, low confidence
GPT-6 Sol60.5%64.8%estimated ± 4.9 pp, medium confidence
Kimi K2.64.6%27.7%estimated ± 4.9 pp, low confidence
MiniMax M34.6%27.7%estimated ± 4.9 pp, low confidence
Qwen3.7 Plus2.8%19.9%estimated ± 4.9 pp, low confidence