benchgap
Calibration

OSWorld 2.0 → APEX-Agents-AA

APEX-Agents-AA is estimated from OSWorld 2.0 with a Michaelis–Menten + offset curve fitted on 5 models measured on both: y = 0.0000 + 0.3993·x / (0.01920 + x), R² = 0.91, cross-validated error 2.9 pp. It is used for 18 estimates.

Estimated modelOSWorld 2.0APEX-Agents-AASource
Claude Fable 5.141.7%38.2%estimated ± 2.9 pp, medium confidence
Claude Opus 4.713.9%35.1%estimated ± 2.9 pp, medium confidence
Claude Opus 4.7 (Adaptive)18.2%36.1%estimated ± 2.9 pp, medium confidence
Claude Opus 4.820.6%36.5%estimated ± 2.9 pp, medium confidence
Claude Opus 570.6%38.9%estimated ± 2.9 pp, low confidence
Claude Opus 5.548.7%38.4%estimated ± 2.9 pp, medium confidence
Claude Sonnet 4.68.3%32.4%estimated ± 2.9 pp, medium confidence
Gemini 3.7 Flash47.9%38.4%estimated ± 2.9 pp, medium confidence
Gemini 3.8 Flash59.0%38.7%estimated ± 2.9 pp, low confidence
Gemini 4 Argon69.2%38.9%estimated ± 2.9 pp, low confidence
GPT-5.6 Sol62.6%38.7%estimated ± 2.9 pp, low confidence
GPT-6 Astra72.6%38.9%estimated ± 2.9 pp, low confidence
GPT-6 Sol60.5%38.7%estimated ± 2.9 pp, low confidence
MiniMax M34.6%28.2%estimated ± 2.9 pp, medium confidence
Muse Spark 1.114.2%35.2%estimated ± 2.9 pp, medium confidence
Muse Spark 1.366.9%38.8%estimated ± 2.9 pp, low confidence
Qwen3.8-Flash-Next19.4%36.3%estimated ± 2.9 pp, medium confidence
Qwen3.8 Max19.4%36.3%estimated ± 2.9 pp, medium confidence