benchgap
Calibration

JobBench → APEX-Agents-AA

APEX-Agents-AA is estimated from JobBench with a Michaelis–Menten curve fitted on 5 models measured on both: y = 0.6741·x / (0.38149 + x), R² = 0.97, cross-validated error 3.6 pp. It is used for 18 estimates.

Estimated modelJobBenchAPEX-Agents-AASource
Atria Dawn Preview50.3%38.3%estimated ± 3.6 pp, medium confidence
Claude 4.1 Opus21.9%24.6%estimated ± 3.6 pp, medium confidence
Claude 4 Sonnet18.4%21.9%estimated ± 3.6 pp, medium confidence
Claude Haiku 4.516.0%19.9%estimated ± 3.6 pp, medium confidence
Claude Opus 4.532.3%30.9%estimated ± 3.6 pp, medium confidence
Claude Sonnet 4.527.7%28.4%estimated ± 3.6 pp, medium confidence
Gemini 3 Flash11.4%15.5%estimated ± 3.6 pp, medium confidence
Gemini 3 Pro11.4%15.5%estimated ± 3.6 pp, medium confidence
GPT-5.1-Codex26.2%27.4%estimated ± 3.6 pp, medium confidence
GPT-5.234.3%31.9%estimated ± 3.6 pp, medium confidence
GPT-5.2-Codex26.0%27.3%estimated ± 3.6 pp, medium confidence
GPT-5.3 Codex33.7%31.6%estimated ± 3.6 pp, medium confidence
GPT-5 (high)8.5%12.3%estimated ± 3.6 pp, low confidence
Hy4 preview61.7%41.7%estimated ± 3.6 pp, low confidence
MiMo-V2.6-Flash61.2%41.5%estimated ± 3.6 pp, low confidence
MiMo-V2.6-Pro62.0%41.7%estimated ± 3.6 pp, low confidence
Qwen3.5 Plus18.5%22.0%estimated ± 3.6 pp, medium confidence
Qwen3.8-27B33.4%31.5%estimated ± 3.6 pp, medium confidence