Calibration
JobBench → APEX-Agents-AA
APEX-Agents-AA is estimated from JobBench with a Michaelis–Menten curve fitted on 5 models measured on both: y = 0.6741·x / (0.38149 + x), R² = 0.97, cross-validated error 3.6 pp. It is used for 18 estimates.
| Estimated model | JobBench | APEX-Agents-AA | Source |
|---|---|---|---|
| Atria Dawn Preview | 50.3% | 38.3% | estimated ± 3.6 pp, medium confidence |
| Claude 4.1 Opus | 21.9% | 24.6% | estimated ± 3.6 pp, medium confidence |
| Claude 4 Sonnet | 18.4% | 21.9% | estimated ± 3.6 pp, medium confidence |
| Claude Haiku 4.5 | 16.0% | 19.9% | estimated ± 3.6 pp, medium confidence |
| Claude Opus 4.5 | 32.3% | 30.9% | estimated ± 3.6 pp, medium confidence |
| Claude Sonnet 4.5 | 27.7% | 28.4% | estimated ± 3.6 pp, medium confidence |
| Gemini 3 Flash | 11.4% | 15.5% | estimated ± 3.6 pp, medium confidence |
| Gemini 3 Pro | 11.4% | 15.5% | estimated ± 3.6 pp, medium confidence |
| GPT-5.1-Codex | 26.2% | 27.4% | estimated ± 3.6 pp, medium confidence |
| GPT-5.2 | 34.3% | 31.9% | estimated ± 3.6 pp, medium confidence |
| GPT-5.2-Codex | 26.0% | 27.3% | estimated ± 3.6 pp, medium confidence |
| GPT-5.3 Codex | 33.7% | 31.6% | estimated ± 3.6 pp, medium confidence |
| GPT-5 (high) | 8.5% | 12.3% | estimated ± 3.6 pp, low confidence |
| Hy4 preview | 61.7% | 41.7% | estimated ± 3.6 pp, low confidence |
| MiMo-V2.6-Flash | 61.2% | 41.5% | estimated ± 3.6 pp, low confidence |
| MiMo-V2.6-Pro | 62.0% | 41.7% | estimated ± 3.6 pp, low confidence |
| Qwen3.5 Plus | 18.5% | 22.0% | estimated ± 3.6 pp, medium confidence |
| Qwen3.8-27B | 33.4% | 31.5% | estimated ± 3.6 pp, medium confidence |