Calibration
CyberGym → Agents' Last Exam
Agents' Last Exam is estimated from CyberGym with a offset logistic curve fitted on 8 models measured on both: y = 0.2401 + (0.3035 − 0.2401) / (1 + exp(−182.53·(x − 0.8388))), R² = 0.80, cross-validated error 2.1 pp. It is used for 20 estimates.
| Estimated model | CyberGym | Agents' Last Exam | Source |
|---|---|---|---|
| Atria Dawn Preview | 86.5% | 30.3% | estimated ± 2.1 pp, high confidence |
| Claude Mythos 5 | 83.8% | 27.0% | estimated ± 2.1 pp, high confidence |
| Claude Mythos Preview | 83.1% | 25.2% | estimated ± 2.1 pp, high confidence |
| Claude Opus 4.5 | 50.6% | 24.0% | estimated ± 2.1 pp, medium confidence |
| Claude Opus 4.6 | 66.6% | 24.0% | estimated ± 2.1 pp, medium confidence |
| Claude Opus 4.7 (Adaptive) | 73.1% | 24.0% | estimated ± 2.1 pp, medium confidence |
| Claude Sonnet 4.6 | 65.2% | 24.0% | estimated ± 2.1 pp, medium confidence |
| Gemini 3.5 Flash Cyber | 83.2% | 25.4% | estimated ± 2.1 pp, high confidence |
| Gemini 3.8 Flash Cyber | 86.2% | 30.3% | estimated ± 2.1 pp, high confidence |
| GLM-5 | 43.2% | 24.0% | estimated ± 2.1 pp, medium confidence |
| GLM-5.1 | 68.7% | 24.0% | estimated ± 2.1 pp, medium confidence |
| GPT-5.4 | 79.0% | 24.0% | estimated ± 2.1 pp, high confidence |
| GPT-5.5 | 81.8% | 24.2% | estimated ± 2.1 pp, high confidence |
| GPT-5.6 Luna | 77.9% | 24.0% | estimated ± 2.1 pp, high confidence |
| GPT-5.6 Sol | 84.5% | 28.8% | estimated ± 2.1 pp, high confidence |
| GPT-5.6 Terra | 81.8% | 24.2% | estimated ± 2.1 pp, high confidence |
| Ling 3.1 Flash | 87.9% | 30.3% | estimated ± 2.1 pp, high confidence |
| Muse Spark | 43.5% | 24.0% | estimated ± 2.1 pp, medium confidence |
| Muse Spark 1.1 | 59.0% | 24.0% | estimated ± 2.1 pp, medium confidence |
| Fugu Cyber | 86.9% | 30.3% | estimated ± 2.1 pp, high confidence |