benchgap
Calibration

AA AutomationBench → CWE-bench v1

CWE-bench v1 is estimated from AA AutomationBench with a Hill curve fitted on 10 models measured on both: y = 0.3702 + (0.7855 − 0.3702)·x^4.86 / (0.61807^4.86 + x^4.86), R² = 0.75, cross-validated error 6.0 pp. It is used for 15 estimates.

Estimated modelAA AutomationBenchCWE-bench v1Source
Claude Haiku 5.535.4%39.6%estimated ± 6.0 pp, medium confidence
Claude Sonnet 5.571.8%65.0%estimated ± 6.0 pp, medium confidence
Gemini 3.8 Flash59.9%56.2%estimated ± 6.0 pp, medium confidence
GLM-5.362.2%58.1%estimated ± 6.0 pp, medium confidence
GLM-5.3-Flash60.4%56.6%estimated ± 6.0 pp, medium confidence
GPT-6.1 Sol64.9%60.2%estimated ± 6.0 pp, medium confidence
GPT-6 Luna53.2%50.5%estimated ± 6.0 pp, medium confidence
Kimi K358.3%54.9%estimated ± 6.0 pp, medium confidence
MiMo-V2.6-Pro58.6%55.1%estimated ± 6.0 pp, medium confidence
MiniMax M321.3%37.2%estimated ± 6.0 pp, medium confidence
Mistral Large 459.9%56.2%estimated ± 6.0 pp, medium confidence
Muse Glimmer 30B6.8%37.0%estimated ± 6.0 pp, medium confidence
Nemotron 3 Ultra3.0%37.0%estimated ± 6.0 pp, low confidence
Qwen3.8-27B48.2%46.6%estimated ± 6.0 pp, medium confidence
Step 5 Preview51.0%48.7%estimated ± 6.0 pp, medium confidence