benchgap
Calibration

AA Coding Index → AA-SciCode

AA-SciCode is estimated from AA Coding Index with a linear curve fitted on 78 models measured on both: y = 0.4473·x + 0.2248, R² = 0.87, cross-validated error 3.2 pp. It is used for 21 estimates.

Estimated modelAA Coding IndexAA-SciCodeSource
Claude 3 Opus19.5%31.2%estimated ± 3.2 pp, high confidence
Claude Opus 4.7 (Adaptive)73.6%55.4%estimated ± 3.2 pp, high confidence
Gemini 1.5 Pro23.6%33.1%estimated ± 3.2 pp, high confidence
Gemma 4 12B31.0%36.3%estimated ± 3.2 pp, high confidence
Gemma 4 E2B7.2%25.7%estimated ± 3.2 pp, medium confidence
GLM-4.745.3%42.7%estimated ± 3.2 pp, high confidence
GPT-4.1 mini20.2%31.5%estimated ± 3.2 pp, high confidence
GPT-4.1 nano11.1%27.5%estimated ± 3.2 pp, high confidence
GPT-4 Turbo21.5%32.1%estimated ± 3.2 pp, high confidence
GPT-4o mini11.4%27.6%estimated ± 3.2 pp, high confidence
GPT-5.149.4%44.6%estimated ± 3.2 pp, high confidence
GPT-5.471.1%54.3%estimated ± 3.2 pp, high confidence
GPT-5 (high)37.8%39.4%estimated ± 3.2 pp, high confidence
K-Exaone32.1%36.8%estimated ± 3.2 pp, high confidence
Kimi K2.5 (Reasoning)46.8%43.4%estimated ± 3.2 pp, high confidence
MiMo-V2-Flash49.8%44.8%estimated ± 3.2 pp, high confidence
Muse Spark58.6%48.7%estimated ± 3.2 pp, high confidence
o139.7%40.2%estimated ± 3.2 pp, high confidence
o1-preview34.1%37.7%estimated ± 3.2 pp, high confidence
Qwen3.6 Plus54.5%46.9%estimated ± 3.2 pp, high confidence
Ultravox v0.6 Llama 3.3 70B11.9%27.8%estimated ± 3.2 pp, high confidence