benchgap
Calibration

AA Coding Index → PostTrainBench v1.1

PostTrainBench v1.1 is estimated from AA Coding Index with a Hill curve fitted on 11 models measured on both: y = 0.0831 + (1.2000 − 0.0831)·x^6.00 / (0.93738^6.00 + x^6.00), R² = 0.55, cross-validated error 6.1 pp. It is used for 46 estimates.

Estimated modelAA Coding IndexPostTrainBench v1.1Source
Apodex 1.160.8%16.0%estimated ± 6.1 pp, low confidence
Apodex 1.1 Mini60.8%16.0%estimated ± 6.1 pp, low confidence
Celeris-114.4%8.3%estimated ± 6.1 pp, low confidence
Claude 3 Opus19.5%8.3%estimated ± 6.1 pp, low confidence
Command A+27.9%8.4%estimated ± 6.1 pp, low confidence
DeepSeek V323.0%8.3%estimated ± 6.1 pp, low confidence
Gemini 1.5 Pro23.6%8.3%estimated ± 6.1 pp, low confidence
Gemma 3 27B10.1%8.3%estimated ± 6.1 pp, low confidence
Gemma 4 12B31.0%8.5%estimated ± 6.1 pp, low confidence
Gemma 4 26B A4B39.3%8.9%estimated ± 6.1 pp, low confidence
Gemma 4 31B43.4%9.4%estimated ± 6.1 pp, low confidence
Gemma 4 E2B7.2%8.3%estimated ± 6.1 pp, low confidence
Gemma 4 E4B9.4%8.3%estimated ± 6.1 pp, low confidence
GPT-4.1 mini20.2%8.3%estimated ± 6.1 pp, low confidence
GPT-4.1 nano11.1%8.3%estimated ± 6.1 pp, low confidence
GPT-4 Turbo21.5%8.3%estimated ± 6.1 pp, low confidence
GPT-4o mini11.4%8.3%estimated ± 6.1 pp, low confidence
GPT-5.149.4%10.7%estimated ± 6.1 pp, low confidence
GPT-5 (high)37.8%8.8%estimated ± 6.1 pp, low confidence
GPT-OSS 120B30.4%8.4%estimated ± 6.1 pp, low confidence
GPT-OSS 20B20.7%8.3%estimated ± 6.1 pp, low confidence
Hy358.8%14.7%estimated ± 6.1 pp, low confidence
Hy3 Preview58.8%14.7%estimated ± 6.1 pp, low confidence
K-Exaone32.1%8.5%estimated ± 6.1 pp, low confidence
Kimi K2.5 (Reasoning)46.8%10.0%estimated ± 6.1 pp, low confidence
LFM2.5-2.6B7.7%8.3%estimated ± 6.1 pp, low confidence
Ling 2.6 Flash25.3%8.4%estimated ± 6.1 pp, low confidence
Ling 3.0 Flash FP850.6%11.0%estimated ± 6.1 pp, low confidence
Llama 4 Maverick16.3%8.3%estimated ± 6.1 pp, low confidence
Llama 4 Scout8.2%8.3%estimated ± 6.1 pp, low confidence
MiMo-V2-Flash49.8%10.8%estimated ± 6.1 pp, low confidence
Mistral Large 320.1%8.3%estimated ± 6.1 pp, low confidence
Mistral Small 426.6%8.4%estimated ± 6.1 pp, low confidence
Mistral Small 4 (Reasoning)26.6%8.4%estimated ± 6.1 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP426.8%8.4%estimated ± 6.1 pp, low confidence
Nemotron 3 Nano 30B14.4%8.3%estimated ± 6.1 pp, low confidence
Nemotron 3 Nano Omni 30B A3B13.8%8.3%estimated ± 6.1 pp, low confidence
Nemotron 3 Super 100B37.7%8.8%estimated ± 6.1 pp, low confidence
o139.7%9.0%estimated ± 6.1 pp, low confidence
o1-preview34.1%8.6%estimated ± 6.1 pp, low confidence
Quasar 438B61.2%16.3%estimated ± 6.1 pp, low confidence
Qwen3.5-122B-A10B45.7%9.8%estimated ± 6.1 pp, low confidence
Qwen3.8 Max Preview71.8%27.1%estimated ± 6.1 pp, medium confidence
Trinity-Large-Preview25.8%8.4%estimated ± 6.1 pp, low confidence
Trinity-Large-Thinking25.8%8.4%estimated ± 6.1 pp, low confidence
Ultravox v0.6 Llama 3.3 70B11.9%8.3%estimated ± 6.1 pp, low confidence