benchgap
Calibration

AA Coding Index → React Native Evals

React Native Evals is estimated from AA Coding Index with a linear curve fitted on 8 models measured on both: y = 0.2513·x + 0.6407, R² = 0.74, cross-validated error 3.5 pp. It is used for 50 estimates.

Estimated modelAA Coding IndexReact Native EvalsSource
Apodex 1.160.8%79.3%estimated ± 3.5 pp, high confidence
Apodex 1.1 Mini60.8%79.3%estimated ± 3.5 pp, high confidence
Celeris-114.4%67.7%estimated ± 3.5 pp, medium confidence
Claude 3 Opus19.5%69.0%estimated ± 3.5 pp, medium confidence
Claude Fable 5.181.6%84.6%estimated ± 3.5 pp, medium confidence
Claude Opus 4.7 (Adaptive)73.6%82.6%estimated ± 3.5 pp, high confidence
Command A+27.9%71.1%estimated ± 3.5 pp, high confidence
DeepSeek V323.0%69.9%estimated ± 3.5 pp, high confidence
Gemini 1.5 Pro23.6%70.0%estimated ± 3.5 pp, high confidence
Gemma 3 27B10.1%66.6%estimated ± 3.5 pp, medium confidence
Gemma 4 12B31.0%71.9%estimated ± 3.5 pp, high confidence
Gemma 4 26B A4B39.3%74.0%estimated ± 3.5 pp, high confidence
Gemma 4 E2B7.2%65.9%estimated ± 3.5 pp, medium confidence
Gemma 4 E4B9.4%66.4%estimated ± 3.5 pp, medium confidence
GPT-4.1 mini20.2%69.2%estimated ± 3.5 pp, medium confidence
GPT-4.1 nano11.1%66.9%estimated ± 3.5 pp, medium confidence
GPT-4 Turbo21.5%69.5%estimated ± 3.5 pp, high confidence
GPT-4o mini11.4%66.9%estimated ± 3.5 pp, medium confidence
GPT-6 Astra76.9%83.4%estimated ± 3.5 pp, medium confidence
Granite 4.2 8B22.4%69.7%estimated ± 3.5 pp, high confidence
Hy358.8%78.8%estimated ± 3.5 pp, high confidence
Hy3 Preview58.8%78.8%estimated ± 3.5 pp, high confidence
K-Exaone32.1%72.1%estimated ± 3.5 pp, high confidence
LFM2.5-2.6B7.7%66.0%estimated ± 3.5 pp, medium confidence
Ling 2.6 Flash25.3%70.4%estimated ± 3.5 pp, high confidence
Ling 3.0 Flash FP850.6%76.8%estimated ± 3.5 pp, high confidence
Llama 4 Maverick16.3%68.2%estimated ± 3.5 pp, medium confidence
Llama 4 Scout8.2%66.1%estimated ± 3.5 pp, medium confidence
MiMo-V2-Flash49.8%76.6%estimated ± 3.5 pp, high confidence
Mistral Large 320.1%69.1%estimated ± 3.5 pp, medium confidence
Mistral Small 426.6%70.8%estimated ± 3.5 pp, high confidence
Mistral Small 4 (Reasoning)26.6%70.8%estimated ± 3.5 pp, high confidence
Muse Glimmer 30B49.0%76.4%estimated ± 3.5 pp, high confidence
Muse Spark 1.375.8%83.1%estimated ± 3.5 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP426.8%70.8%estimated ± 3.5 pp, high confidence
Nemotron 3 Nano 30B14.4%67.7%estimated ± 3.5 pp, medium confidence
Nemotron 3 Nano Omni 30B A3B13.8%67.5%estimated ± 3.5 pp, medium confidence
Nemotron 3 Super 100B37.7%73.6%estimated ± 3.5 pp, high confidence
o139.7%74.1%estimated ± 3.5 pp, high confidence
o1-preview34.1%72.6%estimated ± 3.5 pp, high confidence
Quasar 438B61.2%79.5%estimated ± 3.5 pp, high confidence
Qwen3.5-122B-A10B45.7%75.6%estimated ± 3.5 pp, high confidence
Qwen3.6-35B-A3B41.9%74.6%estimated ± 3.5 pp, high confidence
Qwen3.7 Plus55.9%78.1%estimated ± 3.5 pp, high confidence
Qwen3.8-Flash-Next73.1%82.4%estimated ± 3.5 pp, high confidence
Qwen3.8 Max Preview71.8%82.1%estimated ± 3.5 pp, high confidence
Step 3.7 Flash39.6%74.0%estimated ± 3.5 pp, high confidence
Trinity-Large-Preview25.8%70.5%estimated ± 3.5 pp, high confidence
Trinity-Large-Thinking25.8%70.5%estimated ± 3.5 pp, high confidence
Ultravox v0.6 Llama 3.3 70B11.9%67.1%estimated ± 3.5 pp, medium confidence