benchgap
Calibration

AA-SciCode → React Native Evals

React Native Evals is estimated from AA-SciCode with a linear curve fitted on 6 models measured on both: y = 0.4248·x + 0.5543, R² = 0.56, cross-validated error 4.7 pp. It is used for 24 estimates.

Estimated modelAA-SciCodeReact Native EvalsSource
A.X K241.0%72.8%estimated ± 4.7 pp, medium confidence
Claude Haiku 5.555.0%78.8%estimated ± 4.7 pp, medium confidence
Claude Opus 5.566.9%83.8%estimated ± 4.7 pp, low confidence
Claude Sonnet 5.561.0%81.3%estimated ± 4.7 pp, low confidence
DeepSeek V3 032439.0%72.0%estimated ± 4.7 pp, medium confidence
DeepSeek V4.1 Flash51.9%77.5%estimated ± 4.7 pp, medium confidence
GPT-6.1 Sol54.2%78.5%estimated ± 4.7 pp, medium confidence
GPT-6 Luna54.6%78.6%estimated ± 4.7 pp, medium confidence
GPT-6 Sol57.6%79.9%estimated ± 4.7 pp, medium confidence
Granite 4.2 30B37.8%71.5%estimated ± 4.7 pp, medium confidence
Granite 4.2 3B25.3%66.2%estimated ± 4.7 pp, low confidence
Grok 4.757.4%79.8%estimated ± 4.7 pp, medium confidence
K-EXAONE 2.042.0%73.3%estimated ± 4.7 pp, medium confidence
Ling 3.0 Flash VL44.2%74.2%estimated ± 4.7 pp, medium confidence
Ling 3.0 Tiny24.2%65.7%estimated ± 4.7 pp, low confidence
Ling 3.1 Flash54.1%78.4%estimated ± 4.7 pp, medium confidence
Mercury 2.538.5%71.8%estimated ± 4.7 pp, medium confidence
MiMo-V2.6-Flash51.3%77.2%estimated ± 4.7 pp, medium confidence
MiMo-V2.6-Pro60.9%81.3%estimated ± 4.7 pp, low confidence
MiniCPM5-2B26.3%66.6%estimated ± 4.7 pp, low confidence
North Mini Code38.8%71.9%estimated ± 4.7 pp, medium confidence
Solar Pro 325.5%66.3%estimated ± 4.7 pp, low confidence
Solar Pro 444.6%74.4%estimated ± 4.7 pp, medium confidence
Step 5 Preview58.9%80.5%estimated ± 4.7 pp, low confidence