benchgap
Calibration

React Native Evals → Vals SWE-bench

Vals SWE-bench is estimated from React Native Evals with a Michaelis–Menten curve fitted on 5 models measured on both: y = 2.0000·x / (1.22077 + x), R² = 0.87, cross-validated error 1.4 pp. It is used for 11 estimates.

Estimated modelReact Native EvalsVals SWE-benchSource
Claude Opus 4.684.1%81.6%estimated ± 1.4 pp, medium confidence
Composer 296.1%88.1%estimated ± 1.4 pp, low confidence
Composer 2 Fast94.9%87.5%estimated ± 1.4 pp, low confidence
DeepSeek V3.271.5%73.9%estimated ± 1.4 pp, medium confidence
Gemma 4 31B75.2%76.2%estimated ± 1.4 pp, medium confidence
GLM-574.8%76.0%estimated ± 1.4 pp, medium confidence
GPT-5.485.3%82.3%estimated ± 1.4 pp, low confidence
GPT-OSS 120B71.6%73.9%estimated ± 1.4 pp, medium confidence
GPT-OSS 20B71.0%73.5%estimated ± 1.4 pp, low confidence
Grok 472.6%74.6%estimated ± 1.4 pp, medium confidence
Kimi K2.577.2%77.5%estimated ± 1.4 pp, medium confidence