benchgap
Calibration

SWE-bench Verified → React Native Evals

React Native Evals is estimated from SWE-bench Verified with a offset logistic curve fitted on 6 models measured on both: y = 0.7137 + (1.2000 − 0.7137) / (1 + exp(−33.76·(x − 0.8396))), R² = 0.94, cross-validated error 2.3 pp. It is used for 84 estimates.

Estimated modelSWE-bench VerifiedReact Native EvalsSource
Claude Fable 595.0%100.0%estimated ± 2.3 pp, low confidence
Claude Mythos 595.5%100.0%estimated ± 2.3 pp, low confidence
Claude Opus 4.7 (Adaptive)87.6%100.0%estimated ± 2.3 pp, low confidence
Claude Opus 4.888.6%100.0%estimated ± 2.3 pp, low confidence
Claude Opus 596.0%100.0%estimated ± 2.3 pp, low confidence
Claude Sonnet 585.2%100.0%estimated ± 2.3 pp, low confidence
Ember-192.2%100.0%estimated ± 2.3 pp, low confidence
Ornith-1.5-397B86.0%100.0%estimated ± 2.3 pp, low confidence
GPT-5.3 Codex85.0%99.9%estimated ± 2.3 pp, low confidence
Ornith-1.0-397B82.4%89.4%estimated ± 2.3 pp, low confidence
Claude Opus 4.580.9%84.1%estimated ± 2.3 pp, low confidence
Beam80.9%84.1%estimated ± 2.3 pp, low confidence
DeepSeek V4 Pro 081380.6%83.2%estimated ± 2.3 pp, medium confidence
MiniMax M380.5%82.9%estimated ± 2.3 pp, medium confidence
Qwen3.7 Max80.4%82.6%estimated ± 2.3 pp, medium confidence
Inkling-Small80.2%82.0%estimated ± 2.3 pp, medium confidence
Kimi K2.680.2%82.0%estimated ± 2.3 pp, medium confidence
GPT-5.280.0%81.5%estimated ± 2.3 pp, medium confidence
DeepSeek V4 Flash 073179.0%79.0%estimated ± 2.3 pp, medium confidence
Ornith-1.5-35B-A3B79.0%79.0%estimated ± 2.3 pp, medium confidence
Qwen3.6 Plus78.8%78.6%estimated ± 2.3 pp, medium confidence
BTL-478.4%77.8%estimated ± 2.3 pp, medium confidence
dots3-note Preview78.4%77.8%estimated ± 2.3 pp, medium confidence
MiMo-V2-Pro78.0%77.1%estimated ± 2.3 pp, medium confidence
Apodex 1.177.7%76.6%estimated ± 2.3 pp, medium confidence
Qwen3.7 Plus77.7%76.6%estimated ± 2.3 pp, medium confidence
Inkling77.6%76.5%estimated ± 2.3 pp, medium confidence
Mistral Medium 3.5 128B77.6%76.5%estimated ± 2.3 pp, medium confidence
Muse Spark77.4%76.2%estimated ± 2.3 pp, medium confidence
Claude Sonnet 4.577.2%75.9%estimated ± 2.3 pp, medium confidence
Qwen3.6-27B77.2%75.9%estimated ± 2.3 pp, medium confidence
Kimi K2.5 (Reasoning)76.8%75.3%estimated ± 2.3 pp, medium confidence
Grok 4.2076.7%75.2%estimated ± 2.3 pp, medium confidence
Qwen3.5 397B76.2%74.7%estimated ± 2.3 pp, medium confidence
Muse Glimmer 30B76.0%74.5%estimated ± 2.3 pp, medium confidence
Ornith-1.0-35B75.6%74.1%estimated ± 2.3 pp, medium confidence
MiMo-V2-Omni74.8%73.5%estimated ± 2.3 pp, medium confidence
Laguna M.174.6%73.3%estimated ± 2.3 pp, medium confidence
Claude 4.1 Opus74.5%73.3%estimated ± 2.3 pp, medium confidence
Hy3 Preview74.4%73.2%estimated ± 2.3 pp, medium confidence
Step 3.5 Flash74.4%73.2%estimated ± 2.3 pp, medium confidence
GLM-4.773.8%72.9%estimated ± 2.3 pp, medium confidence
MAI-Thinking-173.5%72.8%estimated ± 2.3 pp, medium confidence
MiMo-V2-Flash73.4%72.7%estimated ± 2.3 pp, medium confidence
Qwen3.6-35B-A3B73.4%72.7%estimated ± 2.3 pp, medium confidence
Claude Haiku 4.573.3%72.7%estimated ± 2.3 pp, medium confidence
Claude 4 Sonnet72.7%72.4%estimated ± 2.3 pp, medium confidence
MAI-Code-1.1-Flash72.6%72.4%estimated ± 2.3 pp, medium confidence
Qwen3.5-27B72.4%72.3%estimated ± 2.3 pp, medium confidence
Qwen3.5-122B-A10B72.0%72.2%estimated ± 2.3 pp, medium confidence
Nemotron 3 Ultra71.9%72.2%estimated ± 2.3 pp, medium confidence
Laguna XS 2.170.9%72.0%estimated ± 2.3 pp, medium confidence
Grok Code Fast 170.8%71.9%estimated ± 2.3 pp, medium confidence
Ornith-1.5-9B70.6%71.9%estimated ± 2.3 pp, medium confidence
Solar Pro 470.6%71.9%estimated ± 2.3 pp, medium confidence
Solar Open 270.4%71.9%estimated ± 2.3 pp, medium confidence
Laguna XS.269.9%71.8%estimated ± 2.3 pp, medium confidence
Ornith-1.0-9B69.4%71.7%estimated ± 2.3 pp, medium confidence
Qwen3.5-35B-A3B69.2%71.7%estimated ± 2.3 pp, medium confidence
K-EXAONE 2.068.2%71.6%estimated ± 2.3 pp, medium confidence
LongCat-Flash-Lite-Sparse68.2%71.6%estimated ± 2.3 pp, medium confidence
DeepSeek V3.166.0%71.5%estimated ± 2.3 pp, medium confidence
Kimi K265.8%71.5%estimated ± 2.3 pp, medium confidence
Gemini 2.5 Pro63.8%71.4%estimated ± 2.3 pp, medium confidence
Ternary Bonsai 2 27B60.8%71.4%estimated ± 2.3 pp, medium confidence
Nemotron 3 Super 100B60.5%71.4%estimated ± 2.3 pp, low confidence
GLM-4.7-Flash59.2%71.4%estimated ± 2.3 pp, low confidence
Granite 4.2 30B57.0%71.4%estimated ± 2.3 pp, low confidence
MiniMax M1 80k56.0%71.4%estimated ± 2.3 pp, low confidence
GPT-4.154.6%71.4%estimated ± 2.3 pp, low confidence
ZAYA1-74B-Preview53.2%71.4%estimated ± 2.3 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP452.8%71.4%estimated ± 2.3 pp, low confidence
K-Exaone49.4%71.4%estimated ± 2.3 pp, low confidence
o3-mini49.3%71.4%estimated ± 2.3 pp, low confidence
LLaDA2.2-flash49.3%71.4%estimated ± 2.3 pp, low confidence
DeepSeek-R149.2%71.4%estimated ± 2.3 pp, low confidence
Claude 3.5 Sonnet49.0%71.4%estimated ± 2.3 pp, low confidence
Granite 4.2 8B47.7%71.4%estimated ± 2.3 pp, low confidence
Mellum2.1-12B-A2.5B-Thinking47.0%71.4%estimated ± 2.3 pp, low confidence
MiniCPM5-2B46.4%71.4%estimated ± 2.3 pp, low confidence
Sarvam 105B45.0%71.4%estimated ± 2.3 pp, low confidence
DeepSeek V342.0%71.4%estimated ± 2.3 pp, low confidence
Sarvam 30B34.0%71.4%estimated ± 2.3 pp, low confidence
GPT-4.1 mini23.6%71.4%estimated ± 2.3 pp, low confidence