benchgap
Calibration

Claw-Eval → ResearchClawBench

ResearchClawBench is estimated from Claw-Eval with a Michaelis–Menten curve fitted on 12 models measured on both: y = 2.0000·x / (6.71213 + x), R² = 0.69, cross-validated error 1.3 pp. It is used for 27 estimates.

Estimated modelClaw-EvalResearchClawBenchSource
Claude Opus 4.559.6%16.3%estimated ± 1.3 pp, high confidence
Claude Sonnet 4.667.8%18.3%estimated ± 1.3 pp, high confidence
DeepSeek V3.240.2%11.3%estimated ± 1.3 pp, medium confidence
dots3-note Preview73.4%19.7%estimated ± 1.3 pp, high confidence
Gemini 3 Flash49.2%13.7%estimated ± 1.3 pp, medium confidence
GLM-557.7%15.8%estimated ± 1.3 pp, high confidence
GLM-5-Turbo55.8%15.4%estimated ± 1.3 pp, high confidence
GLM-5V-Turbo53.8%14.8%estimated ± 1.3 pp, high confidence
K-EXAONE 2.077.7%20.8%estimated ± 1.3 pp, medium confidence
LFM2.5-2.6B62.8%17.1%estimated ± 1.3 pp, high confidence
LLaDA2.2-flash64.2%17.5%estimated ± 1.3 pp, high confidence
LLaDA2.2-mini57.2%15.7%estimated ± 1.3 pp, high confidence
MiMo-V2.5-Pro63.8%17.4%estimated ± 1.3 pp, high confidence
MiMo-V2-Omni45.2%12.6%estimated ± 1.3 pp, medium confidence
MiniMax M2.748.7%13.5%estimated ± 1.3 pp, medium confidence
Muse Spark63.8%17.4%estimated ± 1.3 pp, high confidence
Nemotron 3 Super 100B5.5%1.6%estimated ± 1.3 pp, medium confidence
Ornith-1.0-35B69.8%18.8%estimated ± 1.3 pp, high confidence
Ornith-1.0-397B77.1%20.6%estimated ± 1.3 pp, medium confidence
Ornith-1.0-9B63.1%17.2%estimated ± 1.3 pp, high confidence
Ornith-1.5-35B-A3B72.5%19.5%estimated ± 1.3 pp, high confidence
Ornith-1.5-397B81.4%21.6%estimated ± 1.3 pp, medium confidence
Ornith-1.5-9B66.5%18.0%estimated ± 1.3 pp, high confidence
Qwen3.6-27B72.4%19.5%estimated ± 1.3 pp, high confidence
Qwen3.6-35B-A3B68.7%18.6%estimated ± 1.3 pp, high confidence
Qwen3.7 Plus62.7%17.1%estimated ± 1.3 pp, high confidence
Step 3.7 Flash67.1%18.2%estimated ± 1.3 pp, high confidence