benchgap
Calibration

CharXiv → ERQA

ERQA is estimated from CharXiv with a Michaelis–Menten curve fitted on 12 models measured on both: y = 2.0000·x / (1.62569 + x), R² = 0.68, cross-validated error 3.7 pp. It is used for 18 estimates.

Estimated modelCharXivERQASource
Claude Mythos 593.5%73.0%estimated ± 3.7 pp, high confidence
Claude Opus 4.7 (Adaptive)91.0%71.8%estimated ± 3.7 pp, high confidence
Claude Opus 4.889.9%71.2%estimated ± 3.7 pp, high confidence
Claude Sonnet 4.677.4%64.5%estimated ± 3.7 pp, high confidence
Claude Sonnet 588.3%70.4%estimated ± 3.7 pp, high confidence
Command A+52.7%49.0%estimated ± 3.7 pp, medium confidence
Gemini 3.1 Flash-Lite73.2%62.1%estimated ± 3.7 pp, high confidence
Gemini 3.5 Flash84.2%68.2%estimated ± 3.7 pp, high confidence
Gemini 3.7 Flash88.7%70.6%estimated ± 3.7 pp, high confidence
GLM-5.3-Flash89.4%71.0%estimated ± 3.7 pp, high confidence
Inkling82.0%67.1%estimated ± 3.7 pp, high confidence
Inkling-Small81.3%66.7%estimated ± 3.7 pp, high confidence
MiMo-V2.581.0%66.5%estimated ± 3.7 pp, high confidence
Muse Glimmer 30B78.8%65.3%estimated ± 3.7 pp, high confidence
Muse Spark 1.188.4%70.4%estimated ± 3.7 pp, high confidence
Nemotron 3 Nano Omni 30B A3B76.3%63.9%estimated ± 3.7 pp, high confidence
Sakana Fugu85.1%68.7%estimated ± 3.7 pp, high confidence
Sakana Fugu-Ultra86.6%69.5%estimated ± 3.7 pp, high confidence