benchgap
Calibration

CharXiv → SimpleVQA

SimpleVQA is estimated from CharXiv with a Michaelis–Menten curve fitted on 8 models measured on both: y = 2.0000·x / (1.60507 + x), R² = 0.42, cross-validated error 7.8 pp. It is used for 18 estimates.

Estimated modelCharXivSimpleVQASource
Claude Mythos 593.5%73.6%estimated ± 7.8 pp, low confidence
Claude Opus 4.7 (Adaptive)91.0%72.4%estimated ± 7.8 pp, low confidence
Claude Sonnet 4.677.4%65.1%estimated ± 7.8 pp, low confidence
Claude Sonnet 588.3%71.0%estimated ± 7.8 pp, low confidence
Command A+52.7%49.4%estimated ± 7.8 pp, low confidence
Gemini 3.1 Flash-Lite73.2%62.6%estimated ± 7.8 pp, low confidence
Gemini 3.5 Flash84.2%68.8%estimated ± 7.8 pp, low confidence
Gemini 3.7 Flash88.7%71.2%estimated ± 7.8 pp, low confidence
GLM-5.3-Flash89.4%71.5%estimated ± 7.8 pp, low confidence
GPT-5.282.1%67.7%estimated ± 7.8 pp, low confidence
Inkling82.0%67.6%estimated ± 7.8 pp, low confidence
Inkling-Small81.3%67.2%estimated ± 7.8 pp, low confidence
Kimi K2.680.4%66.7%estimated ± 7.8 pp, low confidence
MiMo-V2.581.0%67.1%estimated ± 7.8 pp, low confidence
Muse Spark 1.188.4%71.0%estimated ± 7.8 pp, low confidence
Qwen3.5-122B-A10B77.2%65.0%estimated ± 7.8 pp, low confidence
Sakana Fugu85.1%69.3%estimated ± 7.8 pp, low confidence
Sakana Fugu-Ultra86.6%70.1%estimated ± 7.8 pp, low confidence