benchgap
Calibration

CharXiv → MathVision w/ Python

MathVision w/ Python is estimated from CharXiv with a offset logistic curve fitted on 5 models measured on both: y = 0.7681 + (0.9770 − 0.7681) / (1 + exp(−138.50·(x − 0.8894))), R² = 0.80, cross-validated error 0.9 pp. It is used for 24 estimates.

Estimated modelCharXivMathVision w/ PythonSource
Claude Mythos 593.5%97.7%estimated ± 0.9 pp, medium confidence
Claude Opus 4.7 (Adaptive)91.0%96.6%estimated ± 0.9 pp, medium confidence
Claude Opus 4.889.9%93.3%estimated ± 0.9 pp, low confidence
Claude Sonnet 4.677.4%76.8%estimated ± 0.9 pp, low confidence
Claude Sonnet 588.3%82.9%estimated ± 0.9 pp, low confidence
Command A+52.7%76.8%estimated ± 0.9 pp, low confidence
Gemini 3.1 Flash-Lite73.2%76.8%estimated ± 0.9 pp, low confidence
Gemini 3.1 Pro80.2%76.8%estimated ± 0.9 pp, low confidence
Gemini 3.5 Flash84.2%76.8%estimated ± 0.9 pp, low confidence
Gemini 3.7 Flash88.7%85.5%estimated ± 0.9 pp, low confidence
GLM-5.3-Flash89.4%90.4%estimated ± 0.9 pp, low confidence
GPT-5.482.8%76.8%estimated ± 0.9 pp, low confidence
Grok 4.2060.9%76.8%estimated ± 0.9 pp, low confidence
Inkling82.0%76.8%estimated ± 0.9 pp, low confidence
Inkling-Small81.3%76.8%estimated ± 0.9 pp, low confidence
MiMo-V2.581.0%76.8%estimated ± 0.9 pp, low confidence
Muse Glimmer 30B78.8%76.8%estimated ± 0.9 pp, low confidence
Muse Spark86.4%77.4%estimated ± 0.9 pp, low confidence
Muse Spark 1.188.4%83.5%estimated ± 0.9 pp, low confidence
Nemotron 3 Nano Omni 30B A3B76.3%76.8%estimated ± 0.9 pp, low confidence
Qwen3.6-27B78.4%76.8%estimated ± 0.9 pp, low confidence
Qwen3.6-35B-A3B78.0%76.8%estimated ± 0.9 pp, low confidence
Sakana Fugu85.1%76.9%estimated ± 0.9 pp, low confidence
Sakana Fugu-Ultra86.6%77.6%estimated ± 0.9 pp, low confidence