Calibration
CharXiv → SimpleVQA
SimpleVQA is estimated from CharXiv with a Michaelis–Menten curve fitted on 8 models measured on both: y = 2.0000·x / (1.60507 + x), R² = 0.42, cross-validated error 7.8 pp. It is used for 18 estimates.
| Estimated model | CharXiv | SimpleVQA | Source |
|---|---|---|---|
| Claude Mythos 5 | 93.5% | 73.6% | estimated ± 7.8 pp, low confidence |
| Claude Opus 4.7 (Adaptive) | 91.0% | 72.4% | estimated ± 7.8 pp, low confidence |
| Claude Sonnet 4.6 | 77.4% | 65.1% | estimated ± 7.8 pp, low confidence |
| Claude Sonnet 5 | 88.3% | 71.0% | estimated ± 7.8 pp, low confidence |
| Command A+ | 52.7% | 49.4% | estimated ± 7.8 pp, low confidence |
| Gemini 3.1 Flash-Lite | 73.2% | 62.6% | estimated ± 7.8 pp, low confidence |
| Gemini 3.5 Flash | 84.2% | 68.8% | estimated ± 7.8 pp, low confidence |
| Gemini 3.7 Flash | 88.7% | 71.2% | estimated ± 7.8 pp, low confidence |
| GLM-5.3-Flash | 89.4% | 71.5% | estimated ± 7.8 pp, low confidence |
| GPT-5.2 | 82.1% | 67.7% | estimated ± 7.8 pp, low confidence |
| Inkling | 82.0% | 67.6% | estimated ± 7.8 pp, low confidence |
| Inkling-Small | 81.3% | 67.2% | estimated ± 7.8 pp, low confidence |
| Kimi K2.6 | 80.4% | 66.7% | estimated ± 7.8 pp, low confidence |
| MiMo-V2.5 | 81.0% | 67.1% | estimated ± 7.8 pp, low confidence |
| Muse Spark 1.1 | 88.4% | 71.0% | estimated ± 7.8 pp, low confidence |
| Qwen3.5-122B-A10B | 77.2% | 65.0% | estimated ± 7.8 pp, low confidence |
| Sakana Fugu | 85.1% | 69.3% | estimated ± 7.8 pp, low confidence |
| Sakana Fugu-Ultra | 86.6% | 70.1% | estimated ± 7.8 pp, low confidence |