benchgap
Calibration

AA-MMMU-Pro → ScreenSpot Pro

ScreenSpot Pro is estimated from AA-MMMU-Pro with a Hill curve fitted on 12 models measured on both: y = 0.5565 + (1.2000 − 0.5565)·x^6.00 / (0.87742^6.00 + x^6.00), R² = 0.48, cross-validated error 10.7 pp. It is used for 60 estimates.

Estimated modelAA-MMMU-ProScreenSpot ProSource
Apodex 1.179.2%78.2%estimated ± 10.7 pp, low confidence
Apodex 1.1 Mini79.2%78.2%estimated ± 10.7 pp, low confidence
Claude 3 Haiku30.8%55.8%estimated ± 10.7 pp, low confidence
Claude 4.1 Opus Thinking67.9%67.0%estimated ± 10.7 pp, low confidence
Claude 4 Sonnet62.4%63.0%estimated ± 10.7 pp, low confidence
Claude Opus 4.5 Thinking74.0%72.7%estimated ± 10.7 pp, low confidence
Claude Opus 4.6 (Adaptive)75.4%74.1%estimated ± 10.7 pp, low confidence
Claude Opus 4.776.4%75.2%estimated ± 10.7 pp, low confidence
Claude Opus 584.7%84.4%estimated ± 10.7 pp, low confidence
Claude Opus 5.587.7%87.8%estimated ± 10.7 pp, low confidence
DeepSeek V4.1 Flash77.0%75.8%estimated ± 10.7 pp, low confidence
Gemini 1.5 Pro55.0%59.3%estimated ± 10.7 pp, low confidence
Gemini 2.5 Flash65.5%65.1%estimated ± 10.7 pp, low confidence
Gemini 2.5 Pro74.9%73.6%estimated ± 10.7 pp, low confidence
Gemini 3.5 Flash-Lite79.0%78.0%estimated ± 10.7 pp, low confidence
Gemini 3.6 Flash83.2%82.7%estimated ± 10.7 pp, low confidence
Gemini 3.8 Flash85.6%85.4%estimated ± 10.7 pp, low confidence
Gemini 3 Flash78.6%77.6%estimated ± 10.7 pp, low confidence
Gemma 3 27B48.0%57.3%estimated ± 10.7 pp, low confidence
Gemma 4 E2B44.6%56.7%estimated ± 10.7 pp, low confidence
Gemma 4 E4B51.4%58.1%estimated ± 10.7 pp, low confidence
GLM-5V-Turbo72.8%71.5%estimated ± 10.7 pp, low confidence
GPT-4.161.2%62.3%estimated ± 10.7 pp, low confidence
GPT-4.1 mini58.7%60.9%estimated ± 10.7 pp, low confidence
GPT-4.1 nano40.1%56.2%estimated ± 10.7 pp, low confidence
GPT-4o mini41.5%56.4%estimated ± 10.7 pp, low confidence
GPT-5.175.5%74.2%estimated ± 10.7 pp, low confidence
GPT-5.1-Codex72.5%71.2%estimated ± 10.7 pp, low confidence
GPT-5.1-Codex-Max72.5%71.2%estimated ± 10.7 pp, low confidence
GPT-5.2-Codex76.3%75.1%estimated ± 10.7 pp, low confidence
GPT-5.3 Codex78.5%77.5%estimated ± 10.7 pp, low confidence
GPT-5 (high)74.2%72.9%estimated ± 10.7 pp, low confidence
GPT-5 (medium)74.3%73.0%estimated ± 10.7 pp, low confidence
GPT-6.1 Sol86.0%85.9%estimated ± 10.7 pp, low confidence
GPT-6 Luna79.7%78.8%estimated ± 10.7 pp, low confidence
GPT-6 Sol82.9%82.4%estimated ± 10.7 pp, low confidence
Grok 468.8%67.8%estimated ± 10.7 pp, low confidence
Grok 4.1 Fast48.4%57.4%estimated ± 10.7 pp, low confidence
Grok 4.1 Fast (Reasoning)63.3%63.6%estimated ± 10.7 pp, low confidence
Grok 4.580.4%79.6%estimated ± 10.7 pp, low confidence
Grok 4 Fast (Reasoning)61.8%62.6%estimated ± 10.7 pp, low confidence
LFM2.5-VL-1.6B-Extract26.5%55.7%estimated ± 10.7 pp, low confidence
Ling 3.0 Flash VL79.0%78.0%estimated ± 10.7 pp, low confidence
Llama 4 Maverick62.1%62.8%estimated ± 10.7 pp, low confidence
Llama 4 Scout52.9%58.6%estimated ± 10.7 pp, low confidence
MiMo-V2.6-Flash73.1%71.8%estimated ± 10.7 pp, low confidence
MiMo-V2-Omni69.9%68.7%estimated ± 10.7 pp, low confidence
Mistral Large 355.7%59.6%estimated ± 10.7 pp, low confidence
Mistral Large 476.4%75.2%estimated ± 10.7 pp, low confidence
Mistral Medium 353.0%58.6%estimated ± 10.7 pp, low confidence
Mistral Medium 3.5 128B64.9%64.7%estimated ± 10.7 pp, low confidence
Mistral Small 456.8%60.1%estimated ± 10.7 pp, low confidence
Mistral Small 4 (Reasoning)56.8%60.1%estimated ± 10.7 pp, low confidence
Nova Pro44.3%56.7%estimated ± 10.7 pp, low confidence
o370.1%68.9%estimated ± 10.7 pp, low confidence
Phi-4 Multimodal Instruct14.5%55.6%estimated ± 10.7 pp, low confidence
Qwen3.5 397B (Reasoning)52.7%58.5%estimated ± 10.7 pp, low confidence
Qwen3.8 Max Preview82.8%82.3%estimated ± 10.7 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct55.5%59.5%estimated ± 10.7 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking60.2%61.7%estimated ± 10.7 pp, low confidence