benchgap
OpenAI · model

GPT-5.5 (xhigh) benchmark scores

As of 2026-10-07, GPT-5.5 (xhigh) (OpenAI) has measured scores on 3 benchmarks and estimated scores on 4 more.

BenchmarkScoreSource
τ³-Bench Banking47.0%estimated ± 2.6 pp, medium confidence
AA-AnalystAgent50.0%measured
Harvey LAB93.4%estimated ± 1.6 pp, low confidence
AutomationBench61.2%estimated ± 11.2 pp, low confidence
EnterpriseOps-Gym51.6%estimated ± 1.8 pp, low confidence
APEX-Agents37.7%measured
ITBench SRE45.8%measured