benchgap
Calibration

Gert Labs → AA Agentic Index

AA Agentic Index is estimated from Gert Labs with a inverse Michaelis–Menten curve fitted on 18 models measured on both: y = 0.18310·(x − 0.0000) / (0.0000 + 1.0746 − x), R² = 0.72, cross-validated error 6.4 pp. It is used for 17 estimates.

Estimated modelGert LabsAA Agentic IndexSource
Claude 4 Sonnet39.7%10.7%estimated ± 6.4 pp, medium confidence
Claude Opus 4.765.6%28.7%estimated ± 6.4 pp, medium confidence
Claude Sonnet 4.548.5%15.1%estimated ± 6.4 pp, medium confidence
Gemini 3.1 Flash-Lite38.5%10.2%estimated ± 6.4 pp, medium confidence
Gemini 3 Pro63.2%26.2%estimated ± 6.4 pp, medium confidence
GPT-4.125.7%5.7%estimated ± 6.4 pp, low confidence
GPT-5.1-Codex49.7%15.7%estimated ± 6.4 pp, medium confidence
GPT-5.246.5%14.0%estimated ± 6.4 pp, medium confidence
GPT-5.2-Codex51.8%17.0%estimated ± 6.4 pp, medium confidence
GPT-5.3 Codex57.5%21.1%estimated ± 6.4 pp, medium confidence
Grok 442.3%11.9%estimated ± 6.4 pp, medium confidence
Grok 4.1 Fast47.3%14.4%estimated ± 6.4 pp, medium confidence
Grok 4.2038.4%10.2%estimated ± 6.4 pp, medium confidence
Grok Build 0.149.2%15.4%estimated ± 6.4 pp, medium confidence
Qwen3.5-27B39.4%10.6%estimated ± 6.4 pp, medium confidence
Qwen3.5-35B-A3B29.0%6.8%estimated ± 6.4 pp, low confidence
Qwen3 Max43.7%12.6%estimated ± 6.4 pp, medium confidence