Calibration
Gert Labs → AA Agentic Index
AA Agentic Index is estimated from Gert Labs with a inverse Michaelis–Menten curve fitted on 18 models measured on both: y = 0.18310·(x − 0.0000) / (0.0000 + 1.0746 − x), R² = 0.72, cross-validated error 6.4 pp. It is used for 17 estimates.
| Estimated model | Gert Labs | AA Agentic Index | Source |
|---|---|---|---|
| Claude 4 Sonnet | 39.7% | 10.7% | estimated ± 6.4 pp, medium confidence |
| Claude Opus 4.7 | 65.6% | 28.7% | estimated ± 6.4 pp, medium confidence |
| Claude Sonnet 4.5 | 48.5% | 15.1% | estimated ± 6.4 pp, medium confidence |
| Gemini 3.1 Flash-Lite | 38.5% | 10.2% | estimated ± 6.4 pp, medium confidence |
| Gemini 3 Pro | 63.2% | 26.2% | estimated ± 6.4 pp, medium confidence |
| GPT-4.1 | 25.7% | 5.7% | estimated ± 6.4 pp, low confidence |
| GPT-5.1-Codex | 49.7% | 15.7% | estimated ± 6.4 pp, medium confidence |
| GPT-5.2 | 46.5% | 14.0% | estimated ± 6.4 pp, medium confidence |
| GPT-5.2-Codex | 51.8% | 17.0% | estimated ± 6.4 pp, medium confidence |
| GPT-5.3 Codex | 57.5% | 21.1% | estimated ± 6.4 pp, medium confidence |
| Grok 4 | 42.3% | 11.9% | estimated ± 6.4 pp, medium confidence |
| Grok 4.1 Fast | 47.3% | 14.4% | estimated ± 6.4 pp, medium confidence |
| Grok 4.20 | 38.4% | 10.2% | estimated ± 6.4 pp, medium confidence |
| Grok Build 0.1 | 49.2% | 15.4% | estimated ± 6.4 pp, medium confidence |
| Qwen3.5-27B | 39.4% | 10.6% | estimated ± 6.4 pp, medium confidence |
| Qwen3.5-35B-A3B | 29.0% | 6.8% | estimated ± 6.4 pp, low confidence |
| Qwen3 Max | 43.7% | 12.6% | estimated ± 6.4 pp, medium confidence |