Calibration
Artificial Analysis Intelligence Index → BioMysteryBench (human-solvable)
BioMysteryBench (human-solvable) is estimated from Artificial Analysis Intelligence Index with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (0.8988 − 0.0000)·x^6.00 / (0.21204^6.00 + x^6.00), R² = 0.70, cross-validated error 1.0 pp. It is used for 13 estimates.
| Estimated model | Artificial Analysis Intelligence Index | BioMysteryBench (human-solvable) | Source |
|---|---|---|---|
| Claude 3 Opus | 8.7% | 0.4% | estimated ± 1.0 pp, low confidence |
| Claude 4.1 Opus | 18.6% | 27.9% | estimated ± 1.0 pp, low confidence |
| DeepSeek R1 Distill Qwen 32B | 8.4% | 0.3% | estimated ± 1.0 pp, low confidence |
| Gemini 1.0 Pro | 5.3% | 0.0% | estimated ± 1.0 pp, low confidence |
| Gemini 1.5 Pro | 7.9% | 0.2% | estimated ± 1.0 pp, low confidence |
| GPT-4 Turbo | 7.0% | 0.1% | estimated ± 1.0 pp, low confidence |
| GPT-4o mini | 6.7% | 0.1% | estimated ± 1.0 pp, low confidence |
| o1-preview | 11.4% | 2.1% | estimated ± 1.0 pp, low confidence |
| o1-pro | 12.4% | 3.5% | estimated ± 1.0 pp, low confidence |
| o3-mini | 12.5% | 3.6% | estimated ± 1.0 pp, low confidence |
| o3-pro | 21.9% | 49.1% | estimated ± 1.0 pp, low confidence |
| Phi-4 Multimodal Instruct | 5.8% | 0.0% | estimated ± 1.0 pp, low confidence |
| Qwen2.5 Coder 32B Instruct | 6.7% | 0.1% | estimated ± 1.0 pp, low confidence |