benchgap
Calibration

Artificial Analysis Intelligence Index → BioMysteryBench (human-solvable)

BioMysteryBench (human-solvable) is estimated from Artificial Analysis Intelligence Index with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (0.8988 − 0.0000)·x^6.00 / (0.21204^6.00 + x^6.00), R² = 0.70, cross-validated error 1.0 pp. It is used for 13 estimates.

Estimated modelArtificial Analysis Intelligence IndexBioMysteryBench (human-solvable)Source
Claude 3 Opus8.7%0.4%estimated ± 1.0 pp, low confidence
Claude 4.1 Opus18.6%27.9%estimated ± 1.0 pp, low confidence
DeepSeek R1 Distill Qwen 32B8.4%0.3%estimated ± 1.0 pp, low confidence
Gemini 1.0 Pro5.3%0.0%estimated ± 1.0 pp, low confidence
Gemini 1.5 Pro7.9%0.2%estimated ± 1.0 pp, low confidence
GPT-4 Turbo7.0%0.1%estimated ± 1.0 pp, low confidence
GPT-4o mini6.7%0.1%estimated ± 1.0 pp, low confidence
o1-preview11.4%2.1%estimated ± 1.0 pp, low confidence
o1-pro12.4%3.5%estimated ± 1.0 pp, low confidence
o3-mini12.5%3.6%estimated ± 1.0 pp, low confidence
o3-pro21.9%49.1%estimated ± 1.0 pp, low confidence
Phi-4 Multimodal Instruct5.8%0.0%estimated ± 1.0 pp, low confidence
Qwen2.5 Coder 32B Instruct6.7%0.1%estimated ± 1.0 pp, low confidence