benchgap
Calibration

Artificial Analysis Intelligence Index → BioMysteryBench (human-difficult)

BioMysteryBench (human-difficult) is estimated from Artificial Analysis Intelligence Index with a offset logistic curve fitted on 6 models measured on both: y = 0.3523 + (0.5012 − 0.3523) / (1 + exp(−200.00·(x − 0.3889))), R² = 0.71, cross-validated error 8.1 pp. It is used for 13 estimates.

Estimated modelArtificial Analysis Intelligence IndexBioMysteryBench (human-difficult)Source
Claude 3 Opus8.7%35.2%estimated ± 8.1 pp, low confidence
Claude 4.1 Opus18.6%35.2%estimated ± 8.1 pp, low confidence
DeepSeek R1 Distill Qwen 32B8.4%35.2%estimated ± 8.1 pp, low confidence
Gemini 1.0 Pro5.3%35.2%estimated ± 8.1 pp, low confidence
Gemini 1.5 Pro7.9%35.2%estimated ± 8.1 pp, low confidence
GPT-4 Turbo7.0%35.2%estimated ± 8.1 pp, low confidence
GPT-4o mini6.7%35.2%estimated ± 8.1 pp, low confidence
o1-preview11.4%35.2%estimated ± 8.1 pp, low confidence
o1-pro12.4%35.2%estimated ± 8.1 pp, low confidence
o3-mini12.5%35.2%estimated ± 8.1 pp, low confidence
o3-pro21.9%35.2%estimated ± 8.1 pp, low confidence
Phi-4 Multimodal Instruct5.8%35.2%estimated ± 8.1 pp, low confidence
Qwen2.5 Coder 32B Instruct6.7%35.2%estimated ± 8.1 pp, low confidence