benchgap
Calibration

FrontierCode 1.1 Extended → VulcanBench v3

VulcanBench v3 is estimated from FrontierCode 1.1 Extended with a offset logistic curve fitted on 5 models measured on both: y = 0.0000 + (0.8707 − 0.0000) / (1 + exp(−200.00·(x − 0.5308))), R² = 0.94, cross-validated error 0.7 pp. It is used for 1 estimate.

Estimated modelFrontierCode 1.1 ExtendedVulcanBench v3Source
GPT-6 Astra64.5%87.1%estimated ± 0.7 pp, low confidence