Method
How the gaps are filled, and when not to trust it
benchgap is an LLM benchmark leaderboard that fills in the missing scores. Most models are only ever run on a handful of benchmarks, so benchgap calibrates benchmarks against each other on the models measured on both, then estimates each missing score with its cross-validated error and a confidence level. Measured and estimated scores are always marked apart.
- Measured scores come from public evaluation leaderboards and are never altered; each benchmark is calibrated only against benchmarks of the same capability.
- For every ordered pair of benchmarks with at least 5 models measured on both, 6 monotone curves are fitted and the one with the lowest leave-one-out cross-validated error is kept. That error, in percentage points, is the ± shown with every estimate.
- A pair keeps no calibration unless its best curve reaches R² ≥ 0.3 and an error of at most 15 pp; such gaps stay empty.
- A missing score is estimated from the best calibration out of a benchmark the model was measured on. Estimates are never used to make further estimates.
- Confidence: high for an error up to 5 pp, medium up to 10 pp, low above; extrapolation, a fit on fewer than 8 models or R² below 0.5 each lower it by one level.
- Estimates are predictions, not measurements, and hold for the source leaderboard's evaluation setup only.