benchgap
Method

How the gaps are filled, and when not to trust it

benchgap is an LLM benchmark leaderboard that fills in the missing scores. Most models are only ever run on a handful of benchmarks, so benchgap calibrates benchmarks against each other on the models measured on both, then estimates each missing score with its cross-validated error and a confidence level. Measured and estimated scores are always marked apart.