benchgap
Publications

The research behind the gapfilling

benchgap is one entry in an active research line: predicting a model's benchmark scores without running every evaluation. These are the papers closest to what this site does, and how they relate to it.

PaperPredicts fromApproachHow benchgap relates
You Don't Need to Run Every Eval
Zeng & Papailiopoulos, 2026
a model's other benchmark scoresthe model × benchmark score matrix is nearly rank-2; matrix completion in logit space (BenchPress)the closest relative: the same gapfilling problem with one global factor model. benchgap keeps local, explicit calibrations per benchmark pair, each with its own cross-validated error, and refuses to transfer across capabilities.
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
Polo et al., 2024
training compute and latent skillsscaling laws over low-dimensional skill factors, within and across model familiespredicts hypothetical models and needs training metadata. benchgap maps an existing model from its measured scores alone, which is all closed API models publish.
Observational Scaling Laws and the Predictability of Language Model Performance
Ruan et al., 2024
simple benchmarks and computea latent capability variable regressed onto downstream benchmarksthe same score-from-scores idea, anchored to compute. benchgap is compute-agnostic, so it also works for models whose training details are unknown.
From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation
Maimon et al., 2025
a subset of a model's scorespsychometric low-rank factorization; profiling a model from a few tasksclosest in the fill-the-profile goal, but in latent space. benchgap stays in observable benchmark space and shows the fitted curve for every pair.
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
Bowyer et al., 2026
a small coreset of benchmark itemsfeature selection plus regression to predict full-benchmark scoresthe item-level analogue of the multivariate view's elastic net, whose lasso part selects the useful benchmarks.
metabench: A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Kipnis et al., 2024
a sparse (~3%) subset of itemsitem-level distillation that preserves scores and rankingsitem-level. benchgap works from published aggregate scores, so it needs no access to benchmark items at all.
Look Before you Leap: Estimating LLM Benchmark Scores from Descriptions
Park et al., 2025
a redacted text description of the taskan LLM as the regressor (the PRECOG corpus); no evaluation runs at allpredicts before any evaluation exists; benchgap predicts after a model has some measured scores. Complementary ends of the pipeline.
How predictable is language model benchmark performance?
Owen, 2024
training computeempirical analysis of benchmark performance across five orders of magnitude of computea different input: predictability against compute, not scores from scores.
How Benchmark Prediction from Fewer Data Misses the Mark
Zhang et al., 2025
—a systematic evaluation of 11 score-prediction methods across 19 benchmarksthe caution this site's guardrails are built around: predictors fail on models unlike their calibration set.
PredictaBoard: Benchmarking LLM Score Predictability
Pacchiardi et al., 2025
—benchmarks score predictability itself, via assessors that anticipate a model's errorsinstance-level predictability rather than score-level estimation; a complementary lens on the same uncertainty.

The multivariate predictions

Predicting one benchmark from several others jointly is where benchgap meets the papers above head-on: BenchPress, Sloth and From Benchmarks to Skills do the same thing through latent factors over the whole score matrix. The multivariate view does it with explicit features - the measured benchmarks themselves, named in every fit - selected per target benchmark.

The selection echoes Efficient Benchmarking Is Just Feature Selection and Multiple Regression, one level up: they select items, benchgap selects benchmarks. An elastic net whose lasso part zeroes the useless candidates runs alongside a greedy forward search trying every one; on the features they find, a linear fit and a multivariate Michaelis–Menten curve compete by cross-validated error.

What a joint model has and a per-target fit does not is strength borrowed across all benchmarks at once - BenchPress finds most of the score matrix is two numbers per model. benchgap trades that for fits a reader can check: every feature is a real benchmark, and every fit carries its own cross-validated error and a measured-vs-predicted scatter. These mappings are the deterministic precursor of that joint model: a Bayesian network over benchmark scores, imputing every gap with one coherent posterior, is where the roadmap points.

Where benchgap differs

  • Capability gating. A factor model imputes between any two benchmarks. benchgap calibrates only within a capability group, so a model never evaluated on vision keeps that gap instead of inheriting an estimate from text benchmarks.
  • No estimate recursion. Every input to an estimate is a measured score; an estimate never feeds another estimate. A factor model completes a matrix that already contains its own outputs.
  • Per-cell error. Each estimate carries its own leave-one-out error and confidence level, and a pair whose best curve still fits poorly keeps no mapping at all. The papers above report one aggregate error over held-out cells.
  • The shared limit. As How Benchmark Prediction from Fewer Data Misses the Mark shows, every method in this line misestimates models unlike its calibration set. Confidence levels flag the known risk factors - extrapolation, small fits, weak R² - but nothing here detects a genuinely novel model.