You Don't Need to Run Every Eval Zeng & Papailiopoulos, 2026 | a model's other benchmark scores | the model × benchmark score matrix is nearly rank-2; matrix completion in logit space (BenchPress) | the closest relative: the same gapfilling problem with one global factor model. benchgap keeps local, explicit calibrations per benchmark pair, each with its own cross-validated error, and refuses to transfer across capabilities. |
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families Polo et al., 2024 | training compute and latent skills | scaling laws over low-dimensional skill factors, within and across model families | predicts hypothetical models and needs training metadata. benchgap maps an existing model from its measured scores alone, which is all closed API models publish. |
Observational Scaling Laws and the Predictability of Language Model Performance Ruan et al., 2024 | simple benchmarks and compute | a latent capability variable regressed onto downstream benchmarks | the same score-from-scores idea, anchored to compute. benchgap is compute-agnostic, so it also works for models whose training details are unknown. |
From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation Maimon et al., 2025 | a subset of a model's scores | psychometric low-rank factorization; profiling a model from a few tasks | closest in the fill-the-profile goal, but in latent space. benchgap stays in observable benchmark space and shows the fitted curve for every pair. |
Efficient Benchmarking Is Just Feature Selection and Multiple Regression Bowyer et al., 2026 | a small coreset of benchmark items | feature selection plus regression to predict full-benchmark scores | the item-level analogue of the multivariate view's elastic net, whose lasso part selects the useful benchmarks. |
metabench: A Sparse Benchmark of Reasoning and Knowledge in Large Language Models Kipnis et al., 2024 | a sparse (~3%) subset of items | item-level distillation that preserves scores and rankings | item-level. benchgap works from published aggregate scores, so it needs no access to benchmark items at all. |
Look Before you Leap: Estimating LLM Benchmark Scores from Descriptions Park et al., 2025 | a redacted text description of the task | an LLM as the regressor (the PRECOG corpus); no evaluation runs at all | predicts before any evaluation exists; benchgap predicts after a model has some measured scores. Complementary ends of the pipeline. |
How predictable is language model benchmark performance? Owen, 2024 | training compute | empirical analysis of benchmark performance across five orders of magnitude of compute | a different input: predictability against compute, not scores from scores. |
How Benchmark Prediction from Fewer Data Misses the Mark Zhang et al., 2025 | — | a systematic evaluation of 11 score-prediction methods across 19 benchmarks | the caution this site's guardrails are built around: predictors fail on models unlike their calibration set. |
PredictaBoard: Benchmarking LLM Score Predictability Pacchiardi et al., 2025 | — | benchmarks score predictability itself, via assessors that anticipate a model's errors | instance-level predictability rather than score-level estimation; a complementary lens on the same uncertainty. |