Calibration
Which benchmarks predict which
For each ordered pair of same-capability benchmarks with at least 5 shared models, several monotone curves are fitted and the one with the lowest leave-one-out error is kept, if it passes the quality gate (R² ≥ 0.3, error ≤ 15 pp).