Pith. sign in

REVIEW 3 cited by

Generalization error of min-norm interpolators in transfer learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.13944 v2 pith:FF4XWSKT submitted 2024-06-20 math.ST cs.LGstat.MEstat.MLstat.TH

classification math.STcs.LGstat.MEstat.MLstat.TH
keywords learningshiftinterpolationdatamin-normmodelpooledresults
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This paper establishes the generalization error of pooled min-$\ell_2$-norm interpolation in transfer learning, where data from diverse distributions are available. Min-norm interpolators arise naturally as implicit regularized limits of modern machine learning algorithms. Prior work has characterized their out-of-distribution risk when samples from the test distribution are unavailable during training. In many applications, however, limited test samples may be available at training time, yet properties of min-norm interpolation in this regime remain poorly understood. We address this gap by characterizing the bias and variance of pooled min-$\ell_2$-norm interpolation under both covariate shift and model shift. Our results yield several important implications. In certain cases under model shift, we show that adding data always hurts when the signal-to-noise ratio (SNR) is low. At higher SNR levels, transfer learning is beneficial provided the shift-to-signal ratio falls below a threshold that we characterize explicitly. Under covariate shift, we find that when the source sample size is small relative to the dimension, greater heterogeneity between domains reduces risk, and vice versa. While our model shift results are initially established for Gaussian designs, we extend them to more general designs through a universality argument. To illustrate the broader applicability of our technical tools beyond interpolation learning, we characterize the risk of a bias-corrected estimator that uses the pooled interpolator as an initialization and corrects the resulting bias with target data. On the technical side, we develop a novel anisotropic local law and a Lindeberg-swapping argument, yielding tools that may be of independent interest in random matrix theory and universality analysis. Finally, we supplement our theory with simulations demonstrating the finite-sample efficacy of our results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression

    math.ST 2026-07 conditional novelty 7.0 of 10

    The zero pattern of a design's covariance—read as a bipartite matching problem—exactly locates the bias support and variance peaks that create multiple descent in ridge regression.

  2. Optimal Regularization for Performative Learning

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Optimal ridge regularization in performative linear regression is set by the mean performative strength, and over-parameterization can turn a self-reinforcing performative effect into a lower optimally tuned risk.

  3. Multi-Environment GLAMP: Approximate Message Passing for Transfer Learning with Applications to Lasso-based Estimators

    math.ST 2025-05 conditional novelty 6.0 of 10

    Multi-environment GLAMP yields exact asymptotic risk formulas for three Lasso-based transfer learning estimators under Gaussian designs, validated by simulations.

Pith tools