REVIEW 2 cited by
Optimal Ridge Regularization for Out-of-Distribution Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We study the behavior of optimal ridge regularization and optimal ridge risk for out-of-distribution prediction, where the test distribution deviates arbitrarily from the train distribution. We establish general conditions that determine the sign of the optimal regularization level under covariate and regression shifts. These conditions capture the alignment between the covariance and signal structures in the train and test data and reveal stark differences compared to the in-distribution setting. For example, a negative regularization level can be optimal under covariate shift or regression shift, even when the training features are isotropic or the design is underparameterized. Furthermore, we prove that the optimally-tuned risk is monotonic in the data aspect ratio, even in the out-of-distribution setting and when optimizing over negative regularization levels. In general, our results do not make any modeling assumptions for the train or the test distributions, except for moment bounds, and allow for arbitrary shifts and the widest possible range of (negative) regularization levels.
Forward citations
Cited by 2 Pith papers
-
A Simple Approximation to the Distribution of the Ridge Regression Estimator
Under local-to-target asymptotics, √n times the ridge estimation error is approximately Gaussian with mean −λ(Σ+λI)^{-1}b and variance (Σ+λI)^{-1}Ω(Σ+λI)^{-1}, yielding closed-form tuning rules for isotropic features.
-
Multi-Environment GLAMP: Approximate Message Passing for Transfer Learning with Applications to Lasso-based Estimators
Multi-environment GLAMP yields exact asymptotic risk formulas for three Lasso-based transfer learning estimators under Gaussian designs, validated by simulations.
Discussion (0). Continue with ORCID to comment.