Pith. sign in

REVIEW 2 cited by

On the Generalization for Transfer Learning: An Information-Theoretic Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.05377 v2 pith:T6ENK36P submitted 2022-07-12 cs.IT cs.LGmath.IT

classification cs.ITcs.LGmath.IT
keywords boundsdatalearningalgorithmriskalgorithmsanalysisexcess
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Transfer learning, or domain adaptation, is concerned with machine learning problems in which training and testing data come from possibly different probability distributions. In this work, we give an information-theoretic analysis of the generalization error and excess risk of transfer learning algorithms. Our results suggest, perhaps as expected, that the Kullback-Leibler (KL) divergence $D(\mu\|\mu')$ plays an important role in the characterizations where $\mu$ and $\mu'$ denote the distribution of the training data and the testing data, respectively. Specifically, we provide generalization error and excess risk upper bounds for learning algorithms where data from both distributions are available in the training phase. Recognizing that the bounds could be sub-optimal in general, we provide improved excess risk upper bounds for a certain class of algorithms, including the empirical risk minimization (ERM) algorithm, by making stronger assumptions through the \textit{central condition}. To demonstrate the usefulness of the bounds, we further extend the analysis to the Gibbs algorithm and the noisy stochastic gradient descent method. We then generalize the mutual information bound with other divergences such as $\phi$-divergence and Wasserstein distance, which may lead to tighter bounds and can handle the case when $\mu$ is not absolutely continuous with respect to $\mu'$. Several numerical results are provided to demonstrate our theoretical findings. Lastly, to address the problem that the bounds are often not directly applicable in practice due to the absence of the distributional knowledge of the data, we develop an algorithm (called InfoBoost) that dynamically adjusts the importance weights for both source and target data based on certain information measures. The empirical results show the effectiveness of the proposed algorithm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Subgroups Matter for Robust Bias Mitigation

    cs.LG 2025-05 accept novelty 7.0 of 10

    Subgroup choice determines whether bias mitigation helps or hurts, and the minimum KL divergence to the unbiased test distribution predicts success.

  2. Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis

    cs.LG 2025-06 reject novelty 5.0 of 10

    The authors derive information-theoretic generalization bounds for VAEs and diffusion models that expose a trade-off in the diffusion time T, and propose using the computable bound to select T and regularize training.

Pith tools