Pith. sign in

REVIEW 2 cited by

Understanding the role of importance weighting for deep learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.15209 v1 pith:QKGFVDFW submitted 2021-03-28 cs.LG

classification cs.LG
keywords importancelearningweightingdeepimpactmodelmodelsrole
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The recent paper by Byrd & Lipton (2019), based on empirical observations, raises a major concern on the impact of importance weighting for the over-parameterized deep learning models. They observe that as long as the model can separate the training data, the impact of importance weighting diminishes as the training proceeds. Nevertheless, there lacks a rigorous characterization of this phenomenon. In this paper, we provide formal characterizations and theoretical justifications on the role of importance weighting with respect to the implicit bias of gradient descent and margin-based learning theory. We reveal both the optimization dynamics and generalization performance under deep learning models. Our work not only explains the various novel phenomenons observed for importance weighting in deep learning, but also extends to the studies where the weights are being optimized as part of the model, which applies to a number of topics under active research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Thumb on the Scale: Optimal Loss Weighting in Last Layer Retraining

    cs.LG 2025-06 conditional novelty 6.0 of 10

    For square-loss weighted ERM in the proportional asymptotic regime, the class weight that equalizes per-class errors is ρ̃ = π−/π+ + (π−/π+ − 1) δ/(2π+ − δ), exceeding the ratio of priors and growing with δ.

  2. CaTE Data Curation for Trustworthy AI

    cs.LG 2025-08 accept novelty 4.0 of 10

    A synthesis of data curation practices for trustworthy AI, framed around an actionable definition of trustworthiness and a decision tree.

Pith tools