Pith. sign in

REVIEW 3 cited by

Optimizing ML Training with Metagradient Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.13751 v1 pith:ZEHUMQBG submitted 2025-03-17 stat.ML cs.AIcs.LG

classification stat.MLcs.AIcs.LG
keywords trainingmodeldescentintroducelearningmetagradientmetagradientsaccuracy-degrading
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A major challenge in training large-scale machine learning models is configuring the training process to maximize model performance, i.e., finding the best training setup from a vast design space. In this work, we unlock a gradient-based approach to this problem. We first introduce an algorithm for efficiently calculating metagradients -- gradients through model training -- at scale. We then introduce a "smooth model training" framework that enables effective optimization using metagradients. With metagradient descent (MGD), we greatly improve on existing dataset selection methods, outperform accuracy-degrading data poisoning attacks by an order of magnitude, and automatically find competitive learning rate schedules.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. (A)iSpy: Parasitic Trojans for Machine Learning Infrastructure

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A runtime-extension Trojan turns a single poisoned sample into a 97%+ backdoor via replay/amplification and leaks training hyperparameters through watermarked weights or innocuous text codewords.

  2. Ambient Diffusion Omni: Training Good Models with Bad Data

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Ambient Diffusion Omni trains diffusion models on mixed-quality data by learning when corrupted images can be treated as clean, improving generation quality and diversity.

  3. Rescaled Influence Functions: Accurate Data Attribution in High Dimension

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Rescaled influence functions, which account for the change in the Hessian when a sample is removed, dramatically improve leave-T-out effect estimates compared to standard influence functions in high-dimensional logist...

Pith tools