Pith. sign in

REVIEW 3 cited by

On the Benefits of Invariance in Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.00178 v1 pith:IUUVOERL submitted 2020-05-01 cs.LG stat.ML

classification cs.LGstat.ML
keywords dataaugmentationinvariancegeneralizationmodelstrainingaveragingbenefits
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many real world data analysis problems exhibit invariant structure, and models that take advantage of this structure have shown impressive empirical performance, particularly in deep learning. While the literature contains a variety of methods to incorporate invariance into models, theoretical understanding is poor and there is no way to assess when one method should be preferred over another. In this work, we analyze the benefits and limitations of two widely used approaches in deep learning in the presence of invariance: data augmentation and feature averaging. We prove that training with data augmentation leads to better estimates of risk and gradients thereof, and we provide a PAC-Bayes generalization bound for models trained with data augmentation. We also show that compared to data augmentation, feature averaging reduces generalization error when used with convex losses, and tightens PAC-Bayes bounds. We provide empirical support of these theoretical results, including a demonstration of why generalization may not improve by training with data augmentation: the `learned invariance' fails outside of the training distribution.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EquiFusion: Kinematics-Agnostic Human Motion Prediction via Equivariant Latent Diffusion

    cs.CV 2026-07 accept novelty 7.5 of 10

    A permutation-equivariant latent diffusion model treats skeleton connectivity as input, enabling the first kinematics-agnostic stochastic human motion predictor that generalizes zero-shot to unseen and partial skeletons.

  2. Symmetries in PAC-Bayesian Learning

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Averaging a model over a symmetry group with a data-dependent kernel shrinks the KL term in PAC-Bayes bounds, giving tighter guarantees for non-compact groups and non-invariant data.

  3. PAC--Bayes Bounds on Quotient Parameter Spaces: Geometry-induced Implicit-Bias Priors

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Identifying parameters that define the same predictor reduces the PAC-Bayes KL complexity term, and a geometry-tilted 'implicit-bias' prior can tighten the certificate when it is closer to the learned posterior.

Pith tools