Pith. sign in

REVIEW 1 cited by

Implicit regularization of deep residual networks towards neural ODEs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.01213 v3 pith:MGWIPGPK submitted 2023-09-03 stat.ML cs.LG

classification stat.MLcs.LG
keywords networksneuralresidualdeepodestrainingconditiondiscretization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Residual neural networks are state-of-the-art deep learning models. Their continuous-depth analog, neural ordinary differential equations (ODEs), are also widely used. Despite their success, the link between the discrete and continuous models still lacks a solid mathematical foundation. In this article, we take a step in this direction by establishing an implicit regularization of deep residual networks towards neural ODEs, for nonlinear networks trained with gradient flow. We prove that if the network is initialized as a discretization of a neural ODE, then such a discretization holds throughout training. Our results are valid for a finite training time, and also as the training time tends to infinity provided that the network satisfies a Polyak-Lojasiewicz condition. Importantly, this condition holds for a family of residual networks where the residuals are two-layer perceptrons with an overparameterization in width that is only linear, and implies the convergence of gradient flow to a global minimum. Numerical experiments illustrate our results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transformative or Conservative? Conservation laws for ResNets and Transformers

    cs.LG 2025-06 conditional novelty 7.0 of 10

    Conservation laws for gradient-flow training of conv ResNets and Transformers are characterized for several building blocks, and deep network block laws reduce to laws of isolated blocks.

Pith tools