Pith. sign in

REVIEW 2 cited by

Do Residual Neural Networks discretize Neural Ordinary Differential Equations?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.14612 v2 pith:WIY5ASDW submitted 2022-05-29 cs.LG stat.ML

classification cs.LGstat.ML
keywords neuralmethodresidualdepthadjointfunctionsresnetcontinuous
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural Ordinary Differential Equations (Neural ODEs) are the continuous analog of Residual Neural Networks (ResNets). We investigate whether the discrete dynamics defined by a ResNet are close to the continuous one of a Neural ODE. We first quantify the distance between the ResNet's hidden state trajectory and the solution of its corresponding Neural ODE. Our bound is tight and, on the negative side, does not go to 0 with depth N if the residual functions are not smooth with depth. On the positive side, we show that this smoothness is preserved by gradient descent for a ResNet with linear residual functions and small enough initial loss. It ensures an implicit regularization towards a limit Neural ODE at rate 1 over N, uniformly with depth and optimization time. As a byproduct of our analysis, we consider the use of a memory-free discrete adjoint method to train a ResNet by recovering the activations on the fly through a backward pass of the network, and show that this method theoretically succeeds at large depth if the residual functions are Lipschitz with the input. We then show that Heun's method, a second order ODE integration scheme, allows for better gradient estimation with the adjoint method when the residual functions are smooth with depth. We experimentally validate that our adjoint method succeeds at large depth, and that Heun method needs fewer layers to succeed. We finally use the adjoint method successfully for fine-tuning very deep ResNets without memory consumption in the residual layers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Layers to States: A State Space Model Perspective to Deep Neural Network Layer Dynamics

    cs.LG 2025-02 conditional novelty 6.0 of 10

    S6LA adds a selective state space recurrence across the layers of CNNs and vision transformers, giving consistent accuracy gains on ImageNet classification and COCO detection and segmentation.

  2. Weight-Parameterization in Continuous Time Deep Neural Networks for Surrogate Modeling

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Legendre-polynomial weight parameterization lowers training cost and improves stability in continuous-time network surrogates, but the reported accuracy advantage conflicts with the paper's own error table.

Pith tools