Pith. sign in

REVIEW 4 cited by

Continuous-in-Depth Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.02389 v1 pith:IEGEXXNP submitted 2020-08-05 cs.LG math.DSstat.ML

classification cs.LGmath.DSstat.ML
keywords dynamicalcontinuous-in-depthgraphricherschemessystemschangescomputational
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has attempted to interpret residual networks (ResNets) as one step of a forward Euler discretization of an ordinary differential equation, focusing mainly on syntactic algebraic similarities between the two systems. Discrete dynamical integrators of continuous dynamical systems, however, have a much richer structure. We first show that ResNets fail to be meaningful dynamical integrators in this richer sense. We then demonstrate that neural network models can learn to represent continuous dynamical systems, with this richer structure and properties, by embedding them into higher-order numerical integration schemes, such as the Runge Kutta schemes. Based on these insights, we introduce ContinuousNet as a continuous-in-depth generalization of ResNet architectures. ContinuousNets exhibit an invariance to the particular computational graph manifestation. That is, the continuous-in-depth model can be evaluated with different discrete time step sizes, which changes the number of layers, and different numerical integration schemes, which changes the graph connectivity. We show that this can be used to develop an incremental-in-depth training scheme that improves model quality, while significantly decreasing training time. We also show that, once trained, the number of units in the computational graph can even be decreased, for faster inference with little-to-no accuracy drop.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Recursive Bound-Constrained AdaGrad with Applications to Multilevel and Domain Decomposition Minimization

    math.OC 2025-07 conditional novelty 6.0 of 10

    Two noise-tolerant, bound-constrained AdaGrad variants for multilevel and domain-decomposition problems are proved to find an epsilon-approximate critical point in O(epsilon^-2) iterations with high probability.

  2. From Layers to States: A State Space Model Perspective to Deep Neural Network Layer Dynamics

    cs.LG 2025-02 conditional novelty 6.0 of 10

    S6LA adds a selective state space recurrence across the layers of CNNs and vision transformers, giving consistent accuracy gains on ImageNet classification and COCO detection and segmentation.

  3. A manifold-aware Neural ODE surrogate model for stochastic induction heating with anisotropic electrical conductivity

    cs.CE 2026-08 conditional novelty 5.0 of 10

    A manifold-aware neural ODE surrogate reproduces Monte Carlo temperature statistics of stochastic induction heating, with conductivity scaling uncertainty dominating and Neural ODE integrators improving run-to-run rep...

  4. Deep Learning with Self-Attention and Enhanced Preprocessing for Precise Diagnosis of Acute Lymphoblastic Leukemia from Bone Marrow Smears in Hemato-Oncology

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    A VGG19 network with added multi-head self-attention and focal loss reaches 99.25% accuracy for acute lymphoblastic leukemia diagnosis from bone marrow smears.

Pith tools