REVIEW 4 cited by
Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Attempts from different disciplines to provide a fundamental understanding of deep learning have advanced rapidly in recent years, yet a unified framework remains relatively limited. In this article, we provide one possible way to align existing branches of deep learning theory through the lens of dynamical system and optimal control. By viewing deep neural networks as discrete-time nonlinear dynamical systems, we can analyze how information propagates through layers using mean field theory. When optimization algorithms are further recast as controllers, the ultimate goal of training processes can be formulated as an optimal control problem. In addition, we can reveal convergence and generalization properties by studying the stochastic dynamics of optimization algorithms. This viewpoint features a wide range of theoretical study from information bottleneck to statistical physics. It also provides a principled way for hyper-parameter tuning when optimal control theory is introduced. Our framework fits nicely with supervised learning and can be extended to other learning problems, such as Bayesian learning, adversarial training, and specific forms of meta learning, without efforts. The review aims to shed lights on the importance of dynamics and optimal control when developing deep learning theory.
Forward citations
Cited by 4 Pith papers
-
Generalization Bounds on Optimal Control for Transformer Training and Wasserstein Distributional Robustness
Finite-sample excess-risk bounds for DP-trained Transformers follow from Lipschitz value-function stability plus concentration of empirical laws on a triply quantized doubly-lifted model, with a matching Wasserstein D...
-
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
Smoothed Best-of-N has finite-sample KL and regret bounds under imperfect reward models, and tuning its temperature can make its regret bound beat hard Best-of-N in the overoptimization regime.
-
Adversarial Data Augmentation for Single Domain Generalization via Lyapunov Exponent-Guided Optimization
LEAwareSGD modulates the learning rate with a Lyapunov exponent estimate and claims state-of-the-art accuracy on three single-domain generalization benchmarks, but the method is underspecified and its hyperparameters ...
-
A Statistical Framework for Model Selection in LSTM Networks
A statistical model selection framework for LSTMs is proposed, but it largely recombines existing methods and provides no convincing evidence of improvement.
Discussion (0). Continue with ORCID to comment.