REVIEW 4 major objections 3 minor 54 references
This paper establishes that a residual block trains stably if and only if its velocity grows at most linearly with input magnitude, making q ≤ 1 the sharp threshold.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Residual blocks are claimed to be stable exactly when their velocity-field growth exponent satisfies q≤1, with composition rules for certifying q from architectural primitives.
T0 review reviewed 2026-08-02 challenge →
load-bearing objection The sufficiency half and the exponent arithmetic are real contributions, but the advertised necessity of q≤1 for stable training is not proven, and the paper's own q=5 survivors cut against it. the 4 major comments →
Sharp Stability Threshold and Certification for Designing Stable Residual Architectures
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that the input-magnitude exponent q — the growth rate of a residual block's output norm as a function of its input norm, ∥v(x,t)∥ ≤ c∥x∥^q + b — is the variable that controls training stability, and that q = 1 is the exact threshold. At q ≤ 1 the forward dynamics admit a global trajectory over the whole depth interval and the discrete residual stack inherits explicit forward-state and backward-Jacobian bounds; at any q > 1 there exist admissible velocity fields whose trajectories escape to infinity in finite time, as shown by the velocity v(x) = c∥x∥^{q−1}x with blow-up time t* = ∥X0∥^{1−q}/(c(q−1)). The optimal-control argument closes the loop that pure ODE theo
What carries the argument
The central object is the sublinear-growth class U^q_ad, the set of velocity fields satisfying ∥v(x,t)∥ ≤ c∥x∥^q + b, together with the input-magnitude exponent q(B), the infimum growth rate of a primitive. The carrying identity is the closed-form Hamiltonian maximizer H_q(X,P) = (c∥X∥^q + b)∥P∥ with bang-bang optimal velocity v* = −(c∥X∥^q + b)P/∥P∥, which pins the training optimum to the boundary of the admissible set; the companion facts are the Grönwall/comparison bounds giving global existence at q ≤ 1 and the explicit blow-up time at q > 1. Theorem 4.2 supplies the design arithmetic: sequential composition multiplies exponents, parallel addition takes the maximum, Hadamard product adds
Load-bearing premise
The necessity direction assumes real training is well modeled by the optimal-control problem in which the optimizer maximizes the Hamiltonian pointwise and drives velocity to the boundary of the admissible class; if SGD or Adam never saturate that boundary, the q > 1 blow-up need not occur in practice.
What would settle it
Take the native Mamba free-velocity block (q = 5) without normalization, initialize weights very small and apply strong weight decay so the effective coefficient c in ∥v(x)∥ ≤ c∥x∥^5 stays well below the blow-up threshold through training, at a depth D large enough that the admissible-ceiling per-step amplification would overflow the numerical range; if it converges to a finite validation loss, the claimed necessity of q ≤ 1 is falsified in that regime.
If this is right
- Every residual block in a trainable architecture must have q ≤ 1; any block with q > 1 carries a guaranteed risk of forward overflow that no initialization, optimizer, or learning-rate schedule can remove.
- Layer normalization's stabilizing role is explained as an exponent collapse to q = 0, but normalization is not the unique stabilizer: structural modifications that reach q = 1 are equally safe, as the Mamba linear-growth variant demonstrates.
- Block-level stability can be certified from primitive-level exponents using the five composition rules, making architectural search a matter of arithmetic rather than trial and error.
- The depth capacity of a q = 1 stack is limited by the exponential state bound (1 + cΔt)^D, which exceeds FP16 range at D ≈ 17 and FP32 range at D ≈ 128 for c = 1, so depth and exponent interact directly.
- A strict interior exponent q ∈ (0,1) would give polynomial-in-depth forward bounds, and the paper identifies the search for such a primitive as an open problem.
Where Pith is reading between the lines
- Because the optimal-control theorem is worst-case, a practical q > 1 model that never reaches the admissible-class ceiling might still train; the guarantee bites when training drives velocities toward the boundary, so the criterion should be read as a robust design rule rather than a universal empirical prediction for all q > 1 runs.
- The exponent arithmetic is a natural static-analysis tool: a linter could parse a model's computational graph, assign exponents from the primitive catalogue, and flag supercritical blocks before any training run, an engineering extension the paper sketches but does not implement.
- The empty q ∈ (0,1) slot suggests looking for new primitives such as sublinear activations or projections that interpolate between boundedness and magnitude preservation; if found, they would inherit polynomial depth bounds.
- A direct test of the necessity direction is to train a q > 1 block with strong regularization keeping the effective coefficient c tiny; stable convergence would not refute the theorem but would mark the regime where the idealized optimal-control picture ceases to describe actual optimizers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a sublinear-growth principle for residual architectures: every residual block's velocity field should satisfy ||v(x,t)|| ≤ c||x||^q + b with q ∈ [0,1]. It claims that q = 1 is a sharp necessary-and-sufficient threshold for stable training, established via ODE non-explosion and an HJB/optimal-control argument. It develops an exponent arithmetic for five primitive operations, a certification procedure, and a parameter-free modification of Mamba reducing q from 5 to 1. Experiments on Mamba and PatchTST report blow-up counts and forecasting MSE for q≤1 versus q>1 variants.
Significance. The sufficiency side is sound and useful: q≤1 yields global forward existence and explicit forward/backward bounds (Theorems 3.5 and 3.7, with the technical caveat noted below). The exponent arithmetic (Theorem 4.2, Table 1) is a practical design tool, and the Mamba modification is concrete, parameter-free, and sensible. The experiments usefully dissociate normalization from exponent in the tested regime. However, the paper's central 'necessary' claim is not established and is contradicted by its own Table 2 in the shallow setting. As a sufficient certification criterion the contribution is valuable; as a sharp necessary-and-sufficient threshold it is not.
major comments (4)
- [§3.1, after Eq. (17)] The step from bang-bang saturation ||v*|| = c||X||^q + b to 'the optimum selects a blow-up field' is asserted, not proved. Saturation fixes only the control norm; the direction is -∇u/||∇u||. For terminal cost G = (X(T)-y)^2, ∇u points outward, so v* points inward and the trajectory is bounded even at q>1; the outward field c||X||^{q-1}X of §2.1 is not the OC-optimal direction. No theorem in §3 derives the value function or shows the OC trajectory blows up. The claimed necessity of q≤1 is therefore unsupported.
- [§3.1, Proposition 3.3] The proof that H_q violates (20) at q>1 is used to conclude the HJB equation is not well-posed. Failure of one sufficient linear-growth condition does not imply non-well-posedness; superlinear-in-X Hamiltonians can still have unique viscosity solutions under other hypotheses (e.g., quadratic Hamiltonians in P). The proposition establishes only that a particular sufficient route fails, not that q≤1 is necessary for well-posedness.
- [§5, Table 2] The paper's own data contradict the abstract's necessary-and-sufficient claim under the standard reading that stable training means no overflow: the Mamba free-velocity q=5 block survives 12/24 ETTm1 runs and 10/12 Weather runs at L=3. If q>1 is compatible with stable training in a substantial fraction of runs, the threshold is not necessary. At most, q>1 is a risk factor; the deterministic 'necessary and sufficient' wording is unsupported.
- [§3.1, Eqs. (7)-(9)] The OC analysis optimizes over the full admissible class U_q^ad, but training optimizes over parameters θ realizing a restricted family of velocity fields via SGD/Adam with finite steps and stochastic batches. It is not established that the relaxed OC optimum v* lies in the parameterized family or that training tracks the HJB value function. The inference from 'the OC optimum is bang-bang' to 'training cannot dodge the boundary' is a modeling assumption and is load-bearing for the necessity direction.
minor comments (3)
- [§2.1] The statement that continuous-time analysis is conservative for the discrete recursion is only shown for the specific field v = c||x||^{q-1}x; it is not a general principle. The finite-depth divergence claim also depends on cΔt and the depth; please state it as a heuristic or provide a threshold.
- [Theorem 4.2(i)] The exponent rule for sequential composition is stated with infimum exponents, but the proof is asymptotic. A precise uniform bound requires retaining the bias terms of both primitives; please clarify the exact bound for all x.
- [Table 1] For self-attention, the q=1 entry is stated in a worst-token sense, with the Frobenius norm giving a possibly looser coefficient. This dependence on sequence length should be stated explicitly in the caption or text to avoid misreading.
Circularity Check
No significant circularity: the q=1 threshold is derived from classical ODE existence theory and an independent HJB analysis; experimental predictions are not fitted inputs.
full rationale
The central claim, that q<=1 is a stability threshold, is not circular. Section 2.1 derives the threshold from standard ODE theory: global existence at q<=1 follows from the Grönwall/Peano comparison argument, and finite-time blow-up at q>1 is exhibited by the explicit field v(x)=c||x||^(q-1)x. These are external mathematical facts, not the paper's conclusion repackaged. Section 3.1 adds an optimal-control analysis: Theorem 3.1 derives the bang-bang form of the optimal velocity by a direct Cauchy-Schwarz maximization on the admissible class U^q_ad, so the saturation identity ||v*||=c||X||^q+b is a mathematical consequence of the definition of the class, not an input assumption. The further claim that the q>1 optimum 'selects' a blow-up trajectory is argued rather than fully proved, but that is a correctness/validity gap, not a circular reduction: the paper does not assume the conclusion in deriving the theorem. Theorem 3.5/3.7 provide forward/backward bounds at q<=1 derived from the same growth bound, again by direct Grönwall/Danskin arguments. The arithmetic of exponents in Section 4 and the q-values in Table 1 are definitions and upper-bound calculations, not fitted parameters. The Mamba q=5 exponent is computed from the architecture's primitives via Theorem 4.2 and then tested empirically; the experiments therefore provide an external check rather than a relabeled fit. The only self-citation is [30] for the elementary boundedness of LayerNorm (|LN(z)|<=|beta|+||gamma|| sqrt(d)), which is parameter-free under the stated epsilon-regularized form and independently verifiable; it is not load-bearing in a way that closes the argument by authority. Overall, no step reduces the predicted stability criterion to its own inputs by construction.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption A residual network is a forward-Euler discretization of the ODE dX/dt=v(X,t) with the same velocity field and uniform-in-depth admissible constants c,b.
- domain assumption Training is equivalent to the deterministic optimal-control problem (7)/(9) over the admissible class U_ad, with a pointwise Hamiltonian maximized at every state and time.
- domain assumption The terminal cost G is bounded uniformly continuous (and, for q∈(0,1), b>0 with G globally Lipschitz), or is replaced by a bounded extension on the forward-trajectory ball.
- domain assumption In the Mamba analysis, the state-transition matrix A is Hurwitz so that the discretized transition Ā_t is a contraction with exponent q=0.
- domain assumption Softmax in attention saturates so that self-attention is q=1; standard activations are Lipschitz or saturating; and super-polynomial primitives are excluded or wrapped by a q=0 controller.
Cite this review
Pith. "Pith review of Sharp Stability Threshold and Certification for Designing Stable Residual Architectures." pith.science (2026). https://pith.science/paper/G5GTSN6U
@misc{pith2026260714576,
author = {Pith},
title = {Pith review of: Sharp Stability Threshold and Certification for Designing Stable Residual Architectures},
year = {2026},
howpublished = {\url{https://pith.science/paper/G5GTSN6U}},
note = {Machine review of arXiv:2607.14576}
}
abstract
We propose \emph{the sublinear-growth principle} for deep residual architectures -- a sharp stability threshold on the input-magnitude exponent of every residual block's velocity field: $$\|v(x, t)\| \leq c\,\|x\|^q + b, \qquad q \in [0, 1].$$ The threshold $q = 1$ is established via two independent arguments. Classical ODE theory gives a global forward flow on $[0, T]$ at $q \le 1$ and exhibits divergent velocity fields at any $q > 1$. The optimal-control analysis, via the Hamilton-Jacobi-Bellman equation, sharpens this to a selection statement: the training optimum is bang-bang on the boundary of the admissible class, so the optimum at $q > 1$ blows up while the optimum at $q \le 1$ is safe by construction. The exponent criterion $q \le 1$ is thereby a necessary and sufficient condition for stable training. It clarifies architectural placements that ensure the stability of training and inference, explaining, for instance, the stabilizing role of layer normalization. The sublinear-growth velocity fields form \emph{the right function space} on which forward dynamics, adjoint sensitivity, and architectural composition are all well-controlled. An arithmetic of input-magnitude exponents under the five operations that build residual blocks enables efficient certification of $q_k \le 1$ at the level of architectural primitives, in place of ad hoc trial and error in the search for stable neural architectural designs. A parameter-free modification reduces the supercritical Mamba block from $q = 5$ to $q = 1$ without layer normalization, demonstrating this point. Experiments on Mamba and PatchTST confirm that the $q \le 1$ variants train stably: the criterion is the input-magnitude exponent, not the presence of a normalization layer.
Figures
Reference graph
Works this paper leans on
-
[1]
C. Anil, J. Lucas, and R. Grosse. Sorting out Lipschitz function approximation. In K. Chaudhuri and R. Salakhutdinov, editors,Proceedings of the 36th International Confer- ence on Machine Learning, volume 97 ofProceedings of Machine Learning Research, pages 291–301. PMLR, 09–15 Jun 2019
2019
-
[2]
J. L. Ba, J. R. Kiros, and G. E. Hinton. Layer normalization.arXiv preprint arXiv:1607.06450, 2016
Pith/arXiv arXiv 2016
-
[3]
Bachlechner, B
T. Bachlechner, B. P. Majumder, H. Mao, G. Cottrell, and J. McAuley. Rezero is all you need: fast convergence at large depth. In C. de Campos and M. H. Maathuis, editors,Pro- ceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence, volume 161 ofProceedings of Machine Learning Research, pages 1352–1361. PMLR, 27–30 Jul 2021
2021
-
[4]
Bardi and I
M. Bardi and I. Capuzzo-Dolcetta.Optimal Control and Viscosity Solutions of Hamilton– Jacobi–Bellman Equations. Birkhäuser, Boston, 1997
1997
-
[5]
Barles and P
G. Barles and P. E. Souganidis. Convergence of approximation schemes for fully nonlinear second order equations.Asymptotic Analysis, 4(3):271–283, 1991
1991
-
[6]
D. P. Bertsekas.Dynamic Programming and Optimal Control, volume I. Athena Scientific, Belmont, MA, 4th edition, 2017
2017
-
[7]
Bjorck, C
J. Bjorck, C. Gomes, B. Selman, and K. Q. Weinberger. Understanding batch normaliza- tion. InProceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, page 7705–7716, Red Hook, NY, USA, 2018. Curran Associates Inc
2018
-
[8]
J. S. Bridle. Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition. In F. F. Soulié and J. Hérault, editors, Neurocomputing, pages 227–236, Berlin, Heidelberg, 1990. Springer Berlin Heidelberg
1990
-
[9]
Characterizingsignalpropagationtoclosetheperformance gap in unnormalized resnets
A.Brock, S.De, andS.L.Smith. Characterizingsignalpropagationtoclosetheperformance gap in unnormalized resnets. InInternational Conference on Learning Representations, 2021
2021
-
[10]
Brock, S
A. Brock, S. De, S. L. Smith, and K. Simonyan. High-performance large-scale image recognition without normalization. In M. Meila and T. Zhang, editors,Proceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 1059–1071. PMLR, 2021
2021
-
[11]
H. Byun, Y. Choi, T. S. Kim, S. Park, and K. Song. Bounded hyperbolic tangent: A stable and efficient alternative to pre-layer normalization in large language models. InProceedings of the 43rd International Conference on Machine Learning (ICML), 2026
2026
-
[12]
R. T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. Duvenaud. Neural ordinary differential equations. InProceedings of the 32nd International Conference on Neural Information Pro- cessing Systems, NIPS’18, page 6572–6583, Red Hook, NY, USA, 2018. Curran Associates Inc
2018
-
[13]
International Series in Pure and Applied Mathematics
E.A.CoddingtonandN.Levinson.Theory of Ordinary Differential Equations. International Series in Pure and Applied Mathematics. McGraw-Hill, New York, 1955. 22
1955
-
[14]
M. G. Crandall, H. Ishii, and P.-L. Lions. User’s guide to viscosity solutions of second order partial differential equations.Bulletin of the American Mathematical Society, 27(1):1–67, 1992
1992
-
[15]
Springer-Verlag, Berlin, Heidelberg, 1967
J.M.Danskin.The Theory of Max-Min and Its Application to Weapons Allocation Problems. Springer-Verlag, Berlin, Heidelberg, 1967
1967
-
[16]
X. Dong, Y. Fu, S. Diao, W. Byeon, Z. Chen, A. S. Mahabaleshwarkar, S.-Y. Liu, M. V. keirsbilck, M.-H. Chen, Y. Suhara, Y. C. Lin, J. Kautz, and P. Molchanov. Hymba: A hybrid-head architecture for small language models. InThe Thirteenth International Con- ference on Learning Representations, 2025
2025
-
[17]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. InInternational Conference on Learning Representations, 2021
2021
-
[18]
Elfwing, E
S. Elfwing, E. Uchibe, and K. Doya. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning.Neural Networks, 107:3–11, 2018
2018
-
[19]
W. H. Fleming and H. M. Soner.Controlled Markov processes and viscosity solutions, volume 25 ofStochastic Modelling and Applied Probability. Springer Science & Business Media, 2nd edition, 2006
2006
-
[20]
H. Gouk, E. Frank, B. Pfahringer, and M. J. Cree. Regularisation of neural networks by enforcing lipschitz continuity.Mach. Learn., 110(2):393–416, Feb. 2021
2021
-
[21]
Gu and T
A. Gu and T. Dao. Mamba: Linear-time sequence modeling with selective state spaces. In First Conference on Language Modeling, 2024
2024
-
[22]
Haber and L
E. Haber and L. Ruthotto. Stable architectures for deep neural networks.Inverse Problems, 34(1):014004, 2017
2017
-
[23]
Hartman.Ordinary Differential Equations: Second Edition
P. Hartman.Ordinary Differential Equations: Second Edition. Classics in Applied Mathe- matics. Society for Industrial and Applied Mathematics (SIAM, 3600 Market Street, Floor 6, Philadelphia, PA 19104), 1982
1982
-
[24]
Hendrycks and K
D. Hendrycks and K. Gimpel. Gaussian error linear units (gelus), 2023
2023
-
[25]
Mobilenets: Efficientconvolutionalneuralnetworksformobilevisionapplications, 2017
A.G.Howard, M.Zhu, B.Chen, D.Kalenichenko, W.Wang, T.Weyand, M.Andreetto, and H.Adam. Mobilenets: Efficientconvolutionalneuralnetworksformobilevisionapplications, 2017
2017
-
[26]
Huang, J
L. Huang, J. Qin, Y. Zhou, F. Zhu, L. Liu, and L. Shao. Normalization techniques in training dnns: Methodology, analysis and application.IEEE Trans. Pattern Anal. Mach. Intell., 45(8):10173–10196, Aug. 2023
2023
-
[27]
Ioffe and C
S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by re- ducing internal covariate shift. In F. Bach and D. Blei, editors,Proceedings of the 32nd International Conference on Machine Learning, volume 37 ofProceedings of Machine Learn- ing Research, pages 448–456, Lille, France, 07–09 Jul 2015. PMLR
2015
-
[28]
N. K. Jha and B. Reagen. Aero: Entropy-guided framework for private llm inference, 2025
2025
-
[29]
K. Kan, X. Li, B. Zhang, T. Sahai, S. Osher, and M. Katsoulakis. Optimal control for transformer architectures: Enhancing generalization, robustness and efficiency. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. 23
2025
-
[30]
K. Kan, X. Li, B. J. Zhang, T. Sahai, S. Osher, K. Kumar, and M. A. Katsoulakis. Stability of transformers under layer normalization, 2025
2025
-
[31]
J. Kim, B. Lee, C. Park, Y. Oh, B. Kim, T. Yoo, S. Shin, D. Han, J. Shin, and K. M. Yoo. Peri-LN: Revisiting normalization layer in the transformer architecture. In A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu, editors,Proceedings of the 42nd International Conference on Machine Learning, volume 267 ofPr...
2025
-
[32]
Klambauer, T
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter. Self-normalizing neural net- works. InProceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 972–981, Red Hook, NY, USA, 2017. Curran Associates Inc
2017
-
[33]
Kushner and P
H. Kushner and P. Dupuis.Numerical Methods for Stochastic Control Problems in Contin- uous Time. Number V. 24 in Applications of mathematics. Springer, 2001
2001
-
[34]
Loshchilov, C.-P
I. Loshchilov, C.-P. Hsieh, S. Sun, and B. Ginsburg. nGPT: Normalized transformer with representation learning on the hypersphere. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[35]
Miyato, T
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida. Spectral normalization for generative adversarial networks. InInternational Conference on Learning Representations, 2018
2018
-
[36]
T. Q. Nguyen and J. Salazar. Transformers without tears: Improving the normalization of self-attention. In J. Niehues, R. Cattoni, S. Stüker, M. Negri, M. Turchi, T.-L. Ha, E. Salesky, R. Sanabria, L. Barrault, L. Specia, and M. Federico, editors,Proceedings of the 16th International Conference on Spoken Language Translation, Hong Kong, Nov. 2-3
-
[37]
Y. Nie, N. H. Nguyen, P. Sinthong, and J. Kalagnanam. A time series is worth 64 words: Long-term forecasting with transformers. InThe Eleventh International Conference on Learning Representations, 2023
2023
-
[38]
L. S. Pontryagin, V. G. Boltyanskii, R. V. Gamkrelidze, and E. F. Mishchenko.The Math- ematical Theory of Optimal Processes. Interscience Publishers, a division of John Wiley & Sons, New York, 1962
1962
-
[39]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al. Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019
2019
-
[40]
Santurkar, D
S. Santurkar, D. Tsipras, A. Ilyas, and A. Mądry. How does batch normalization help optimization? InProceedings of the 32nd International Conference on Neural Informa- tion Processing Systems, NIPS’18, page 2488–2498, Red Hook, NY, USA, 2018. Curran Associates Inc
2018
-
[41]
H. V. Tran.Hamilton-Jacobi Equations: Theory and Applications, volume 213 ofGraduate Studies in Mathematics. American Mathematical Society, 2021
2021
-
[42]
Ulyanov, A
D. Ulyanov, A. Vedaldi, and V. Lempitsky. Instance normalization: The missing ingredient for fast stylization, 2017
2017
-
[43]
van den Oord, O
A. van den Oord, O. Vinyals, and K. Kavukcuoglu. Neural discrete representation learn- ing. InProceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 6309–6318, Red Hook, NY, USA, 2017. Curran Associates Inc. 24
2017
-
[44]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors,Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017
2017
-
[45]
H. Wu, J. Xu, J. Wang, and M. Long. Autoformer: decomposition transformers with auto-correlation for long-term series forecasting. InProceedings of the 35th International Conference on Neural Information Processing Systems, NIPS ’21, Red Hook, NY, USA,
-
[46]
Wu and K
Y. Wu and K. He. Group normalization. InComputer Vision – ECCV 2018: 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part XIII, page 3–19, Berlin, Heidelberg, 2018. Springer-Verlag
2018
-
[47]
C. Zeng, Z. Liu, G. Zheng, and L. Kong. Cmamba: Channel correlation enhanced state space models for multivariate time series forecasting, 2024
2024
-
[48]
Zhang and R
B. Zhang and R. Sennrich.Root mean square layer normalization. Curran Associates Inc., Red Hook, NY, USA, 2019
2019
-
[49]
B. J. Zhang and M. A. Katsoulakis. A mean-field games laboratory for generative modeling, 2023
2023
-
[50]
Zhang, Y
H. Zhang, Y. N. Dauphin, and T. Ma. Residual learning without normalization via better initialization. InInternational Conference on Learning Representations, 2019
2019
-
[51]
H. Zhou, S. Zhang, J. Peng, S. Zhang, J. Li, H. Xiong, and W. Zhang. Informer: Beyond efficient transformer for long sequence time-series forecasting. InProceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 11106–11115, 2021
2021
-
[52]
J. Zhu, X. Chen, K. He, Y. LeCun, and Z. Liu. Transformers without normalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025. 25
2025
-
[2019]
Association for Computational Linguistics
-
[2021]
Curran Associates Inc
This paper was first reviewed by deepseek-v4-flash on August 2, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.