REVIEW 3 major objections 4 minor 38 references
Symplectic reservoir updates from Hamiltonians linear in momentum are exactly the maps that preserve Legendre graphs, so they preserve Legendre dynamics.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 14:43 UTC pith:DOEDE3ER
load-bearing objection The normal-form theorem is standard and the LD framing is neat, but Theorem 6.2's final step—graph preservation to strict convexity—fails, and the abstract's experiments are absent. the 3 major comments →
Symplectic Representation of Legendre Dynamics
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central theorem (Theorem 6.2) states that any Symplectic Reservoir update arising from an input-driven Hamiltonian at most linear in the momentum has the form τ_{dχ_t} ∘ f^♯_t — the composition of a cotangent lift of a base diffeomorphism with an exact fibre translation. Consequently it sends the Legendre graph L_ψ to L_{ψ∘f^{-1}+χ}. The geometric characterization (Theorem 5.7) shows these are exactly the symplectomorphisms preserving all Legendre-type Lagrangian submanifolds. The paper also proves that LTI-GPR and OU processes evolve on Legendre graphs, making them instances of Legendre dynamics.
What carries the argument
The central object is the Legendre graph L_ψ = {(q,p) : p = dψ(q)}, a Lagrangian submanifold of the cotangent bundle T^*Q. The normal form is τ_{dχ} ∘ f^♯: a cotangent lift f^♯(q,p) = (f(q), (f^{-1})^*p) composed with an exact fibre translation τ_{dχ}(q,p) = (q, p + dχ(q)). The workhorse theorem characterizes all symplectomorphisms preserving these graphs, and the dynamical realization shows that Hamiltonians of the form H(q,p) = p·A(q) + V(q) generate exactly these maps.
Load-bearing premise
The load-bearing identification is that preserving Legendre type structure (mapping every smooth potential graph to another smooth potential graph) is equivalent to preserving Legendre dynamics (keeping the trajectory inside a dualistic model with strictly convex potential), since the normal form's ψ∘f^{-1}+χ may lose convexity.
What would settle it
Take a Hamiltonian H(q,p) = p·A(q) + V(q) with V chosen so that the χ_t defined by Eq. (6.9) is concave and f_t is a nontrivial diffeomorphism. For a strictly convex ψ, check whether ψ∘f_t^{-1}+χ_t is strictly convex; if for some such f_t, χ_t it is not, then the SR update maps a Legendre graph to a smooth graph that is not Legendre, disproving preservation of Legendre dynamics even though the normal form holds.
If this is right
- Any Symplectic Reservoir built from a Hamiltonian linear in momentum preserves Legendre duality at every time step, so the represented trajectory stays within the family of Legendre graphs.
- The normal form is unique: no other symplectic map preserves all Legendre graphs, so the architecture is forced by the invariant rather than chosen heuristically.
- The result is coordinate-free and valid under reparameterization, aligning with the physical principle that laws should be independent of coordinates.
- The Markov semigroup characterization gives an infinitesimal criterion: a generator preserves an exponential family iff its action on densities is affine in the sufficient statistics, and this criterion is shown to hold for Ornstein-Uhlenbeck dynamics.
- Legendre dynamics provides a common structural umbrella for LTI-GPR and OU, turning regime switches into geometric events such as curvature jumps in the evolving potential.
Where Pith is reading between the lines
- If preservation of Legendre graphs is taken as the definition of structure-preserving representation, then any recurrent system whose core is Hamiltonian with linear momenta is uniquely suited to represent exponential-family dynamics; generic symplectic RNNs would not qualify because they do not enforce this graph invariance.
- The normal form yields ψ' = ψ∘f^{-1}+χ, but this sum need not be strictly convex even when ψ is, so 'preserves smooth potential graphs' and 'preserves Legendre dynamics' may diverge; a convexity-preserving condition on f and χ would be needed for a true exponential-family flow.
- The readout potential ψ' = ψ∘f^{-1}+χ could be interpreted as a learned representation, suggesting that training the readout is equivalent to inferring the Legendre potential that organizes the dual trajectory.
- One concrete test: for a given reservoir, measure empirically whether the distribution of (q,p) stays on a Legendre graph; the normal form predicts exact preservation for linear-in-momentum Hamiltonians and violation for generic symplectic ones.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a geometric framework for reservoir computing in which the internal state update is required to preserve Legendre duality, formalized as invariance of Legendre graphs L_ψ = {(q,p): p=dψ(q)} in T*Q. The authors introduce 'Legendre dynamics' as processes that remain on such graphs, and construct Symplectic Reservoirs (SR) from input-driven Hamiltonians. The central claim is a characterization: symplectomorphisms preserving all Legendre graphs are exactly cotangent lifts of base diffeomorphisms composed with exact fiber translations (Theorem 5.7), and that SR updates generated by Hamiltonians at most linear in momentum have this normal form (Theorem 6.1). From this, the paper asserts (Theorem 6.2) that such SR updates preserve Legendre dynamics. The appendix gives a Markov-semigroup proof that Ornstein-Uhlenbeck processes are Legendre dynamics, plus an equivalence theorem between semigroup, Lagrangian, and infinitesimal characterizations.
Significance. If correct, the paper would provide an elegant bridge between symplectic geometry, information geometry, and reservoir computing, with a concrete architectural design principle: the reservoir update is derived from an invariant rather than heuristically imposed. The identification of Legendre graphs as Lagrangian submanifolds is sound and standard, and the explicit construction of SR from a quadratic input-driven Hamiltonian with exact discretization is a useful engineering contribution. The Appendix A semigroup characterization and its application to OU processes is self-contained and appears correct. However, the central result connecting SR to Legendre dynamics is not established: the map from Legendre-type graph preservation to preservation of strictly convex Legendre potentials is missing, and a simple counterexample shows the claimed preservation can fail. The abstract additionally promises numerical experiments that do not appear anywhere in the manuscript. The paper's main claim therefore does not hold as stated.
major comments (3)
- [§6, Theorem 6.2] The theorem states that an SR update maps L_ψ to L_{ψ∘f^{-1}+χ} and 'thereby preserves Legendre dynamics.' This implication is invalid as written. Definition 2.2 requires the new potential to be C² and strictly convex, while Definition 5.5 and Theorem 5.7 only guarantee smoothness of ψ∘f^{-1}+χ. Strict convexity is not preserved by composition with a diffeomorphism plus an arbitrary function χ. Concretely, take Q=R and H(q,p)=q²/2 (so S=0, L=1, C_q=C_p=0 in the SR Hamiltonian of §6). The flow over time t gives q(t)=q0, p(t)=p0-t q0. Starting on L_{q²/2}, i.e. p0=q0, the image is the graph of ψ'(q)=(1-t)q²/2. For t>1 this is strictly concave, so the update leaves the dualistic model of Definition 2.2. Thus Theorem 6.2 overclaims; at best it preserves smooth Legendre-type graphs, not Legendre dynamics.
- [§5, Theorem 5.6 proof] The (⇒) direction of the proof is circular. After defining Ψ=τ_{-β}∘F, the text asserts 'there exists function P, P′ such that Ψ(q,p)=(f(q),P(q,p))' and then uses this coordinate form to prove that Ψ is fiber-preserving. But the assertion that the first component is f(q), independent of p, is already a form of base-fiber structure that has not been established. The proof of vertical-bundle preservation therefore assumes the conclusion it is intended to derive. The classification may be true (it is a known result), but the proof as written does not establish it.
- [Abstract / full text] The abstract states: 'Numerical experiments confirm the normal-form identities and distinguish Legendre preserving Hamiltonian SRs from generic symplectic and standard reservoir baselines.' The full manuscript contains no experimental section, no figures, no tables, and no data. For a cs.LG paper this is a significant mismatch between claimed and actual content. Either the experiments must be supplied, or the claim must be removed.
minor comments (4)
- [§4/§5 terminology] The paper uses 'Legendre type structure' (Definition 5.5, smooth potential) and 'Legendre dynamics' (Definition 2.2, C² strictly convex potential) almost interchangeably. This ambiguity is not just cosmetic; it is the source of the gap in Theorem 6.2. The authors should state explicitly where convexity enters and whether it is preserved.
- [§5, proof of Theorem 5.4] Typo: 'Lagrangian submnaifolds' should be 'Lagrangian submanifolds'; also 'Let V by a Lagrangian subspace' should be 'Let V be...'.
- [§6, Equation (6.8)] The derivation of p_j(t) involves a chain-rule step that is correct but presented too tersely; the exchange of ∂/∂q_i0 and ∂/∂q_j(t) should be explained or displayed more clearly.
- [Appendix A] The notation η is used both for the natural parameter (η,Λ) and for the expectation parameter η(θ)=∇ψ(θ) in the same appendix (e.g., Theorem A.3 and Theorem A.7). This conflicts with Remark 2.3 and makes the semigroup derivation harder to follow. Also, b(θ) is introduced as a scalar but then written as a function with arguments; aligning notation would improve readability.
Circularity Check
Circular proof of the graph-preservation normal form and definitional equivocation in 'preserves Legendre dynamics'.
specific steps
-
other
[Theorem 5.6 (⇒ direction), proof that Ψ is fiber-preserving, Section 5]
"By construction Ψ = τ_{−β} ◦F, so there exists function P, P′ such that Ψ, Ψ^{-1} are of the following form: Ψ(q, p) = (f(q), P(q, p)); Ψ^{-1}(q′, p′) = (f^{-1}(q′), P′(q′, p′))."
Fiber-preservation means precisely that the first component of Ψ and Ψ^{-1} does not depend on the fiber variables p/p′. The proof has not yet established this; it posits the representation Ψ(q,p)=(f(q),P(q,p)) and its inverse, then computes G_k = q^k ∘ Ψ^{-1} = f^{-1}(q′)_k, concluding ∂G_k/∂p′=0. That conclusion is exactly the assumed fiber-preserving form, not a derivation from the symplectic or graph-preserving hypotheses. Since Theorem 5.6 is the basis for the normal form in Theorem 5.7 and for Theorem 6.2, this circular step is load-bearing.
-
self definitional
[Theorem 6.2, Section 6; Definitions 2.2 and 5.5]
"the SR update map is of the form τ_{dχ_t}∘f^♯_t and hence maps L_ψ to L_{ψ∘f^{-1}+χ}. Thereby preserves Legendre dynamics."
Definition 5.5 defines 'preserves Legendre type structure' as mapping every smooth-potential graph L_ψ to another smooth-potential graph L_{ψ_out}, with no convexity condition. Definition 2.2 defines Legendre dynamics by requiring the next potential ψ′ to be C² and strictly convex so that the law remains in the exponential-family dualistic model. The proof establishes only the smooth-graph statement. The jump 'Thereby preserves Legendre dynamics' identifies these two different notions without proof: ψ∘f^{-1}+χ need not be strictly convex (e.g. H(q,p)=Lq²/2 with L>1 maps p=q to p=(1−L)q, a concave graph). Thus the final preservation conclusion is imported by equivocation between 'Legendre type' and 'Legendre dynamics', not derived from the proved equations.
full rationale
The paper contains no fitted parameters and no load-bearing self-citation chain; the explicit computation that a Hamiltonian at most linear in p generates τ_{dχ_t}∘f^♯_t (Theorem 6.1) and the SR construction are derived in an independent, parametrized way. However, the derivation chain has two junctions where the claimed preservation is not actually derived. First, the proof of Theorem 5.6, the classification of symplectomorphisms preserving graphs of closed 1-forms, assumes the fiber-preserving form Ψ(q,p)=(f(q),P(q,p)) that it is supposed to prove; the subsequent vertical-basis computation only unpacks this assumed form. Since Theorem 5.7 and Theorem 6.2 both invoke Theorem 5.6, this circular proof step is load-bearing. Second, Theorem 6.2 concludes 'Thereby preserves Legendre dynamics' from the normal form, but the normal form only preserves smooth Legendre graphs (Definition 5.5), not the strictly convex potentials required by Definition 2.2. No convexity of ψ∘f^{-1}+χ is proved, and it can fail within the paper's own SR class. The central preservation claim therefore rests partly on a definitional equivocation rather than on the proved equations. Because the core symplectic computations are still substantive and independently checkable, the circularity is partial rather than total; however, the final 'preserves Legendre dynamics' assertion is not supported by the derivation chain as written.
Axiom & Free-Parameter Ledger
free parameters (1)
- Reservoir Hamiltonian matrices M,C (or S,L,C_q,C_p) =
arbitrary (hyperparameters)
axioms (6)
- domain assumption Hamiltonian vector field A(q) is complete so the flow f_t exists globally
- ad hoc to paper The transformed potential ψ∘f^{-1}+χ stays in the dualistic class (C² strictly convex)
- standard math Corollary 3.31 of Bates-Weinstein: zero-section and fiber-preserving symplectomorphism is a cotangent lift
- standard math Dom(L) dense in C0 and Riesz-Markov identification of measures
- domain assumption Gaussian exponential family is minimal and regular, θ→pθ injective
- domain assumption Zero-order hold: input constant over each Δt and exact matrix exponential discretization
read the original abstract
Modern learning systems act on internal representations of data, yet how these representations encode underlying physical or statistical structure is often left implicit. In physics, symplecticity keeps Hamiltonian systems faithful to their phase-space geometry. Recent learning methods impose such geometric structure either in the dynamics or through training losses. Here we ask a different question: what would it mean for the representation itself to obey a symplectic conservation law? We pose this representation-level constraint through Legendre duality: the relation $p = d\psi(q)$ between primal and dual coordinates, which in exponential family models is the information-geometric pairing of natural and expectation parameters. We formalize Legendre dynamics as stochastic processes whose trajectories remain on Legendre graphs, where the evolving primal-dual parameters stay Legendre dual. We show that this class includes linear time-invariant Gaussian process regression and Ornstein-Uhlenbeck dynamics. Geometrically, we characterize the symplectomorphisms of cotangent bundles that preserve all Legendre graphs. We show that these maps are exactly cotangent lifts of base diffeomorphisms followed by exact fibre translations. This gives an explicit normal form for Legendre-preserving representation updates. Dynamically, we prove that the normal form is realized by Hamiltonians that are at most linear in the momentum. This realization principle is used to construct linear and nonlinear Hamiltonian Symplectic Reservoirs (SR) whose recurrent updates preserve Legendre graphs by construction. This is the only normal form that preserves Legendre duality, so the architecture follows from the invariant. Numerical experiments confirm the normal-form identities and distinguish Legendre preserving Hamiltonian SRs from generic symplectic and standard reservoir baselines.
Reference graph
Works this paper leans on
-
[1]
Methods of information geometry , volume 191
Shun-ichi Amari and Hiroshi Nagaoka. Methods of information geometry , volume 191. American Mathematical Soc., 2000
2000
-
[2]
Mathematical methods of classical mechanics , volume 60
Vladimir Igorevich Arnol'd. Mathematical methods of classical mechanics , volume 60. Springer Science & Business Media, 2013
2013
-
[3]
On invariance and selectivity in representation learning
Fabio Anselmi, Lorenzo Rosasco, and Tomaso Poggio. On invariance and selectivity in representation learning. Information and Inference: A Journal of the IMA , 5(2):134--158, 2016
2016
-
[4]
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives. IEEE transactions on pattern analysis and machine intelligence , 35(8):1798--1828, 2013
2013
-
[5]
On explaining the surprising success of reservoir computing forecaster of chaos? the universal machine learning dynamical system with contrast to var and dmd
Erik Bollt. On explaining the surprising success of reservoir computing forecaster of chaos? the universal machine learning dynamical system with contrast to var and dmd. Chaos: An Interdisciplinary Journal of Nonlinear Science , 31(1), 2021
2021
-
[6]
Projecting the fokker-planck equation onto a finite dimensional exponential family
Damiano Brigo and Giovanni Pistone. Projecting the fokker-planck equation onto a finite dimensional exponential family. arXiv preprint arXiv:0901.1308 , 2009
Pith/arXiv arXiv 2009
-
[7]
Lectures on the Geometry of Quantization , volume 8
Sean Bates and Alan Weinstein. Lectures on the Geometry of Quantization , volume 8. American Mathematical Soc., 1997
1997
-
[8]
Group equivariant convolutional networks
Taco Cohen and Max Welling. Group equivariant convolutional networks. In International conference on machine learning , pages 2990--2999. PMLR, 2016
2016
-
[9]
Symplectic recurrent neural networks
Zhengdao Chen, Jianyu Zhang, Martin Arjovsky, and L \'e on Bottou. Symplectic recurrent neural networks. arXiv preprint arXiv:1909.13334 , 2019
Pith/arXiv arXiv 1909
-
[10]
Information processing capacity of dynamical systems
Joni Dambre, David Verstraeten, Benjamin Schrauwen, and Serge Massar. Information processing capacity of dynamical systems. Scientific reports , 2(1):514, 2012
2012
-
[11]
Universality of real minimal complexity reservoir
Robert Simon Fong, Boyu Li, and Peter Tino. Universality of real minimal complexity reservoir. Proceedings of the AAAI Conference on Artificial Intelligence , 39(16):16622--16629, 2025
2025
-
[12]
Linear simple cycle reservoirs at the edge of stability perform fourier decomposition of the input driving signals
Robert Simon Fong, Boyu Li, and Peter Tiňo. Linear simple cycle reservoirs at the edge of stability perform fourier decomposition of the input driving signals. Chaos , 35(4):043109, 2025
2025
-
[13]
Euler state networks: Non-dissipative reservoir computing
Claudio Gallicchio. Euler state networks: Non-dissipative reservoir computing. Neurocomputing , 579:127411, 2024
2024
-
[14]
Echo state networks are universal
Lyudmila Grigoryeva and Juan-Pablo Ortega. Echo state networks are universal. Neural Networks , 108:495--508, 2018
2018
-
[15]
Universal discrete-time reservoir computers with stochastic inputs and linear readouts using non-homogeneous state-affine systems
Lyudmila Grigoryeva and Juan-Pablo Ortega. Universal discrete-time reservoir computers with stochastic inputs and linear readouts using non-homogeneous state-affine systems. Journal of Machine Learning Research , 19(24):1--40, 2018
2018
-
[16]
Reservoir computing universality with stochastic inputs
Lukas Gonon and Juan-Pablo Ortega. Reservoir computing universality with stochastic inputs. IEEE transactions on neural networks and learning systems , 31(1):100--112, 2019
2019
-
[17]
Kalman filtering and smoothing solutions to temporal gaussian process regression models
Jouni Hartikainen and Simo S \"a rkk \"a . Kalman filtering and smoothing solutions to temporal gaussian process regression models. In 2010 IEEE international workshop on machine learning for signal processing , pages 379--384. IEEE, 2010
2010
-
[18]
Reservoir computing beyond memory-nonlinearity trade-off
Masanobu Inubushi and Kazuyuki Yoshimura. Reservoir computing beyond memory-nonlinearity trade-off. Scientific reports , 7(1):10199, 2017
2017
-
[19]
Short term memory in echo state networks
Herbert Jaeger. Short term memory in echo state networks. 2001
2001
-
[20]
echo state
Herbert Jaeger. The “echo state” approach to analysing and training recurrent neural networks. Bonn, Germany: German national research center for information technology gmd technical report , 148(34):13, 2001
2001
-
[21]
Tutorial on training recurrent neural networks, covering bppt, rtrl, ekf and the echo state network approach
Herbert Jaeger. Tutorial on training recurrent neural networks, covering bppt, rtrl, ekf and the echo state network approach. 5(1), 2002
2002
-
[22]
Harnessing nonlinearity: Predicting chaotic systems and saving energy in wireless communication
Herbert Jaeger and Harald Haas. Harnessing nonlinearity: Predicting chaotic systems and saving energy in wireless communication. science , 304(5667):78--80, 2004
2004
-
[23]
An introduction to probabilistic graphical models, 2003
Michael I Jordan. An introduction to probabilistic graphical models, 2003
2003
-
[24]
Markov-modulated affine processes
Kevin Kurt and R \"u diger Frey. Markov-modulated affine processes. Stochastic Processes and their Applications , 153:391--422, 2022
2022
-
[25]
On translation invariance in cnns: Convolutional layers can exploit absolute spatial location
Osman Semih Kayhan and Jan C van Gemert. On translation invariance in cnns: Convolutional layers can exploit absolute spatial location. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 14274--14285, 2020
2020
-
[26]
Metric learning: A survey
Brian Kulis. Metric learning: A survey. Foundations and Trends® in Machine Learning , 5(4):287--364, 2013
2013
-
[27]
Introduction to smooth manifolds
John M Lee. Introduction to smooth manifolds . Springer, 2003
2003
-
[28]
Simple Cycle Reservoirs are Universal
Boyu Li, Robert Simon Fong, and Peter Tino. Simple Cycle Reservoirs are Universal . Journal of Machine Learning Research , 25(158):1--28, 2024
2024
-
[29]
Reservoir computing approaches to recurrent neural network training
Mantas Lukosevicius and Herbert Jaeger. Reservoir computing approaches to recurrent neural network training. Computer science review , 3(3):127--149, 2009
2009
-
[30]
Real-time computing without stable states: A new framework for neural computation based on perturbations
Wolfgang Maass, Thomas Natschl \"a ger, and Henry Markram. Real-time computing without stable states: A new framework for neural computation based on perturbations. Neural computation , 14(11):2531--2560, 2002
2002
-
[31]
Stochastic processes and applications
Grigorios A Pavliotis. Stochastic processes and applications. Texts in applied mathematics , 60, 2014
2014
-
[32]
Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations
Maziar Raissi, Paris Perdikaris, and George E Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational physics , 378:686--707, 2019
2019
-
[33]
Minimum complexity echo state network
Ali Rodan and Peter Ti n o. Minimum complexity echo state network. IEEE transactions on neural networks , 22(1):131--144, 2010
2010
-
[34]
Bayesian filtering and smoothing , volume 17
Simo S \"a rkk \"a and Lennart Svensson. Bayesian filtering and smoothing , volume 17. Cambridge university press, 2023
2023
-
[35]
Predicting the future of discrete sequences from fractal representations of the past
Peter Tino and Georg Dorffner. Predicting the future of discrete sequences from fractal representations of the past. Machine Learning , 45(2):187--217, 2001
2001
-
[36]
Dynamical systems as temporal feature spaces
Peter Tino. Dynamical systems as temporal feature spaces. Journal of Machine Learning Research , 21(44):1--42, 2020
2020
-
[37]
Universal time-series representation learning: A survey
Patara Trirat, Yooju Shin, Junhyeok Kang, Youngeun Nam, Jihye Na, Minyoung Bae, Joeun Kim, Byunghyun Kim, and Jae-Gil Lee. Universal time-series representation learning: A survey. arXiv preprint arXiv:2401.03717 , 2024
Pith/arXiv arXiv 2024
-
[38]
Recent advances in physical reservoir computing: A review
Gouhei Tanaka, Toshiyuki Yamane, Jean Benoit H \'e roux, Ryosho Nakane, Naoki Kanazawa, Seiji Takeda, Hidetoshi Numata, Daiju Nakano, and Akira Hirose. Recent advances in physical reservoir computing: A review. Neural Networks , 115:100--123, 2019
2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.