Pith. sign in

REVIEW 2 major objections 4 minor 83 references

SHiPPO makes channel interaction part of online polynomial memory itself: stored HiPPO coefficients live in a moving frame and obey Sylvester dynamics driven by a right-transport path.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 05:11 UTC pith:W7T2L4IB

load-bearing objection Clean HiPPO lift to Sylvester transported memory with honest mechanism diagnostics; group-local scan is a real restriction, not a free lunch. the 2 major comments →

arxiv 2607.03055 v1 pith:W7T2L4IB submitted 2026-07-03 cs.LG

SHiPPO: Recurrent Memory with Transported Polynomial Projections

classification cs.LG
keywords SHiPPOHiPPOstate space modelsselective SSMsSylvester dynamicstransported polynomial projectionsrecurrent memoryblock-affine scan
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

HiPPO gives recurrent states a clear meaning as coefficients of an online polynomial projection of history, but only in fixed channel coordinates. Modern selective state-space models add token-dependent control and channel mixing, yet those interactions usually sit outside the memory semantics. SHiPPO (Sylvester HiPPO) lifts the projection problem into a moving channel frame: for any fixed or realized right-transport path it transports both the approximation family and the channel metric, so the state is ordinary HiPPO in a tied moving frame and follows Sylvester coefficient dynamics that keep the left online-memory operator while adding a right-action term on already-written memory. For selective-SSM use the paper derives a restricted group-local cell with controller-compatible right actions, exponential-adjusted updates, and an exact block-affine scan, plus a criterion for when right transports collapse to static mixing plus independent banks. Controlled diagnostics show that raising current-token write rank improves ordinary error but cannot recover order-sensitive changes to stored memory, whereas transported variants recover that signal—and lose it when the transport path is removed—supporting a mechanistically grounded transported-memory prior rather than a claim of broad sequence-modeling dominance.

Core claim

Coupling a transported approximation family with a transported channel metric yields Sylvester coefficient dynamics ˙C = A_L C + B_L f^⊤ + C A_R. Conditional on any realized right-transport path, the state is ordinary HiPPO in a tied moving channel frame: the left operator is inherited from the online projection problem while the right action is external path transport of already-written memory coordinates.

What carries the argument

The SHiPPO online approximation problem (Definition 2.1) and its Sylvester dynamics (Theorem 2.2): jointly transporting family and metric produces the right-action gauge term C A_R on stored coefficients, with a pathwise lift of any closed one-sided HiPPO equation (Corollary 2.3) and a scan-compatible group-local realization for selective SSMs.

Load-bearing premise

That a restricted group-local, controller-compatible right action—kept independent of the main memory state so the block-affine scan stays exact—still carries the operator-level transported-memory meaning the theory assigns to a general right path.

What would settle it

On the paired noncommutative diagnostic, train a transported model that recovers Pair ΔNMSE near zero, then at evaluation replace every right action R_t by the identity while freezing all other weights; if the paired-difference signal does not return to near one (as reported), the recovered signal is not mediated by future right transport of already-written memory.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper introduces SHiPPO, a transported online-projection memory prior that lifts HiPPO coefficient dynamics into a moving channel frame. Conditional on any fixed or realized right-transport path, the approximation family and channel metric are transported together, so the coefficient state is ordinary HiPPO in a tied moving frame and obeys the Sylvester ODE ˙C = A_L C + B_L f^⊤ + C A_R (Theorem 2.2, Corollary 2.3). For selective-SSM execution the authors derive a restricted group-local, controller-compatible cell with exponential-adjusted updates and exact block-affine scan, plus a simultaneous-reducibility collapse criterion for right transports. Controlled paired noncommutative diagnostics separate high-rank current-token writes from future right action on already-written memory; transported variants recover the order-sensitive signal, which vanishes under an evaluation-time R_t → I intervention. Transport-MQAR supplies complementary autoregressive evidence while leaving the preferred right-action realization open. Claims are scoped to memory mechanisms rather than broad sequence-modeling dominance.

Significance. If the pathwise lift and the write-versus-transport separation hold—as the derivations and interventions indicate—this is a genuine contribution to inductive-bias design for recurrent memory: channel interaction is made part of the online-approximation semantics rather than an architectural mixer. Strengths include complete normal-equation, conjugacy, gauge, scan-closure, and discretization arguments (Appendices A–C), an explicit collapse criterion (Proposition 3.3), and falsifiable diagnostics with seed-level means/stds and a clean R→I ablation. The work is carefully scoped and does not overclaim empirical dominance. Residual open questions (preferred right-action backend; group-local fidelity) are largely acknowledged and do not erase the operator-level contribution.

major comments (2)
  1. [Section 3; Appendix D.6, Table 6] Section 3 presents the group-local, controller-compatible cell as the practical SHiPPO lift for selective SSMs, yet Appendix D.6 shows that oracle-R with group size 2 yields PairΔNMSE = 1 (complete failure) while group size 4 (full transport for d=4) succeeds. The formal separation claim is stated for full transport, but contribution (iii) and the selective-SSM narrative rest on the restricted cell. The manuscript should either (a) demonstrate a group-local design that recovers the paired-difference signal without relying on front-end/readout recoding, or (b) demote the group-local cell more explicitly to a computational approximation with quantified fidelity loss relative to the operator-level prior.
  2. [Section 4.3; Table 1; Appendix E.5] Table 1 and Appendix E.4–E.5: Transport-MQAR gains over no-right/static-basis controls are modest (e.g., coordinate accuracy 0.110 vs 0.102 at length 4096), exact-accuracy leadership flips between DirectGen-SingleExp and StructGen-Split, and the suffix-zeroing counterfactual only shows that controller coordinates are used—not that a noncommutative transport geometry was learned. As complementary evidence this is acceptable under the paper’s scope, but the main-text framing that these results support “learned right-action pathways” should be tightened to match the weaker identification actually achieved.
minor comments (4)
  1. [Figure 1] Figure 1 is useful but dense; a short caption callout distinguishing the abstract operator (left) from the scan-compatible restriction (right) would help readers who skip Section 3.
  2. [Section 3.4, Eq. (3)] Notation for the discrete cell mixes L_t, R_t, bU_t and λ_t; a one-line glossary at the start of Section 3.4 would reduce back-references to (1)–(3).
  3. [Section 5; Appendix D.7] Appendix D.7 notes excluded exploratory diagnostics; a single sentence in the main Discussion pointing readers there would improve transparency without expanding the main text.
  4. [References] Several related-work citations are 2025–2026 arXiv preprints; ensure final versions or stable identifiers are used at camera-ready if available.

Circularity Check

0 steps flagged

No significant circularity: Sylvester dynamics are derived from an explicit transported projection objective, and diagnostics use held-out interventions rather than fitted-as-prediction.

full rationale

The load-bearing theoretical chain is Definition 2.1 (transported approximation family + coupled channel metric) → normal equation (Prop. A.3 / Thm. 2.2) → Sylvester coefficient ODE by Leibniz differentiation under HiPPO closure. That is a standard variational derivation, not a recurrence postulated and then re-labeled as projection memory. Corollary 2.3 is a pathwise lift of any closed one-sided HiPPO-style equation and is explicitly conditional on an external right-transport path; it does not smuggle the target dynamics into the definition of the left operator. The scan-compatible cell (Sec. 3) is openly a restricted realization (group-tied diagonal left, controller-compatible right) with proved block-affine scan closure, not a claim that the restriction is forced by uniqueness. Empirically, the paired noncommutative diagnostic and the evaluation-time R_t → I intervention freeze other weights and remove the transport pathway; the paired-difference signal disappears, so the result is not forced by construction or by a fitted normalization. No self-citation uniqueness theorem, no fitted parameter renamed as prediction, and no mere renaming of a known empirical pattern under new coordinates. Residual open questions (preferred right-action realization, group-local fidelity) are scoped limitations, not circular reductions.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 3 invented entities

The central claim rests on the classical HiPPO projection setup plus a new right-transport path that is treated as an external modeling choice; the scan-compatible realization further assumes controller-compatible (state-independent) group-local actions so that an exact block-affine scan exists. Free parameters appear mainly in the experimental cells and generator families; the invented entities are the SHiPPO operator itself and the diagnostic tasks used to isolate the mechanism.

free parameters (4)
  • group width P / number of groups G
    Chosen by hand for the scan-compatible cell; the paper notes that full channelwise lift loses the small scan algebra, so P is a computational free parameter that also affects expressivity (Appendix D.6 audit).
  • source rank r and write-rank baselines
    Hand-chosen ranks {1,2,4,8} in the no-right baseline; the separation claim depends on showing that even large r fails to recover PairΔNMSE.
  • right-generator library (skew, nilpotent, low-rank, diagonal)
    Structured family in Eq. (2) and the 13 finite-field operations in Transport-MQAR are design choices that determine whether transports are non-reducible.
  • discretization weights λ_t and step sizes Δ_t
    Two-point exponential source rule parameters; they affect local truncation order but are not fitted to the diagnostic targets.
axioms (4)
  • domain assumption HiPPO closure condition ∂_t ψ = A_L ψ (τ < t) and invertible Gram matrix G(t)
    Inherited from Gu et al.; required for the left operator and the normal equation (Appendix A.2).
  • standard math A_R is integrable so the state-transition family P(t,τ) exists and lies in GL(d)
    Standard linear ODE theory (Coddington–Levinson); used throughout Section 2 and Appendix B.
  • ad hoc to paper Right transport is controller-compatible (independent of main memory H) so the step map remains affine and the block-affine scan algebra closes
    Definition 3.1 and Proposition 3.2; state-coupled transport is allowed pathwise but breaks the finite scan summary used by the selective cell.
  • standard math Simultaneous block-reducibility of the generator family implies collapse to static mixing plus independent banks
    Proposition 3.3; algebraic fact used to motivate non-reducible split-flow designs.
invented entities (3)
  • SHiPPO online approximation problem (transported family G_SH_t + metric M_P) independent evidence
    purpose: Defines the coefficient matrix as the minimizer of a jointly transported projection objective, yielding Sylvester dynamics.
    Core new object (Definition 2.1); independent evidence is the conjugacy to ordinary HiPPO and the diagnostic recovery of order-sensitive signals.
  • Scan-compatible group-local SHiPPO cell with exponential-adjusted updates no independent evidence
    purpose: Restricted realization that preserves exact block-affine scan and recurrent decoding for selective SSMs.
    Computational restriction of the abstract prior (Section 3); not claimed to be the unique or optimal realization.
  • Paired noncommutative transport diagnostic and Transport-MQAR no independent evidence
    purpose: Synthetic tasks that isolate future right transport of already-written memory from high-rank current writes.
    Purpose-built diagnostics (Section 4, Appendices D–E); they supply the empirical separation but are not external benchmarks.

pith-pipeline@v1.1.0-grok45 · 43876 in / 3336 out tokens · 38774 ms · 2026-07-12T05:11:04.370481+00:00 · methodology

0 comments
read the original abstract

HiPPO gives recurrent states memory semantics as coefficients of online polynomial projections, but in fixed channel coordinates. Modern selective SSMs, by contrast, rely on token-dependent control and channel interaction. We introduce SHiPPO (Sylvester HiPPO), a transported projection-memory prior that lifts HiPPO coefficient memories into a moving channel frame. For any fixed or realized right-transport path, SHiPPO transports the approximation family and channel metric together; conditional on that path, the state is ordinary HiPPO in a tied moving frame and follows Sylvester coefficient dynamics, preserving the left online-memory operator while adding right-action transport. For selective-SSM execution, we derive a restricted group-local realization with controller-compatible right actions, exponential-adjusted updates, exact block-affine scan, and recurrent decoding. We also give a simultaneous-reducibility criterion identifying when right transports collapse to static mixing plus independent scalar or blockwise banks. Controlled diagnostics show that larger current-token write rank improves ordinary prediction error but cannot recover order-sensitive changes to already-written memory; transported-memory variants recover this signal, which disappears when the transport pathway is removed. A finite-field associative-recall diagnostic with interleaved bindings, operations, and queries provides complementary autoregressive evidence while leaving the preferred right-action realization open. Taken together, these results support SHiPPO as a mechanistically grounded transported-memory prior, with evidence focused on memory mechanisms rather than broad sequence-modeling dominance.

Figures

Figures reproduced from arXiv: 2607.03055 by Bum Jun Kim, Tomoya Mizuguchi.

Figure 1
Figure 1. Figure 1: Overview of the SHiPPO lift. Ordinary HiPPO gives one-sided online projection memory [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Main paired-transport diagnostic. (a) Paired examples share the same payload and op [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

83 extracted references · 1 canonical work pages

  1. [1]

    On the expressiveness of state space models via temporal logics, 2026

    Eric Alsmann, Lowejatan Noori, and Martin Lange. On the expressiveness of state space models via temporal logics, 2026. URLhttps://arxiv.org/abs/2601.19467

  2. [2]

    Zoology: Measuring and improving recall in efficient language models

    Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, and Christopher Ré. Zoology: Measuring and improving recall in efficient language models. InInternational Conference on Learning Representations, 2024. URLhttps:// openreview.net/forum?id=LY3ukUANko

  3. [3]

    xL- STM: Extended long short-term memory

    Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prud- nikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. xL- STM: Extended long short-term memory. InAdvances in Neural Information Processing Sys- tems, 2024. URLhttps://openreview.net/forum?id=ARAxPPIAhq

  4. [4]

    Solution formulas for differential sylvester and lyapunov equations.Calcolo, 56:51, 2019

    Maximilian Behr, Peter Benner, and Jan Heiland. Solution formulas for differential sylvester and lyapunov equations.Calcolo, 56:51, 2019. doi: 10.1007/s10092-019-0348-x

  5. [5]

    Blelloch

    Guy E. Blelloch. Prefix sums and their applications. Technical Report CMU-CS-90-190, School of Computer Science, Carnegie Mellon University, 1990. URLhttps://www.cs. cmu.edu/~scandal/papers/CMU-CS-90-190.html

  6. [6]

    Bo Chang, Minmin Chen, Eldad Haber, and Ed H. Chi. AntisymmetricRNN: A dynamical system view on recurrent neural networks. InInternational Conference on Learning Represen- tations, 2019. URLhttps://openreview.net/forum?id=ryxepo0cFX

  7. [7]

    Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neu- ral ordinary differential equations. InAdvances in Neural Information Process- ing Systems, volume 31, 2018. URLhttps://papers.neurips.cc/paper/ 7892-neural-ordinary-differential-equations

  8. [8]

    Learning phrase representations using RNN encoder– decoder for statistical machine translation

    Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder– decoder for statistical machine translation. InProceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1724–1734. Asso- ciation for Compu...

  9. [9]

    Nicola Muca Cirone, Antonio Orvieto, Benjamin Walker, Cristopher Salvi, and Terry J. Lyons. Theoretical foundations of deep selective state-space models. InAdvances in Neural Informa- tion Processing Systems, 2024. URLhttps://openreview.net/forum?id=3SzrqwupUx

  10. [10]

    Coddington and Norman Levinson.Theory of Ordinary Differential Equations

    Earl A. Coddington and Norman Levinson.Theory of Ordinary Differential Equations. McGraw-Hill, New York, 1955

  11. [11]

    Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality

    Tri Dao and Albert Gu. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Machine Learning Research, pages 10041– 10071. PMLR, 2024. URLhttps://proceedings.mlr.press/v235/dao24a.html. 10

  12. [12]

    Reza Ebrahimi and Roland Memisevic

    M. Reza Ebrahimi and Roland Memisevic. Revisiting bi-linear state transitions in recurrent neural networks. InAdvances in Neural Information Processing Systems, 2025. URLhttps: //arxiv.org/abs/2505.21749

  13. [13]

    Benjamin Erichson, Omri Azencot, Alejandro Queiruga, Liam Hodgkinson, and Michael W

    N. Benjamin Erichson, Omri Azencot, Alejandro Queiruga, Liam Hodgkinson, and Michael W. Mahoney. Lipschitz recurrent neural networks. InInternational Conference on Learning Rep- resentations, 2021. URLhttps://openreview.net/forum?id=-N7PBXqOUJZ

  14. [14]

    Priors in bayesian deep learning: A review.International Statistical Review, 90(3):563–591, 2022

    Vincent Fortuin. Priors in bayesian deep learning: A review.International Statistical Review, 90(3):563–591, 2022. doi: 10.1111/insr.12502

  15. [15]

    Fu, Tri Dao, Khaled K

    Daniel Y . Fu, Tri Dao, Khaled K. Saab, Armin W. Thomas, Atri Rudra, and Christopher Ré. Hungry hungry hippos: Towards language modeling with state space models. InInternational Conference on Learning Representations, 2023. URLhttps://openreview.net/forum? id=COZDy0WYGg

  16. [16]

    Jack Goffinet, Casey Hanks, and David E. Carlson. HiPPO Zoo: Explicit memory mechanisms for interpretable state space models, 2026. URLhttps://arxiv.org/abs/2602.21340

  17. [17]

    Gordon and Marie desJardins

    Diana F. Gordon and Marie desJardins. Evaluation and selection of biases in machine learning. Machine Learning, 20(1–2):5–22, 1995. doi: 10.1023/A:1022630017346

  18. [18]

    Riccardo Grazzi, Julien Siems, Jörg K. H. Franke, Arber Zela, Frank Hutter, and Massimiliano Pontil. Unlocking state-tracking in linear RNNs through negative eigenvalues.arXiv preprint arXiv:2411.12537, 2024. URLhttps://arxiv.org/abs/2411.12537

  19. [19]

    Mamba: Linear-time sequence modeling with selective state spaces

    Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. InThe First Conference on Language Modeling, 2024. URLhttps://openreview.net/ forum?id=tEYskw1VY2

  20. [20]

    HiPPO: Recurrent memory with optimal polynomial projections

    Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré. HiPPO: Recurrent memory with optimal polynomial projections. InAdvances in Neural Information Pro- cessing Systems, 2020. URLhttps://proceedings.neurips.cc/paper/2020/hash/ 102f0bb6efb3a6128a3c750dd16729be-Abstract.html

  21. [21]

    Efficiently modeling long sequences with struc- tured state spaces

    Albert Gu, Karan Goel, and Christopher Ré. Efficiently modeling long sequences with struc- tured state spaces. InInternational Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=uYLFoz1vlAC

  22. [22]

    On the parameterization and initialization of diagonal state space models

    Albert Gu, Ankit Gupta, Karan Goel, and Christopher Ré. On the parameterization and initialization of diagonal state space models. InAdvances in Neural Information Process- ing Systems, 2022. URLhttps://papers.nips.cc/paper_files/paper/2022/hash/ e9a32fade47b906de908431991440f7c-Abstract-Conference.html

  23. [23]

    How to train your HiPPO: State space models with generalized orthogonal basis projections

    Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Ré. How to train your HiPPO: State space models with generalized orthogonal basis projections. InInternational Conference on Learning Representations, 2023. URLhttps://openreview.net/forum? id=klK17OQ3KB

  24. [24]

    Diagonal state spaces are as effective as struc- tured state spaces

    Ankit Gupta, Albert Gu, and Jonathan Berant. Diagonal state spaces are as effective as struc- tured state spaces. InAdvances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=RjS0j6tsSrf

  25. [25]

    Springer, 2 edition,

    Ernst Hairer, Christian Lubich, and Gerhard Wanner.Geometric Numerical Integration: Structure-Preserving Algorithms for Ordinary Differential Equations. Springer, 2 edition,

  26. [26]

    doi: 10.1007/3-540-30666-8

  27. [27]

    Liquid time-constant networks

    Ramin Hasani, Mathias Lechner, Alexander Amini, Daniela Rus, and Radu Grosu. Liquid time-constant networks. InProceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 7657–7666, 2021. URLhttps://ojs.aaai.org/index.php/AAAI/ article/view/16936. 11

  28. [28]

    Liquid structural state-space models

    Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine, Alexander Amini, and Daniela Rus. Liquid structural state-space models. InInternational Conference on Learning Representations, 2023. URLhttps://openreview.net/forum?id=g4OTKRKfS7R

  29. [29]

    Higham.Functions of Matrices: Theory and Computation

    Nicholas J. Higham.Functions of Matrices: Theory and Computation. SIAM, 2008. doi: 10.1137/1.9780898717778

  30. [30]

    Understanding input selectivity in mamba: Impact on approximation power, memorization, and associative recall capacity, 2025

    Ningyuan Huang, Miguel Sarabia, Abhinav Moudgil, Pau Rodriguez, Luca Zappella, and Fed- erico Danieli. Understanding input selectivity in mamba: Impact on approximation power, memorization, and associative recall capacity, 2025. URLhttps://arxiv.org/abs/2506. 11891

  31. [31]

    Anderson Keller, Carmen Amo Alonso, Terrence J

    Arjun Karuvally, Franz Nowak, T. Anderson Keller, Carmen Amo Alonso, Terrence J. Se- jnowski, and Hava T. Siegelmann. Bridging expressivity and scalability with adaptive uni- tary SSMs. InAdvances in Neural Information Processing Systems, 2025. URLhttps: //arxiv.org/abs/2507.05238

  32. [32]

    Neural controlled differen- tial equations for irregular time series

    Patrick Kidger, James Morrill, James Foster, and Terry Lyons. Neural controlled differen- tial equations for irregular time series. InAdvances in Neural Information Processing Sys- tems, volume 33, 2020. URLhttps://proceedings.neurips.cc/paper/2020/hash/ 4a5876b450b45371f6cfe5047ac8cd45-Abstract.html

  33. [33]

    Kimi linear: An expressive, efficient attention architecture.arXiv preprint arXiv:2510.26692, 2025

    Kimi Team, Yu Zhang, Zongyu Lin, Xingcheng Yao, Jiaxi Hu, et al. Kimi linear: An expressive, efficient attention architecture.arXiv preprint arXiv:2510.26692, 2025. URL https://arxiv.org/abs/2510.26692

  34. [34]

    Li, Berlin Chen, Caitlin Wang, Aviv Bick, J

    Aakash Lahoti, Kevin Y . Li, Berlin Chen, Caitlin Wang, Aviv Bick, J. Zico Kolter, Tri Dao, and Albert Gu. Mamba-3: Improved sequence modeling using state space principles. In International Conference on Learning Representations, 2026. URLhttps://openreview. net/forum?id=HwCvaJOiCj

  35. [35]

    UnHiPPO: Uncertainty-aware initialization for state space models

    Marten Lienen, Abdullah Saydemir, and Stephan Günnemann. UnHiPPO: Uncertainty-aware initialization for state space models. InInternational Conference on Machine Learning, 2025. URLhttps://openreview.net/forum?id=U8GUmxnzXn

  36. [36]

    Longhorn: State space models are amortized online learners

    Bo Liu, Rui Wang, Lemeng Wu, Yihao Feng, Peter Stone, and Qiang Liu. Longhorn: State space models are amortized online learners. InInternational Conference on Learning Repre- sentations, 2025. URLhttps://openreview.net/forum?id=8jOqCcLzeO

  37. [37]

    Autocorrelation matters: Understanding the role of initialization schemes for state space models

    Fusheng Liu and Qianxiao Li. Autocorrelation matters: Understanding the role of initialization schemes for state space models. InInternational Conference on Learning Representations,

  38. [38]

    URLhttps://openreview.net/forum?id=sZJNkorXMk

  39. [39]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. InInternational Conference on Learning Representations, 2019. URLhttps://openreview.net/forum? id=Bkg6RiCqY7

  40. [40]

    Parallelizing linear recurrent neural nets over sequence length

    Eric Martin and Chris Cundy. Parallelizing linear recurrent neural nets over sequence length. In International Conference on Learning Representations, 2018. URLhttps://openreview. net/forum?id=HyUNwulC-

  41. [41]

    The illusion of state in state-space models

    William Merrill, Jackson Petty, and Ashish Sabharwal. The illusion of state in state-space models. InProceedings of the 41st International Conference on Machine Learning, volume 235, 2024. URLhttps://proceedings.mlr.press/v235/merrill24a.html

  42. [42]

    M 2RNN: Non- linear RNNs with matrix-valued states for scalable language modeling.arXiv preprint arXiv:2603.14360, 2026

    Mayank Mishra, Shawn Tan, Ion Stoica, Joseph Gonzalez, and Tri Dao. M 2RNN: Non- linear RNNs with matrix-valued states for scalable language modeling.arXiv preprint arXiv:2603.14360, 2026. URLhttps://arxiv.org/abs/2603.14360

  43. [43]

    Mitchell

    Tom M. Mitchell. The need for biases in learning generalizations. Technical Report CBM- TR-117, Department of Computer Science, Rutgers University, 1980. URLhttps://www. cs.cmu.edu/~tom/pubs/NeedForBias_1980.pdf. 12

  44. [44]

    Fixed-point RNNs: Interpolating from diagonal to dense, 2025

    Sajad Movahedi, Felix Sarnthein, Nicola Muca Cirone, and Antonio Orvieto. Fixed-point RNNs: Interpolating from diagonal to dense, 2025. URLhttps://arxiv.org/abs/2503. 10799

  45. [45]

    Roussel Desmond Nzoyem, Nawid Keshtmand, Enrique Crespo Fernandez, Idriss Tsayem, Raul Santos-Rodriguez, David A. W. Barton, and Tom Deakin. Weight-space linear recur- rent neural networks. InInternational Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=zHdKaF3ZM7

  46. [46]

    Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Raz- van Pascanu, and Soham De

    Antonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Raz- van Pascanu, and Soham De. Resurrecting recurrent neural networks for long sequences. InProceedings of the 40th International Conference on Machine Learning, 2023. URL https://proceedings.mlr.press/v202/orvieto23a.html

  47. [47]

    HGRN2: Gated linear RNNs with state expansion

    Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong. HGRN2: Gated linear RNNs with state expansion. InConference on Language Modeling,

  48. [48]

    URLhttps://openreview.net/forum?id=y6SqbJfCSk

  49. [49]

    Yulia Rubanova, Ricky T. Q. Chen, and David K. Duvenaud. Latent ordinary differen- tial equations for irregularly-sampled time series. InAdvances in Neural Information Processing Systems, volume 32, 2019. URLhttps://papers.neurips.cc/paper/ 8773-latent-ordinary-differential-equations-for-irregularly-sampled-time-series

  50. [50]

    Konstantin Rusch and Siddhartha Mishra

    T. Konstantin Rusch and Siddhartha Mishra. Coupled oscillatory recurrent neural network (coRNN): An accurate and (gradient) stable architecture for learning long time dependen- cies. InInternational Conference on Learning Representations, 2021. URLhttps:// openreview.net/forum?id=F3s69XzWOia

  51. [51]

    Konstantin Rusch and Daniela Rus

    T. Konstantin Rusch and Daniela Rus. Oscillatory state-space models. InInternational Con- ference on Learning Representations, 2025. URLhttps://openreview.net/forum?id= GRMfXcAAFh. Oral presentation

  52. [52]

    The expressive capacity of state space models: A formal language perspective

    Yash Sarrof, Yana Veitsman, and Michael Hahn. The expressive capacity of state space models: A formal language perspective. InAdvances in Neural Information Processing Systems, 2024. URLhttps://openreview.net/forum?id=eV5YIrJPdy

  53. [53]

    Linear transformers are secretly fast weight programmers

    Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. Linear transformers are secretly fast weight programmers. InProceedings of the 38th International Conference on Machine Learn- ing, volume 139, pages 9355–9366, 2021. URLhttps://proceedings.mlr.press/v139/ schlag21a.html

  54. [54]

    The ex- pressive limits of diagonal SSMs for state-tracking

    Mehran Shakerinava, Behnoush Khavari, Siamak Ravanbakhsh, and Sarath Chandar. The ex- pressive limits of diagonal SSMs for state-tracking. InInternational Conference on Learning Representations, 2026. URLhttps://openreview.net/forum?id=5bg5Ru5OML

  55. [55]

    Nir Shlezinger and Yonina C. Eldar. Model-based deep learning.Foundations and Trends in Signal Processing, 17(4):291–416, 2023. doi: 10.1561/2000000113

  56. [56]

    Zeilinger, and Antonio Orvieto

    Jerome Sieber, Carmen Amo Alonso, Alexandre Didier, Melanie N. Zeilinger, and Antonio Orvieto. Understanding the differences in foundation models: Attention, state space models, and recurrent neural networks. InAdvances in Neural Information Processing Systems, 2024. URLhttps://arxiv.org/abs/2405.15731

  57. [57]

    Zeilinger, and Carmen Amo Alonso

    Jerome Sieber, Antonio Orvieto, Melanie N. Zeilinger, and Carmen Amo Alonso. Design principles for sequence models via coefficient dynamics, 2025. URLhttps://arxiv.org/ abs/2510.09389

  58. [58]

    Deltaproduct: Increasing the expressivity of deltanet through products of householders

    Julien Siems, Timur Carstensen, Arber Zela, Frank Hutter, Massimiliano Pontil, and Riccardo Grazzi. Deltaproduct: Increasing the expressivity of deltanet through products of householders. arXiv preprint arXiv:2502.10297, 2025. URLhttps://arxiv.org/abs/2502.10297

  59. [59]

    Computational methods for linear matrix equations.SIAM Review, 58(3): 377–441, 2016

    Valeria Simoncini. Computational methods for linear matrix equations.SIAM Review, 58(3): 377–441, 2016. doi: 10.1137/130912839. 13

  60. [60]

    Jimmy T. H. Smith, Andrew Warrington, and Scott W. Linderman. Simplified state space layers for sequence modeling. InInternational Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=Ai8Hw3AXqks

  61. [61]

    Un- covering the spectral bias in diagonal state space models

    Ruben Solozabal, Velibor Bojkovic, Hilal AlQuabeh, Kentaro Inui, and Martin Taká ˇc. Un- covering the spectral bias in diagonal state space models. InAdvances in Neural Information Processing Systems, 2025. URLhttps://arxiv.org/abs/2508.20441

  62. [62]

    Learning to (learn at test time): RNNs with expressive hidden states

    Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xin- lei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, and Carlos Guestrin. Learning to (learn at test time): RNNs with expressive hidden states. InProceedings of the 42nd Inter- national Conference on Machine Learning, 2025. URLhttps://openreview.net/forum? id=...

  63. [63]

    On the expressiveness and length generalization of selective state- space models on regular languages

    Aleksandar Terzi ´c, Michael Hersche, Giacomo Camposampiero, Thomas Hofmann, Abu Se- bastian, and Abbas Rahimi. On the expressiveness and length generalization of selective state- space models on regular languages. InProceedings of the AAAI Conference on Artificial Intel- ligence, 2025. URLhttps://arxiv.org/abs/2412.19350

  64. [64]

    Structured sparse transition matrices to enable state tracking in state-space models

    Aleksandar Terzi ´c, Nicolas Menet, Michael Hersche, Thomas Hofmann, and Abbas Rahimi. Structured sparse transition matrices to enable state tracking in state-space models. InAd- vances in Neural Information Processing Systems, 2025. URLhttps://openreview.net/ forum?id=RDbuSCWhad

  65. [65]

    Zico Kolter, Sanjiv Kumar, and Srinadh Bhojanapalli

    Asher Trockman, Hrayr Harutyunyan, J. Zico Kolter, Sanjiv Kumar, and Srinadh Bhojanapalli. Mimetic initialization helps state space models learn to recall, 2024. URLhttps://arxiv. org/abs/2410.11135. Presented at the ICLR 2025 Workshop on Weight Space Learning

  66. [66]

    On the implicit bias in deep-learning algorithms.Communications of the ACM, 66 (6):86–93, 2023

    Gal Vardi. On the implicit bias in deep-learning algorithms.Communications of the ACM, 66 (6):86–93, 2023. doi: 10.1145/3571070

  67. [67]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Infor- mation Processing Systems, volume 30, 2017. URLhttps://papers.neurips.cc/paper/ 7181-attention-is-all-you-need

  68. [68]

    V oelker, Ivana Kaji´c, and Chris Eliasmith

    Aaron R. V oelker, Ivana Kaji´c, and Chris Eliasmith. Legendre memory units: Continuous-time representation in recurrent neural networks. InAdvances in Neural Information Processing Systems, volume 32, 2019. URLhttps://proceedings.neurips.cc/paper/2019/hash/ 952285b9b7e7a1be5aa7849f32ffff05-Abstract.html

  69. [69]

    Saurous, Charlotte Frenkel, Razvan Pascanu, Blaise Aguera y Arcas, and Joao Sacramento

    Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Max- imilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Guillaume Lajoie, Rif A. Saurous, Charlotte Frenkel, Razvan Pascanu, Blaise Aguera y Arcas, and Joao Sacramento. MesaNet: Sequence modeling by locally optimal test-time training. In...

  70. [70]

    Informed ma- chine learning—a taxonomy and survey of integrating prior knowledge into learning sys- tems.IEEE Transactions on Knowledge and Data Engineering, 35(1):614–633, 2023

    Laura von Rueden, Sebastian Mayer, Katharina Beckh, Bogdan Georgiev, Sven Giessel- bach, Raoul Heese, Birgit Kirsch, Julius Pfrommer, Annika Pick, Rajkumar Ramamurthy, Michał Walczak, Jochen Garcke, Christian Bauckhage, and Jannis Schuecker. Informed ma- chine learning—a taxonomy and survey of integrating prior knowledge into learning sys- tems.IEEE Trans...

  71. [71]

    Struc- tured linear CDEs: Maximally expressive and parallel-in-time sequence models

    Benjamin Walker, Lingyi Yang, Nicola Muca Cirone, Cristopher Salvi, and Terry Lyons. Struc- tured linear CDEs: Maximally expressive and parallel-in-time sequence models. InAdvances in Neural Information Processing Systems, 2025. URLhttps://openreview.net/forum? id=HKDyRDzy1E

  72. [72]

    StableSSM: Alleviating the curse of memory in state-space mod- els through stable reparameterization

    Shida Wang and Qianxiao Li. StableSSM: Alleviating the curse of memory in state-space mod- els through stable reparameterization. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Machine Learning Research, pages 50766– 50793. PMLR, 2024. URLhttps://proceedings.mlr.press/v235/wang24ag.html. 14

  73. [73]

    State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memory

    Shida Wang and Beichen Xue. State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memory. InAdvances in Neural Information Pro- cessing Systems, 2023. URLhttps://proceedings.neurips.cc/paper_files/paper/ 2023/hash/ea8608c6258450e75b3443ec8022fb2e-Abstract-Conference.html

  74. [74]

    Gated linear attention transformers with hardware-efficient training

    Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. Gated linear attention transformers with hardware-efficient training. InProceedings of the 41st Interna- tional Conference on Machine Learning, 2024. URLhttps://openreview.net/forum? id=ia5XvxFUJT

  75. [75]

    Parallelizing linear transformers with the delta rule over sequence length

    Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. Parallelizing linear transformers with the delta rule over sequence length. InAdvances in Neural Information Processing Systems, 2024. URLhttps://openreview.net/forum?id=y8Rm4VNRPH

  76. [76]

    Gated delta networks: Improving mamba2 with delta rule

    Songlin Yang, Jan Kautz, and Ali Hatamizadeh. Gated delta networks: Improving mamba2 with delta rule. InInternational Conference on Learning Representations, 2025. URLhttps: //openreview.net/forum?id=r8H7xhYPwz

  77. [77]

    Mahoney, and N

    Annan Yu, Arnur Nigmetov, Dmitriy Morozov, Michael W. Mahoney, and N. Benjamin Erichson. Robustifying state-space models for long sequences via approximate diagonal- ization. InInternational Conference on Learning Representations, 2024. URLhttps: //openreview.net/forum?id=DjeQ39QoLQ

  78. [78]

    Mahoney, and N

    Annan Yu, Dongwei Lyu, Soon Hoe Lim, Michael W. Mahoney, and N. Benjamin Erichson. Tuning frequency bias of state space models. InInternational Conference on Learning Repre- sentations, 2025. URLhttps://openreview.net/forum?id=wkHcXDv7cv

  79. [79]

    Mahoney, and N

    Annan Yu, Michael W. Mahoney, and N. Benjamin Erichson. HOPE for a robust parameter- ization of long-memory state space models. InInternational Conference on Learning Repre- sentations, 2025. URLhttps://openreview.net/forum?id=RZwtbg3qYD

  80. [80]

    Explaining modern gated-linear RNNs via a unified implicit attention formulation, 2024

    Itamar Zimerman, Ameen Ali, and Lior Wolf. Explaining modern gated-linear RNNs via a unified implicit attention formulation, 2024. URLhttps://arxiv.org/abs/2405.16504. 15 A Derivations for Section 2 This appendix supports the operator-level claims of Section 2. We first recall the ordinary vector- valued HiPPO variational equations, then derive the SHiPPO...

Showing first 80 references.