REVIEW 2 major objections 4 minor 83 references
SHiPPO makes channel interaction part of online polynomial memory itself: stored HiPPO coefficients live in a moving frame and obey Sylvester dynamics driven by a right-transport path.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 05:11 UTC pith:W7T2L4IB
load-bearing objection Clean HiPPO lift to Sylvester transported memory with honest mechanism diagnostics; group-local scan is a real restriction, not a free lunch. the 2 major comments →
SHiPPO: Recurrent Memory with Transported Polynomial Projections
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Coupling a transported approximation family with a transported channel metric yields Sylvester coefficient dynamics ˙C = A_L C + B_L f^⊤ + C A_R. Conditional on any realized right-transport path, the state is ordinary HiPPO in a tied moving channel frame: the left operator is inherited from the online projection problem while the right action is external path transport of already-written memory coordinates.
What carries the argument
The SHiPPO online approximation problem (Definition 2.1) and its Sylvester dynamics (Theorem 2.2): jointly transporting family and metric produces the right-action gauge term C A_R on stored coefficients, with a pathwise lift of any closed one-sided HiPPO equation (Corollary 2.3) and a scan-compatible group-local realization for selective SSMs.
Load-bearing premise
That a restricted group-local, controller-compatible right action—kept independent of the main memory state so the block-affine scan stays exact—still carries the operator-level transported-memory meaning the theory assigns to a general right path.
What would settle it
On the paired noncommutative diagnostic, train a transported model that recovers Pair ΔNMSE near zero, then at evaluation replace every right action R_t by the identity while freezing all other weights; if the paired-difference signal does not return to near one (as reported), the recovered signal is not mediated by future right transport of already-written memory.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SHiPPO, a transported online-projection memory prior that lifts HiPPO coefficient dynamics into a moving channel frame. Conditional on any fixed or realized right-transport path, the approximation family and channel metric are transported together, so the coefficient state is ordinary HiPPO in a tied moving frame and obeys the Sylvester ODE ˙C = A_L C + B_L f^⊤ + C A_R (Theorem 2.2, Corollary 2.3). For selective-SSM execution the authors derive a restricted group-local, controller-compatible cell with exponential-adjusted updates and exact block-affine scan, plus a simultaneous-reducibility collapse criterion for right transports. Controlled paired noncommutative diagnostics separate high-rank current-token writes from future right action on already-written memory; transported variants recover the order-sensitive signal, which vanishes under an evaluation-time R_t → I intervention. Transport-MQAR supplies complementary autoregressive evidence while leaving the preferred right-action realization open. Claims are scoped to memory mechanisms rather than broad sequence-modeling dominance.
Significance. If the pathwise lift and the write-versus-transport separation hold—as the derivations and interventions indicate—this is a genuine contribution to inductive-bias design for recurrent memory: channel interaction is made part of the online-approximation semantics rather than an architectural mixer. Strengths include complete normal-equation, conjugacy, gauge, scan-closure, and discretization arguments (Appendices A–C), an explicit collapse criterion (Proposition 3.3), and falsifiable diagnostics with seed-level means/stds and a clean R→I ablation. The work is carefully scoped and does not overclaim empirical dominance. Residual open questions (preferred right-action backend; group-local fidelity) are largely acknowledged and do not erase the operator-level contribution.
major comments (2)
- [Section 3; Appendix D.6, Table 6] Section 3 presents the group-local, controller-compatible cell as the practical SHiPPO lift for selective SSMs, yet Appendix D.6 shows that oracle-R with group size 2 yields PairΔNMSE = 1 (complete failure) while group size 4 (full transport for d=4) succeeds. The formal separation claim is stated for full transport, but contribution (iii) and the selective-SSM narrative rest on the restricted cell. The manuscript should either (a) demonstrate a group-local design that recovers the paired-difference signal without relying on front-end/readout recoding, or (b) demote the group-local cell more explicitly to a computational approximation with quantified fidelity loss relative to the operator-level prior.
- [Section 4.3; Table 1; Appendix E.5] Table 1 and Appendix E.4–E.5: Transport-MQAR gains over no-right/static-basis controls are modest (e.g., coordinate accuracy 0.110 vs 0.102 at length 4096), exact-accuracy leadership flips between DirectGen-SingleExp and StructGen-Split, and the suffix-zeroing counterfactual only shows that controller coordinates are used—not that a noncommutative transport geometry was learned. As complementary evidence this is acceptable under the paper’s scope, but the main-text framing that these results support “learned right-action pathways” should be tightened to match the weaker identification actually achieved.
minor comments (4)
- [Figure 1] Figure 1 is useful but dense; a short caption callout distinguishing the abstract operator (left) from the scan-compatible restriction (right) would help readers who skip Section 3.
- [Section 3.4, Eq. (3)] Notation for the discrete cell mixes L_t, R_t, bU_t and λ_t; a one-line glossary at the start of Section 3.4 would reduce back-references to (1)–(3).
- [Section 5; Appendix D.7] Appendix D.7 notes excluded exploratory diagnostics; a single sentence in the main Discussion pointing readers there would improve transparency without expanding the main text.
- [References] Several related-work citations are 2025–2026 arXiv preprints; ensure final versions or stable identifiers are used at camera-ready if available.
Circularity Check
No significant circularity: Sylvester dynamics are derived from an explicit transported projection objective, and diagnostics use held-out interventions rather than fitted-as-prediction.
full rationale
The load-bearing theoretical chain is Definition 2.1 (transported approximation family + coupled channel metric) → normal equation (Prop. A.3 / Thm. 2.2) → Sylvester coefficient ODE by Leibniz differentiation under HiPPO closure. That is a standard variational derivation, not a recurrence postulated and then re-labeled as projection memory. Corollary 2.3 is a pathwise lift of any closed one-sided HiPPO-style equation and is explicitly conditional on an external right-transport path; it does not smuggle the target dynamics into the definition of the left operator. The scan-compatible cell (Sec. 3) is openly a restricted realization (group-tied diagonal left, controller-compatible right) with proved block-affine scan closure, not a claim that the restriction is forced by uniqueness. Empirically, the paired noncommutative diagnostic and the evaluation-time R_t → I intervention freeze other weights and remove the transport pathway; the paired-difference signal disappears, so the result is not forced by construction or by a fitted normalization. No self-citation uniqueness theorem, no fitted parameter renamed as prediction, and no mere renaming of a known empirical pattern under new coordinates. Residual open questions (preferred right-action realization, group-local fidelity) are scoped limitations, not circular reductions.
Axiom & Free-Parameter Ledger
free parameters (4)
- group width P / number of groups G
- source rank r and write-rank baselines
- right-generator library (skew, nilpotent, low-rank, diagonal)
- discretization weights λ_t and step sizes Δ_t
axioms (4)
- domain assumption HiPPO closure condition ∂_t ψ = A_L ψ (τ < t) and invertible Gram matrix G(t)
- standard math A_R is integrable so the state-transition family P(t,τ) exists and lies in GL(d)
- ad hoc to paper Right transport is controller-compatible (independent of main memory H) so the step map remains affine and the block-affine scan algebra closes
- standard math Simultaneous block-reducibility of the generator family implies collapse to static mixing plus independent banks
invented entities (3)
-
SHiPPO online approximation problem (transported family G_SH_t + metric M_P)
independent evidence
-
Scan-compatible group-local SHiPPO cell with exponential-adjusted updates
no independent evidence
-
Paired noncommutative transport diagnostic and Transport-MQAR
no independent evidence
read the original abstract
HiPPO gives recurrent states memory semantics as coefficients of online polynomial projections, but in fixed channel coordinates. Modern selective SSMs, by contrast, rely on token-dependent control and channel interaction. We introduce SHiPPO (Sylvester HiPPO), a transported projection-memory prior that lifts HiPPO coefficient memories into a moving channel frame. For any fixed or realized right-transport path, SHiPPO transports the approximation family and channel metric together; conditional on that path, the state is ordinary HiPPO in a tied moving frame and follows Sylvester coefficient dynamics, preserving the left online-memory operator while adding right-action transport. For selective-SSM execution, we derive a restricted group-local realization with controller-compatible right actions, exponential-adjusted updates, exact block-affine scan, and recurrent decoding. We also give a simultaneous-reducibility criterion identifying when right transports collapse to static mixing plus independent scalar or blockwise banks. Controlled diagnostics show that larger current-token write rank improves ordinary prediction error but cannot recover order-sensitive changes to already-written memory; transported-memory variants recover this signal, which disappears when the transport pathway is removed. A finite-field associative-recall diagnostic with interleaved bindings, operations, and queries provides complementary autoregressive evidence while leaving the preferred right-action realization open. Taken together, these results support SHiPPO as a mechanistically grounded transported-memory prior, with evidence focused on memory mechanisms rather than broad sequence-modeling dominance.
Figures
Reference graph
Works this paper leans on
-
[1]
On the expressiveness of state space models via temporal logics, 2026
Eric Alsmann, Lowejatan Noori, and Martin Lange. On the expressiveness of state space models via temporal logics, 2026. URLhttps://arxiv.org/abs/2601.19467
arXiv 2026
-
[2]
Zoology: Measuring and improving recall in efficient language models
Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, and Christopher Ré. Zoology: Measuring and improving recall in efficient language models. InInternational Conference on Learning Representations, 2024. URLhttps:// openreview.net/forum?id=LY3ukUANko
2024
-
[3]
xL- STM: Extended long short-term memory
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prud- nikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. xL- STM: Extended long short-term memory. InAdvances in Neural Information Processing Sys- tems, 2024. URLhttps://openreview.net/forum?id=ARAxPPIAhq
2024
-
[4]
Solution formulas for differential sylvester and lyapunov equations.Calcolo, 56:51, 2019
Maximilian Behr, Peter Benner, and Jan Heiland. Solution formulas for differential sylvester and lyapunov equations.Calcolo, 56:51, 2019. doi: 10.1007/s10092-019-0348-x
-
[5]
Blelloch
Guy E. Blelloch. Prefix sums and their applications. Technical Report CMU-CS-90-190, School of Computer Science, Carnegie Mellon University, 1990. URLhttps://www.cs. cmu.edu/~scandal/papers/CMU-CS-90-190.html
1990
-
[6]
Bo Chang, Minmin Chen, Eldad Haber, and Ed H. Chi. AntisymmetricRNN: A dynamical system view on recurrent neural networks. InInternational Conference on Learning Represen- tations, 2019. URLhttps://openreview.net/forum?id=ryxepo0cFX
2019
-
[7]
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neu- ral ordinary differential equations. InAdvances in Neural Information Process- ing Systems, volume 31, 2018. URLhttps://papers.neurips.cc/paper/ 7892-neural-ordinary-differential-equations
2018
-
[8]
Learning phrase representations using RNN encoder– decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder– decoder for statistical machine translation. InProceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1724–1734. Asso- ciation for Compu...
-
[9]
Nicola Muca Cirone, Antonio Orvieto, Benjamin Walker, Cristopher Salvi, and Terry J. Lyons. Theoretical foundations of deep selective state-space models. InAdvances in Neural Informa- tion Processing Systems, 2024. URLhttps://openreview.net/forum?id=3SzrqwupUx
2024
-
[10]
Coddington and Norman Levinson.Theory of Ordinary Differential Equations
Earl A. Coddington and Norman Levinson.Theory of Ordinary Differential Equations. McGraw-Hill, New York, 1955
1955
-
[11]
Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality
Tri Dao and Albert Gu. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Machine Learning Research, pages 10041– 10071. PMLR, 2024. URLhttps://proceedings.mlr.press/v235/dao24a.html. 10
2024
-
[12]
Reza Ebrahimi and Roland Memisevic
M. Reza Ebrahimi and Roland Memisevic. Revisiting bi-linear state transitions in recurrent neural networks. InAdvances in Neural Information Processing Systems, 2025. URLhttps: //arxiv.org/abs/2505.21749
arXiv 2025
-
[13]
Benjamin Erichson, Omri Azencot, Alejandro Queiruga, Liam Hodgkinson, and Michael W
N. Benjamin Erichson, Omri Azencot, Alejandro Queiruga, Liam Hodgkinson, and Michael W. Mahoney. Lipschitz recurrent neural networks. InInternational Conference on Learning Rep- resentations, 2021. URLhttps://openreview.net/forum?id=-N7PBXqOUJZ
2021
-
[14]
Priors in bayesian deep learning: A review.International Statistical Review, 90(3):563–591, 2022
Vincent Fortuin. Priors in bayesian deep learning: A review.International Statistical Review, 90(3):563–591, 2022. doi: 10.1111/insr.12502
-
[15]
Fu, Tri Dao, Khaled K
Daniel Y . Fu, Tri Dao, Khaled K. Saab, Armin W. Thomas, Atri Rudra, and Christopher Ré. Hungry hungry hippos: Towards language modeling with state space models. InInternational Conference on Learning Representations, 2023. URLhttps://openreview.net/forum? id=COZDy0WYGg
2023
-
[16]
Jack Goffinet, Casey Hanks, and David E. Carlson. HiPPO Zoo: Explicit memory mechanisms for interpretable state space models, 2026. URLhttps://arxiv.org/abs/2602.21340
Pith/arXiv arXiv 2026
-
[17]
Diana F. Gordon and Marie desJardins. Evaluation and selection of biases in machine learning. Machine Learning, 20(1–2):5–22, 1995. doi: 10.1023/A:1022630017346
-
[18]
Riccardo Grazzi, Julien Siems, Jörg K. H. Franke, Arber Zela, Frank Hutter, and Massimiliano Pontil. Unlocking state-tracking in linear RNNs through negative eigenvalues.arXiv preprint arXiv:2411.12537, 2024. URLhttps://arxiv.org/abs/2411.12537
Pith/arXiv arXiv 2024
-
[19]
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. InThe First Conference on Language Modeling, 2024. URLhttps://openreview.net/ forum?id=tEYskw1VY2
2024
-
[20]
HiPPO: Recurrent memory with optimal polynomial projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré. HiPPO: Recurrent memory with optimal polynomial projections. InAdvances in Neural Information Pro- cessing Systems, 2020. URLhttps://proceedings.neurips.cc/paper/2020/hash/ 102f0bb6efb3a6128a3c750dd16729be-Abstract.html
2020
-
[21]
Efficiently modeling long sequences with struc- tured state spaces
Albert Gu, Karan Goel, and Christopher Ré. Efficiently modeling long sequences with struc- tured state spaces. InInternational Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=uYLFoz1vlAC
2022
-
[22]
On the parameterization and initialization of diagonal state space models
Albert Gu, Ankit Gupta, Karan Goel, and Christopher Ré. On the parameterization and initialization of diagonal state space models. InAdvances in Neural Information Process- ing Systems, 2022. URLhttps://papers.nips.cc/paper_files/paper/2022/hash/ e9a32fade47b906de908431991440f7c-Abstract-Conference.html
2022
-
[23]
How to train your HiPPO: State space models with generalized orthogonal basis projections
Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Ré. How to train your HiPPO: State space models with generalized orthogonal basis projections. InInternational Conference on Learning Representations, 2023. URLhttps://openreview.net/forum? id=klK17OQ3KB
2023
-
[24]
Diagonal state spaces are as effective as struc- tured state spaces
Ankit Gupta, Albert Gu, and Jonathan Berant. Diagonal state spaces are as effective as struc- tured state spaces. InAdvances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=RjS0j6tsSrf
2022
-
[25]
Springer, 2 edition,
Ernst Hairer, Christian Lubich, and Gerhard Wanner.Geometric Numerical Integration: Structure-Preserving Algorithms for Ordinary Differential Equations. Springer, 2 edition,
-
[26]
doi: 10.1007/3-540-30666-8
-
[27]
Liquid time-constant networks
Ramin Hasani, Mathias Lechner, Alexander Amini, Daniela Rus, and Radu Grosu. Liquid time-constant networks. InProceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 7657–7666, 2021. URLhttps://ojs.aaai.org/index.php/AAAI/ article/view/16936. 11
2021
-
[28]
Liquid structural state-space models
Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine, Alexander Amini, and Daniela Rus. Liquid structural state-space models. InInternational Conference on Learning Representations, 2023. URLhttps://openreview.net/forum?id=g4OTKRKfS7R
2023
-
[29]
Higham.Functions of Matrices: Theory and Computation
Nicholas J. Higham.Functions of Matrices: Theory and Computation. SIAM, 2008. doi: 10.1137/1.9780898717778
-
[30]
Understanding input selectivity in mamba: Impact on approximation power, memorization, and associative recall capacity, 2025
Ningyuan Huang, Miguel Sarabia, Abhinav Moudgil, Pau Rodriguez, Luca Zappella, and Fed- erico Danieli. Understanding input selectivity in mamba: Impact on approximation power, memorization, and associative recall capacity, 2025. URLhttps://arxiv.org/abs/2506. 11891
2025
-
[31]
Anderson Keller, Carmen Amo Alonso, Terrence J
Arjun Karuvally, Franz Nowak, T. Anderson Keller, Carmen Amo Alonso, Terrence J. Se- jnowski, and Hava T. Siegelmann. Bridging expressivity and scalability with adaptive uni- tary SSMs. InAdvances in Neural Information Processing Systems, 2025. URLhttps: //arxiv.org/abs/2507.05238
arXiv 2025
-
[32]
Neural controlled differen- tial equations for irregular time series
Patrick Kidger, James Morrill, James Foster, and Terry Lyons. Neural controlled differen- tial equations for irregular time series. InAdvances in Neural Information Processing Sys- tems, volume 33, 2020. URLhttps://proceedings.neurips.cc/paper/2020/hash/ 4a5876b450b45371f6cfe5047ac8cd45-Abstract.html
2020
-
[33]
Kimi linear: An expressive, efficient attention architecture.arXiv preprint arXiv:2510.26692, 2025
Kimi Team, Yu Zhang, Zongyu Lin, Xingcheng Yao, Jiaxi Hu, et al. Kimi linear: An expressive, efficient attention architecture.arXiv preprint arXiv:2510.26692, 2025. URL https://arxiv.org/abs/2510.26692
Pith/arXiv arXiv 2025
-
[34]
Li, Berlin Chen, Caitlin Wang, Aviv Bick, J
Aakash Lahoti, Kevin Y . Li, Berlin Chen, Caitlin Wang, Aviv Bick, J. Zico Kolter, Tri Dao, and Albert Gu. Mamba-3: Improved sequence modeling using state space principles. In International Conference on Learning Representations, 2026. URLhttps://openreview. net/forum?id=HwCvaJOiCj
2026
-
[35]
UnHiPPO: Uncertainty-aware initialization for state space models
Marten Lienen, Abdullah Saydemir, and Stephan Günnemann. UnHiPPO: Uncertainty-aware initialization for state space models. InInternational Conference on Machine Learning, 2025. URLhttps://openreview.net/forum?id=U8GUmxnzXn
2025
-
[36]
Longhorn: State space models are amortized online learners
Bo Liu, Rui Wang, Lemeng Wu, Yihao Feng, Peter Stone, and Qiang Liu. Longhorn: State space models are amortized online learners. InInternational Conference on Learning Repre- sentations, 2025. URLhttps://openreview.net/forum?id=8jOqCcLzeO
2025
-
[37]
Autocorrelation matters: Understanding the role of initialization schemes for state space models
Fusheng Liu and Qianxiao Li. Autocorrelation matters: Understanding the role of initialization schemes for state space models. InInternational Conference on Learning Representations,
-
[38]
URLhttps://openreview.net/forum?id=sZJNkorXMk
-
[39]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. InInternational Conference on Learning Representations, 2019. URLhttps://openreview.net/forum? id=Bkg6RiCqY7
2019
-
[40]
Parallelizing linear recurrent neural nets over sequence length
Eric Martin and Chris Cundy. Parallelizing linear recurrent neural nets over sequence length. In International Conference on Learning Representations, 2018. URLhttps://openreview. net/forum?id=HyUNwulC-
2018
-
[41]
The illusion of state in state-space models
William Merrill, Jackson Petty, and Ashish Sabharwal. The illusion of state in state-space models. InProceedings of the 41st International Conference on Machine Learning, volume 235, 2024. URLhttps://proceedings.mlr.press/v235/merrill24a.html
2024
-
[42]
Mayank Mishra, Shawn Tan, Ion Stoica, Joseph Gonzalez, and Tri Dao. M 2RNN: Non- linear RNNs with matrix-valued states for scalable language modeling.arXiv preprint arXiv:2603.14360, 2026. URLhttps://arxiv.org/abs/2603.14360
Pith/arXiv arXiv 2026
-
[43]
Mitchell
Tom M. Mitchell. The need for biases in learning generalizations. Technical Report CBM- TR-117, Department of Computer Science, Rutgers University, 1980. URLhttps://www. cs.cmu.edu/~tom/pubs/NeedForBias_1980.pdf. 12
1980
-
[44]
Fixed-point RNNs: Interpolating from diagonal to dense, 2025
Sajad Movahedi, Felix Sarnthein, Nicola Muca Cirone, and Antonio Orvieto. Fixed-point RNNs: Interpolating from diagonal to dense, 2025. URLhttps://arxiv.org/abs/2503. 10799
2025
-
[45]
Roussel Desmond Nzoyem, Nawid Keshtmand, Enrique Crespo Fernandez, Idriss Tsayem, Raul Santos-Rodriguez, David A. W. Barton, and Tom Deakin. Weight-space linear recur- rent neural networks. InInternational Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=zHdKaF3ZM7
2026
-
[46]
Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Raz- van Pascanu, and Soham De
Antonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Raz- van Pascanu, and Soham De. Resurrecting recurrent neural networks for long sequences. InProceedings of the 40th International Conference on Machine Learning, 2023. URL https://proceedings.mlr.press/v202/orvieto23a.html
2023
-
[47]
HGRN2: Gated linear RNNs with state expansion
Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong. HGRN2: Gated linear RNNs with state expansion. InConference on Language Modeling,
-
[48]
URLhttps://openreview.net/forum?id=y6SqbJfCSk
-
[49]
Yulia Rubanova, Ricky T. Q. Chen, and David K. Duvenaud. Latent ordinary differen- tial equations for irregularly-sampled time series. InAdvances in Neural Information Processing Systems, volume 32, 2019. URLhttps://papers.neurips.cc/paper/ 8773-latent-ordinary-differential-equations-for-irregularly-sampled-time-series
2019
-
[50]
Konstantin Rusch and Siddhartha Mishra
T. Konstantin Rusch and Siddhartha Mishra. Coupled oscillatory recurrent neural network (coRNN): An accurate and (gradient) stable architecture for learning long time dependen- cies. InInternational Conference on Learning Representations, 2021. URLhttps:// openreview.net/forum?id=F3s69XzWOia
2021
-
[51]
Konstantin Rusch and Daniela Rus
T. Konstantin Rusch and Daniela Rus. Oscillatory state-space models. InInternational Con- ference on Learning Representations, 2025. URLhttps://openreview.net/forum?id= GRMfXcAAFh. Oral presentation
2025
-
[52]
The expressive capacity of state space models: A formal language perspective
Yash Sarrof, Yana Veitsman, and Michael Hahn. The expressive capacity of state space models: A formal language perspective. InAdvances in Neural Information Processing Systems, 2024. URLhttps://openreview.net/forum?id=eV5YIrJPdy
2024
-
[53]
Linear transformers are secretly fast weight programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. Linear transformers are secretly fast weight programmers. InProceedings of the 38th International Conference on Machine Learn- ing, volume 139, pages 9355–9366, 2021. URLhttps://proceedings.mlr.press/v139/ schlag21a.html
2021
-
[54]
The ex- pressive limits of diagonal SSMs for state-tracking
Mehran Shakerinava, Behnoush Khavari, Siamak Ravanbakhsh, and Sarath Chandar. The ex- pressive limits of diagonal SSMs for state-tracking. InInternational Conference on Learning Representations, 2026. URLhttps://openreview.net/forum?id=5bg5Ru5OML
2026
-
[55]
Nir Shlezinger and Yonina C. Eldar. Model-based deep learning.Foundations and Trends in Signal Processing, 17(4):291–416, 2023. doi: 10.1561/2000000113
-
[56]
Zeilinger, and Antonio Orvieto
Jerome Sieber, Carmen Amo Alonso, Alexandre Didier, Melanie N. Zeilinger, and Antonio Orvieto. Understanding the differences in foundation models: Attention, state space models, and recurrent neural networks. InAdvances in Neural Information Processing Systems, 2024. URLhttps://arxiv.org/abs/2405.15731
Pith/arXiv arXiv 2024
-
[57]
Zeilinger, and Carmen Amo Alonso
Jerome Sieber, Antonio Orvieto, Melanie N. Zeilinger, and Carmen Amo Alonso. Design principles for sequence models via coefficient dynamics, 2025. URLhttps://arxiv.org/ abs/2510.09389
Pith/arXiv arXiv 2025
-
[58]
Deltaproduct: Increasing the expressivity of deltanet through products of householders
Julien Siems, Timur Carstensen, Arber Zela, Frank Hutter, Massimiliano Pontil, and Riccardo Grazzi. Deltaproduct: Increasing the expressivity of deltanet through products of householders. arXiv preprint arXiv:2502.10297, 2025. URLhttps://arxiv.org/abs/2502.10297
arXiv 2025
-
[59]
Computational methods for linear matrix equations.SIAM Review, 58(3): 377–441, 2016
Valeria Simoncini. Computational methods for linear matrix equations.SIAM Review, 58(3): 377–441, 2016. doi: 10.1137/130912839. 13
-
[60]
Jimmy T. H. Smith, Andrew Warrington, and Scott W. Linderman. Simplified state space layers for sequence modeling. InInternational Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=Ai8Hw3AXqks
2023
-
[61]
Un- covering the spectral bias in diagonal state space models
Ruben Solozabal, Velibor Bojkovic, Hilal AlQuabeh, Kentaro Inui, and Martin Taká ˇc. Un- covering the spectral bias in diagonal state space models. InAdvances in Neural Information Processing Systems, 2025. URLhttps://arxiv.org/abs/2508.20441
Pith/arXiv arXiv 2025
-
[62]
Learning to (learn at test time): RNNs with expressive hidden states
Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xin- lei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, and Carlos Guestrin. Learning to (learn at test time): RNNs with expressive hidden states. InProceedings of the 42nd Inter- national Conference on Machine Learning, 2025. URLhttps://openreview.net/forum? id=...
2025
-
[63]
Aleksandar Terzi ´c, Michael Hersche, Giacomo Camposampiero, Thomas Hofmann, Abu Se- bastian, and Abbas Rahimi. On the expressiveness and length generalization of selective state- space models on regular languages. InProceedings of the AAAI Conference on Artificial Intel- ligence, 2025. URLhttps://arxiv.org/abs/2412.19350
Pith/arXiv arXiv 2025
-
[64]
Structured sparse transition matrices to enable state tracking in state-space models
Aleksandar Terzi ´c, Nicolas Menet, Michael Hersche, Thomas Hofmann, and Abbas Rahimi. Structured sparse transition matrices to enable state tracking in state-space models. InAd- vances in Neural Information Processing Systems, 2025. URLhttps://openreview.net/ forum?id=RDbuSCWhad
2025
-
[65]
Zico Kolter, Sanjiv Kumar, and Srinadh Bhojanapalli
Asher Trockman, Hrayr Harutyunyan, J. Zico Kolter, Sanjiv Kumar, and Srinadh Bhojanapalli. Mimetic initialization helps state space models learn to recall, 2024. URLhttps://arxiv. org/abs/2410.11135. Presented at the ICLR 2025 Workshop on Weight Space Learning
Pith/arXiv arXiv 2024
-
[66]
On the implicit bias in deep-learning algorithms.Communications of the ACM, 66 (6):86–93, 2023
Gal Vardi. On the implicit bias in deep-learning algorithms.Communications of the ACM, 66 (6):86–93, 2023. doi: 10.1145/3571070
doi:10.1145/3571070 2023
-
[67]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Infor- mation Processing Systems, volume 30, 2017. URLhttps://papers.neurips.cc/paper/ 7181-attention-is-all-you-need
2017
-
[68]
V oelker, Ivana Kaji´c, and Chris Eliasmith
Aaron R. V oelker, Ivana Kaji´c, and Chris Eliasmith. Legendre memory units: Continuous-time representation in recurrent neural networks. InAdvances in Neural Information Processing Systems, volume 32, 2019. URLhttps://proceedings.neurips.cc/paper/2019/hash/ 952285b9b7e7a1be5aa7849f32ffff05-Abstract.html
2019
-
[69]
Saurous, Charlotte Frenkel, Razvan Pascanu, Blaise Aguera y Arcas, and Joao Sacramento
Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Max- imilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Guillaume Lajoie, Rif A. Saurous, Charlotte Frenkel, Razvan Pascanu, Blaise Aguera y Arcas, and Joao Sacramento. MesaNet: Sequence modeling by locally optimal test-time training. In...
2026
-
[70]
Laura von Rueden, Sebastian Mayer, Katharina Beckh, Bogdan Georgiev, Sven Giessel- bach, Raoul Heese, Birgit Kirsch, Julius Pfrommer, Annika Pick, Rajkumar Ramamurthy, Michał Walczak, Jochen Garcke, Christian Bauckhage, and Jannis Schuecker. Informed ma- chine learning—a taxonomy and survey of integrating prior knowledge into learning sys- tems.IEEE Trans...
-
[71]
Struc- tured linear CDEs: Maximally expressive and parallel-in-time sequence models
Benjamin Walker, Lingyi Yang, Nicola Muca Cirone, Cristopher Salvi, and Terry Lyons. Struc- tured linear CDEs: Maximally expressive and parallel-in-time sequence models. InAdvances in Neural Information Processing Systems, 2025. URLhttps://openreview.net/forum? id=HKDyRDzy1E
2025
-
[72]
StableSSM: Alleviating the curse of memory in state-space mod- els through stable reparameterization
Shida Wang and Qianxiao Li. StableSSM: Alleviating the curse of memory in state-space mod- els through stable reparameterization. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Machine Learning Research, pages 50766– 50793. PMLR, 2024. URLhttps://proceedings.mlr.press/v235/wang24ag.html. 14
2024
-
[73]
State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memory
Shida Wang and Beichen Xue. State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memory. InAdvances in Neural Information Pro- cessing Systems, 2023. URLhttps://proceedings.neurips.cc/paper_files/paper/ 2023/hash/ea8608c6258450e75b3443ec8022fb2e-Abstract-Conference.html
2023
-
[74]
Gated linear attention transformers with hardware-efficient training
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. Gated linear attention transformers with hardware-efficient training. InProceedings of the 41st Interna- tional Conference on Machine Learning, 2024. URLhttps://openreview.net/forum? id=ia5XvxFUJT
2024
-
[75]
Parallelizing linear transformers with the delta rule over sequence length
Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. Parallelizing linear transformers with the delta rule over sequence length. InAdvances in Neural Information Processing Systems, 2024. URLhttps://openreview.net/forum?id=y8Rm4VNRPH
2024
-
[76]
Gated delta networks: Improving mamba2 with delta rule
Songlin Yang, Jan Kautz, and Ali Hatamizadeh. Gated delta networks: Improving mamba2 with delta rule. InInternational Conference on Learning Representations, 2025. URLhttps: //openreview.net/forum?id=r8H7xhYPwz
2025
-
[77]
Mahoney, and N
Annan Yu, Arnur Nigmetov, Dmitriy Morozov, Michael W. Mahoney, and N. Benjamin Erichson. Robustifying state-space models for long sequences via approximate diagonal- ization. InInternational Conference on Learning Representations, 2024. URLhttps: //openreview.net/forum?id=DjeQ39QoLQ
2024
-
[78]
Mahoney, and N
Annan Yu, Dongwei Lyu, Soon Hoe Lim, Michael W. Mahoney, and N. Benjamin Erichson. Tuning frequency bias of state space models. InInternational Conference on Learning Repre- sentations, 2025. URLhttps://openreview.net/forum?id=wkHcXDv7cv
2025
-
[79]
Mahoney, and N
Annan Yu, Michael W. Mahoney, and N. Benjamin Erichson. HOPE for a robust parameter- ization of long-memory state space models. InInternational Conference on Learning Repre- sentations, 2025. URLhttps://openreview.net/forum?id=RZwtbg3qYD
2025
-
[80]
Explaining modern gated-linear RNNs via a unified implicit attention formulation, 2024
Itamar Zimerman, Ameen Ali, and Lior Wolf. Explaining modern gated-linear RNNs via a unified implicit attention formulation, 2024. URLhttps://arxiv.org/abs/2405.16504. 15 A Derivations for Section 2 This appendix supports the operator-level claims of Section 2. We first recall the ordinary vector- valued HiPPO variational equations, then derive the SHiPPO...
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.