REVIEW 3 major objections 5 minor 57 references
Modular TTT: Rethinking Test-Time Training as Composable Modules
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Modular graphs turn test-time training into a composable design space
desk verdict Useful modular framework and a mostly clean empirical recipe, but the headline 'deeper hurts' result is confounded by an untested init-scale variable that the authors' own math identifies. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the learner DAG: a directed acyclic graph whose nodes are registered primitives (Linear, Gate, Norm, Act, Add, Mul) and whose edges are tensor dependencies, with a loss function attached at the output node. Each primitive carries three local rules, the train-view forward, the train-view backward, and the causal query-view forward; executing them in topological order, then reverse topological order, then topological order again produces the full TTT computation including the fast-weight state transition. The load-bearing identity is the chunkwise dual-form readout for a linear primitive, $O = QW_{\text{start}} - \operatorname{Tril}(Q\hat{K}^\top)d\hat{V}$, where the primitive rules supply the learning-rate-scaled key $\hat{K}$ and the decay matrix. Automatic differentiation supplies the local backward signals, while the query-view rules and fast-weight transition come from the registered primitives, removing the need to hand-derive a global update rule for each topology.
What would settle it
A concrete test: extend the modular primitive library with a two-step inner-update primitive and train the same 160M-scale models on 10B tokens under the same protocol; if the two-step variant lowers validation loss below Linear-SiLU, the paper's conclusion that one-step shallow learners are the strongest configuration fails. A second test would compare a momentum-augmented inner update against the same shallow frontier.
Extended reading notes
Core claim
The central claim is that every TTT variant can be expressed as a DAG of registered primitives, and that the full graph-level TTT computation composes automatically from each primitive's train-view forward, train-view backward, and causal query-view rules. The paper uses this to conduct controlled ablations at 160M and 410M scale on 10B tokens, finding that MSE and inner-product losses are equally competitive, that small learning-rate initialization stabilizes the update by keeping the spectrum of $I - K^\top \operatorname{diag}(\eta) K$ controlled, that scalar decay recovers most of the gain of vector decay, and that a single SiLU layer gives a consistent improvement. Deeper product-form learners fail to beat the shallow frontier, and the paper offers a structural explanation: rescaling the factors of a product $W^{(1)}W^{(2)}$ leaves the represented memory function unchanged but changes the induced one-step update direction, so depth adds an extra optimization burden. Guided by these findings, the selected shallow variant trained on 100B tokens reaches training loss and benchmark performance comparable to Gated DeltaNet at 410M and 1.45B scale.
Load-bearing premise
The framework's promise as a general design space rests on the assumption that every TTT variant that matters can be expressed as a directed graph of its one-step primitives; variants needing multi-step inner updates, momentum, or specialized update normalization lie outside it, and the ablations may not transfer to them.
Editorial extensions
If this is right
- New TTT variants can be assembled by rearranging or replacing registered primitives, so the design space can be explored without hand-deriving a global update for each topology.
- The ablation yields a concrete default recipe: a single linear fast weight, small learning-rate initialization, scalar decay, and a SiLU activation, with deeper, normalized, residual, or gated learners offering no measured benefit under one-step updates.
- At 410M and 1.45B parameters trained on 100B tokens, the selected variant matches Gated DeltaNet on training loss and multiple-choice accuracy, while containment-style tasks and long-context retrieval remain weaker.
- The analytic primitive back-ends give roughly 2-3x higher training throughput than the official hard-coded TTT implementation for the same topologies, and about 1.65-2.62x faster primitive-level backward computation than automatic-differentiation references.
Reading between the lines
- A consequence the paper leaves implicit is that the shallow-frontier result nudges TTT toward linear-attention-like memories with a bounded nonlinear write; if this holds, part of the reported gains in newer TTT variants may come from learning-rate and decay tuning rather than from a richer inner learner.
- Because the framework's primitives assume a single gradient step, a natural testable extension is to register a two-step or momentum primitive; if multi-step inner updates beat Linear-SiLU, the depth penalty would be a limitation of the one-step family, not of TTT in general.
- The finding that vector decay achieves the lowest loss but scalar decay captures most of the gain suggests an untested prediction: feature-selective decay should matter most on long heterogeneous sequences, where selective forgetting is more valuable.
- The speedup of analytic backward operators over automatic differentiation implies that expanding the primitive library with analytic rules for gated and normalized paths could make deeper learners inexpensive enough to revisit the depth conclusion.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Modular TTT, a framework that represents test-time-training inner learners as DAGs of primitive operators (Linear, Gate, Norm, Act, Add, Mul) and automatically composes train-view forward, train-view backward, and causal query-view rules into a chunkwise TTT computation with fast-weight state transitions. The authors report controlled ablations at 160M and 410M scale over loss function, learning-rate initialization, decay, nonlinearity, depth, and normalization, finding that small-lr initialization, scalar decay, and a single SiLU activation are beneficial, MSE and inner-product losses are comparable, deeper fast-weight networks and normalization are fragile, and residual/gated structure gives little benefit. A selected shallow variant is scaled to 410M and 1.45B parameters for 100B tokens and is reported as comparable to Gated DeltaNet on training loss and multiple-choice benchmarks, while weaker on containment and long-context retrieval. The paper also reports 2.2-3.3x training-throughput gains over the official TTT implementation.
Significance. If the results hold, the primitive grammar is a genuinely useful engineering contribution: it removes the need for per-variant derivation of global TTT updates within a broad one-step dual-form family, and the ablations provide one of the most systematic component-level pictures of TTT to date. The released code, the five-seed paired runs for the key learning-rate/decay/activation comparisons (Appendix Table 17), and the explicit stability propositions are strengths. The scale-up comparison with GDN supports the practical relevance of the shallow recommendation. However, the central negative result on deep fast-weight memories currently rests on experiments that do not control the factor-scale variable identified by the paper's own theory, so the claimed structural conclusion is not yet established.
major comments (3)
- [§4.1, Table 6, Eqs. (13)-(17), Appendix D.4] The conclusion that deeper fast-weight memories are structurally harder is underdetermined by the experiments. All deep variants initialize both fast factors with the official Gaussian std 0.02 (Appendix B.1), while the paper's own Eq. (15) shows that the induced update depends on the factor scale c. With d_head = 128, the effective product W^(1)W^(2) has entry std ≈ sqrt(128)·0.02^2 ≈ 4.5e-3, roughly 4-5x smaller than the shallow 0.02, and the terms W2^T W2 and W1 W1^T in the induced update scale as d_head·σ^2 ≈ 0.05, so the deep net starts with a damped effective update. The stabilization sweeps in Tables 22-24 vary mean scaling, epsilon, activation placement, and zero-vs-Gaussian initialization, but never the factor scale or balance that Eqs. (16)-(17) identify as controlling. Without a sweep that matches the product scale of the shallow learner, or that varies c directly, the observed +0.09 to +0.18 validation-loss gaps in Table 21 may be an artifact of the default initialization rather than evidence for the paper's factor-coupling explanation.
- [Abstract, §4.1, Table 6] The abstract's causal explanation ('Deeper fast-weight networks and normalization tend to hurt performance because they induce excessively large activations') is inconsistent with the main-text analysis for the deep linear rows. The Linear-Linear row in Table 6 contains no nonlinearity, so an activation-magnitude mechanism cannot explain its 3.1265 loss; the paper's own §4.1 explanation is a factor-coupling/update-geometry argument, while for Norm the proposed mechanism is the 1/σ gradient amplification in Eq. (10). These are different mechanisms and should not be conflated. Please revise the abstract to state the factor-coupling explanation for depth and the gradient-amplification explanation for normalization separately, or provide activation-norm diagnostics for the deep variants.
- [§1, §3.2, Appendix E] The paper repeatedly calls Modular TTT a 'unified framework' for the TTT design space, but Appendix E concedes that multi-step inner updates, momentum, and Muon-style update normalization (as used in LaCT) do not fit the current fused graph implementation without specialized update paths. Since the ablations and the deep-memory negative results are derived only under the one-step, dual-form update family, the framework is a subfamily of TTT rather than the full design space. The main text should carry the Appendix E scope statement forward, so that the 'unified framework' claim and the transferability of the ablations are not overstated.
minor comments (5)
- [Appendix B.2] The validation set used for all final-loss tables is not described; please state its size, how it is split from the pretraining corpus, and how validation loss is computed.
- [Table 9, Appendix D.5.3] Please clarify whether all external baselines in Table 9 (LLaMA, GDN, LaCT) were trained with the same tokenizer and corpus as the Modular TTT models; the LaCT check in Table 30 explicitly notes tokenizer-matched groups, and the main table should be equally explicit.
- [Table 6, Appendix D.4] Report seed variance for the deep-memory rows; the five-seed protocol is applied only to the shallow comparisons in Table 17, so the deep-memory gaps currently have no reported uncertainty.
- [Figure 1] The 'validation loss' axis in Figure 1(c) is unlabeled in the caption; please add the metric name and note the configuration used.
- [Table 1, Appendix A.5] The decay matrix notation M in Table 1 is defined only through the appendix's vector-decay formulas; a one-line pointer in Table 1 would help readers connect the primitive rules to the token-wise recurrence.
Circularity Check
No significant circularity: the Modular TTT framework, its ablations, and its stability bounds are derived from first-principles equations and external measurements rather than from their own conclusions.
full rationale
Walked the paper's derivation chain and found no load-bearing circular step. Section 3 defines the train-view forward, train-view backward, and query-view forward rules for each primitive and then composes them over a DAG; this composition is explicit algebraic bookkeeping (Eqs. 4-6, Algorithm 1) and does not assume any of the paper's empirical findings. The claims that small learning-rate initialization, scalar decay, and SiLU improve performance are empirical ablations (Tables 2-5) measured against validation loss and external baselines, not quantities fitted to the outcome being predicted. The stability analysis in Appendix C.2 proves that a small initial learning rate bounds the spectral radius of the homogeneous MSE update; this is a conditional mathematical statement that is then tested by experiments, not a prediction manufactured from the experiments themselves. The deep-memory analysis in Section 4.1 and Appendix C.4 derives factor coupling identities (Eqs. 13-17) and uses them as a mechanistic explanation, while the deep-memory results themselves come from controlled sweeps (Tables 6, 21-27); the identities do not assume the empirical outcome. The paper cites prior TTT implementations and baselines, but these citations are external or provide context, not self-citations that carry the central argument. The skeptic's concern that deep runs use a fixed Gaussian init scale is a potential experimental confound or correctness risk, not circularity, because the factor scale was not fitted to the deep-memory losses and the claim is not defined in terms of the conclusion. Overall, the paper's framework and empirical conclusions stand independently of each other, supported by reproducible equations and external comparisons.
Assumptions & free parameters
free parameters (5)
- Inner learning-rate initialization (eta_0 = 1e-3, small-lr init offset b ~ -7.60) =
1e-3
- Scalar decay initial value =
~0.99
- Chunk size =
256
- RMSNorm epsilon =
1e-5 default, swept up to 1.0
- Loss normalization constant s (mean scaling) =
s=1 (mean scaling disabled)
assumptions (4)
- standard math The chunkwise dual form (Eq. A.14) is algebraically equivalent to the token-wise recurrent update (Eqs. A.11-A.13).
- domain assumption The inner learning rate is predicted as eta_t = 2 * Sigmoid(beta_t + b) following prior work (Grazzi et al., 2024).
- domain assumption The inner learner performs exactly one gradient step per chunk (one-step TTT).
- domain assumption The set of primitives (Linear, Gate, Norm, Act, Add, Mul) and losses (inner product, MSE, L1, RMSE) is representative of the TTT design space.
Cite this review
Pith. "Pith review of Modular TTT: Rethinking Test-Time Training as Composable Modules." pith.science (2026). https://pith.science/paper/WU2UENVM
@misc{pith2026260807110,
author = {Pith},
title = {Pith review of: Modular TTT: Rethinking Test-Time Training as Composable Modules},
year = {2026},
howpublished = {\url{https://pith.science/paper/WU2UENVM}},
note = {Machine review of arXiv:2608.07110}
}
read the original abstract
Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard-code each variant separately, which makes it difficult to design new TTT methods and to isolate the role of each component. To address this, we propose Modular TTT, a framework that represents the inner learner as a directed acyclic graph and exposes the fast-weight network, loss function, learning rate, weight decay, and normalization as explicit design dimensions. Modular TTT automatically composes primitive-level train-view forward, train-view backward, and causal query-view rules into the full graph-level TTT computation, including the fast-weight state transition. Using Modular TTT, we systematically ablate the components of TTT and find that small learning-rate initialization, weight decay, and a single-layer nonlinearity improve performance, while MSE and inner-product losses perform similarly. Deeper fast-weight networks and normalization tend to hurt performance because they induce excessively large activations, while residual connections and gating provide little measurable benefit. Guided by these findings, we train the best resulting variant as 410M- and 1.45B-parameter models on 100B tokens, and observe training loss and benchmark performance comparable to Gated DeltaNet.
Reference graph
Works this paper leans on
-
[1]
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan. On the optimization of deep networks: Implicit acceleration by overparameterization. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 244–253, 2018
work page 2018
-
[2]
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo. Implicit regularization in deep matrix factorization. In Advances in Neural Information Processing Systems, volume 32, pages 7411–7422, 2019
work page 2019
-
[3]
Hinton, Volodymyr Mnih, Joel Z
Jimmy Ba, Geoffrey E. Hinton, Volodymyr Mnih, Joel Z. Leibo, and Catalin Ionescu. Using fast weights to attend to the recent past. InAdvances in Neural Information Processing Systems, volume 29, 2016
work page 2016
-
[4]
Pierre Baldi and Kurt Hornik. Neural networks and principal component analysis: Learning from examples without local minima.Neural Networks, 2(1):53–58, 1989. doi: 10.1016/0893-6080(89)90014-2
-
[5]
ATLAS: Learning to optimally memorize the context at test time.arXiv preprint arXiv:2505.23735, 2025
Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, and Vahab Mirrokni. ATLAS: Learning to optimally memorize the context at test time.arXiv preprint arXiv:2505.23735, 2025
arXiv 2025
-
[6]
Ali Behrouz, Meisam Razaviyayn, Peilin Zhong, and Vahab Mirrokni. It’s all connected: A journey through test-time memorization, attentional bias, retention, and online optimization.arXiv preprint arXiv:2504.13173, 2025
arXiv 2025
-
[7]
Nested learning: The illusion of deep learning architectures
Ali Behrouz, Meisam Razaviyayn, Peilin Zhong, and Vahab Mirrokni. Nested learning: The illusion of deep learning architectures. arXiv preprint arXiv:2512.24695, 2025
arXiv 2025
-
[8]
Titans: Learning to memorize at test time.arXiv preprint arXiv:2501.00663, 2025
Ali Behrouz, Peilin Zhong, and Vahab Mirrokni. Titans: Learning to memorize at test time.arXiv preprint arXiv:2501.00663, 2025
arXiv 2025
Show all 57 references
-
[9]
PIQA: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. PIQA: Reasoning about physical commonsense in natural language. InProceedings of the AAAI Conference on Artificial Intelligence, 2020
2020
-
[10]
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. Rethinking attention with performers. InInternational Conference on...
2021
-
[11]
Empirical evaluation of gated recurrent neural networks on sequence modeling.arXiv preprint arXiv:1412.3555, 2014
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. Empirical evaluation of gated recurrent neural networks on sequence modeling.arXiv preprint arXiv:1412.3555, 2014
2014 arXiv
-
[12]
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computatio...
2019
-
[13]
Think you have solved question answering? try arc, the ai2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018
2018 arXiv
-
[14]
Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality
Tri Dao and Albert Gu. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. InProceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 10041–10071. ...
2024
-
[15]
Franke, Arber Zela, Frank Hutter, and Massimiliano Pontil
Riccardo Grazzi, Julien Siems, Jörg K.H. Franke, Arber Zela, Frank Hutter, and Massimiliano Pontil. Unlocking state-tracking in linear RNNs through negative eigenvalues. InNeurIPS 2024 Workshop on Mathematics of Modern Machine Learning, 2024. URLhttps://openreview.net/forum?id...
2024
-
[16]
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. InFirst Conference on Language Modeling, 2024
2024
-
[17]
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré. Efficiently modeling long sequences with structured state spaces. In International Conference on Learning Representations, 2022
2022
-
[18]
Vit3: Unlocking test-time training in vision.arXiv preprint arXiv:2512.01643, 2025
Dongchen Han, Yining Li, Tianyu Li, Zixuan Cao, Ziming Wang, Jun Song, Yu Cheng, Bo Zheng, and Gao Huang. Vit3: Unlocking test-time training in vision.arXiv preprint arXiv:2512.01643, 2025. 13
2025 arXiv
-
[19]
From one tree to a forest: A unified solution for structured web data extraction
Qiang Hao, Rui Cai, Yanwei Pang, and Lei Zhang. From one tree to a forest: A unified solution for structured web data extraction. InProceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 775–784, 2011
2011
-
[20]
Long short-term memory.Neural Computation, 9(8):1735–1780, 1997
Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory.Neural Computation, 9(8):1735–1780, 1997
1997
-
[21]
RULER: What’s the real context size of your long-context language models? InFirst Conference on Language Modeling, 2024
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg. RULER: What’s the real context size of your long-context language models? InFirst Conference on Language Modeling, 2024. URLhttps://openreview.net/forum?id=kIoBbc76Sy
2024
-
[22]
Going beyond linear transformers with recurrent fast weight programmers
Kazuki Irie, Imanol Schlag, Róbert Csordás, and Jürgen Schmidhuber. Going beyond linear transformers with recurrent fast weight programmers. InAdvances in Neural Information Processing Systems, volume 34, 2021
2021
-
[23]
The dual form of neural networks revisited: Connecting test time predictions to training patterns via spotlights of attention
Kazuki Irie, Róbert Csordás, and Jürgen Schmidhuber. The dual form of neural networks revisited: Connecting test time predictions to training patterns via spotlights of attention. InProceedings of the 39th International Conference on Machine Learning, 2022
2022
-
[24]
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. InProceedings of the 37th International Conference on Machine Learning, 2020
2020
-
[25]
Deep learning without poor local minima
Kenji Kawaguchi. Deep learning without poor local minima. InAdvances in Neural Information Processing Systems, volume 29, pages 586–594, 2016
2016
-
[26]
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals. Dynamic evaluation of neural sequence models. In Proceedings of the 35th International Conference on Machine Learning, 2018
2018
-
[27]
TNT: Improving chunkwise training for test-time memorization.arXiv preprintarXiv:2511.07343, 2025
Zeman Li, Ali Behrouz, Yuan Deng, Peilin Zhong, Praneeth Kacham, Mahdi Karami, Meisam Razaviyayn, and Vahab Mirrokni. TNT: Improving chunkwise training for test-time memorization.arXiv preprintarXiv:2511.07343, 2025
2025
-
[28]
Parallelizing linear recurrent neural nets over sequence length
Eric Martin and Chris Cundy. Parallelizing linear recurrent neural nets over sequence length. InInternational Conference on Learning Representations, 2018
2018
-
[29]
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In International Conference on Learning Representations, 2017
2017
-
[30]
Can a suit of armor conduct electricity? A new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? A new dataset for open book question answering. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2618–2628, 2018
2018
-
[31]
Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De
Antonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De. Resurrecting recurrent neural networks for long sequences. InProceedings of the 40th International Conference on Machine Learning, 2023
2023
-
[32]
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. The LAMBADA dataset: Word prediction requiring a broad discourse context. InProceedings of the 54th Annual Meeting of th...
2016
-
[33]
RWKV: Reinventing RNNs for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Leon Derczynski, Xingjian Du, Matteo Grella, Kranthi Gv, Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bartłomiej Koptyra, Hay...
2023
-
[34]
Wind, Tianyi Wu, Daniel Wuttke, and Christian Zhou-Zheng
Bo Peng, Ruichong Zhang, Daniel Goldstein, Eric Alcaide, Xingjian Du, Haowen Hou, Jiaju Lin, Jiaxing Liu, Janna Lu, William Merrill, Guangyu Song, Kaifeng Tan, Saiteja Utpala, Nathan Wilce, Johan S. Wind, Tianyi Wu, Daniel Wuttke, and Christian Zhou-Zheng. RWKV-7 “goose” with ...
-
[35]
Hierarchically gated recurrent neural network for sequence modeling
Zhen Qin, Songlin Yang, and Yiran Zhong. Hierarchically gated recurrent neural network for sequence modeling. In Advances in Neural Information Processing Systems, 2023. 14
2023
-
[36]
Transnormerllm: A faster and better large language model with improved transnormer, 2024
Zhen Qin, Dong Li, Weigao Sun, Weixuan Sun, Xuyang Shen, Xiaodong Han, Yunshen Wei, Baohong Lv, Xiao Luo, Yu Qiao, and Yiran Zhong. Transnormerllm: A faster and better large language model with improved transnormer, 2024
2024
-
[37]
HGRN2: Gated linear RNNs with state expansion
Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong. HGRN2: Gated linear RNNs with state expansion. InFirst Conference on Language Modeling, 2024
2024
-
[38]
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, 2016
2016
-
[39]
WinoGrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. WinoGrande: An adversarial winograd schema challenge at scale. InProceedings of the AAAI Conference on Artificial Intelligence, 2020
2020
-
[40]
Social IQa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. Social IQa: Commonsense reasoning about social interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural ...
2019 doi
-
[41]
Linear transformers are secretly fast weight programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. Linear transformers are secretly fast weight programmers. In Proceedings of the 38th International Conference on Machine Learning, 2021
2021
-
[42]
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber. Learning to control fast-weight memories: An alternative to dynamic recurrent networks. Neural Computation, 4(1):131–139, 1992
1992
-
[43]
GLU variants improve transformer.arXiv preprint arXiv:2002.05202, 2020
Noam Shazeer. GLU variants improve transformer.arXiv preprint arXiv:2002.05202, 2020
2002 arXiv
-
[44]
Learning to (learn at test time): RNNs with expressive hidden states.arXiv preprint arXiv:2407.04620, 2024
Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, and Carlos Guestrin. Learning to (learn at test time): RNNs with expressive hidden states.arXiv preprint arXiv:2407.04620, 2024
2024 arXiv
-
[45]
Retentive network: A successor to transformer for large language models.arXiv preprint arXiv:2307.08621, 2023
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. Retentive network: A successor to transformer for large language models.arXiv preprint arXiv:2307.08621, 2023
2023 arXiv
-
[46]
EleutherAI/lm-evaluation-harness: v0.4.9.1, August 2025
Lintang Sutawika, Hailey Schoelkopf, Leo Gao, Baber Abbasi, Stella Biderman, Jonathan Tow, ben fattori, Charles Lovering, farzanehnakhaee70, Jason Phang, Anish Thite, Fazz, Aflah, Niklas, Thomas Wang, sdtblck, nopperl, gakada, tttyuntian, researcher2, Julen Etxaniz, Chris, Han...
2025
-
[47]
End-to-end test-time training for long context.arXiv preprint arXiv:2512.23675, 2025
Arnuv Tandon, Karan Dalal, Xinhao Li, Daniel Koceja, Marcel Rød, Sam Buchanan, Xiaolong Wang, Jure Leskovec, Sanmi Koyejo, Tatsunori Hashimoto, Carlos Guestrin, Jed McCaleb, Yejin Choi, and Yu Sun. End-to-end test-time training for long context.arXiv preprint arXiv:2512.23675, 2025
2025
-
[48]
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[49]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems, volume 30, 2017
2017
-
[50]
Gated linear attention transformers with hardware-efficient training
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. Gated linear attention transformers with hardware-efficient training. InInternational Conference on Machine Learning, 2024
2024
-
[51]
Parallelizing linear transformers with the delta rule over sequence length
Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. Parallelizing linear transformers with the delta rule over sequence length. In Advances in Neural Information Processing Systems, volume 37,
-
[52]
Gated delta networks: Improving Mamba2 with delta rule
Songlin Yang, Jan Kautz, and Ali Hatamizadeh. Gated delta networks: Improving Mamba2 with delta rule. In International Conference on Learning Representations, 2025
2025
-
[53]
Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019. 15
2019
-
[54]
Freeman, and Hao Tan
Tianyuan Zhang, Sai Bi, Yicong Hong, Kai Zhang, Fujun Luan, Songlin Yang, Kalyan Sunkavalli, William T. Freeman, and Hao Tan. Test-time training done right.arXiv preprint arXiv:2505.23884, 2025
2025 arXiv
-
[55]
Flame: Flash language modeling made easy, January 2025
Yu Zhang and Songlin Yang. Flame: Flash language modeling made easy, January 2025. URLhttps://github. com/fla-org/flame. 16 Appendix Appendix Contents Appendix AMethod details ...........................................................................17 Appendix BExperimental ...
2025
-
[2023]
doi: 10.18653/v1/2023.findings-emnlp.936
Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-emnlp.936. URL https: //aclanthology.org/2023.findings-emnlp.936/
2023 doi
-
[2024]
URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ d13a3eae72366e61dfdc7eea82eeb685-Abstract-Conference.html
doi: 10.52202/079017-3668. URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ d13a3eae72366e61dfdc7eea82eeb685-Abstract-Conference.html
2024 doi
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.