REVIEW 4 major objections 1 minor 2 cited by
A three-axis taxonomy organizes Memory-Augmented Transformers and shows a shift from static caches to adaptive test-time learning.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Memory-augmented Transformer research is organized into a three-axis taxonomy bridging neuroscience memory concepts to network designs, but no new result is produced.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection A coherent survey abstract atop an unreadable body: worth sending back for a clean copy and a methods section, not yet citable. the 4 major comments →
Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that the diverse memory mechanisms bolted onto Transformers form a coherent design space, not a collection of unrelated patches. It proposes a three-axis taxonomy: functional objectives (context extension, reasoning, knowledge integration, adaptation), memory representations (parameter-encoded, state-based, explicit, hybrid), and integration mechanisms (attention fusion, gated control, associative retrieval). Using reading, writing, forgetting, and capacity management as the core memory operations, it argues that the field is shifting from fixed external caches toward adaptive systems that update their own memory at test time. It further claims that neuroscience
What carries the argument
The organizing device is the taxonomy itself: each memory-augmented Transformer is located by a triple—objective, representation, integration. The four memory operations (reading, writing, forgetting, capacity management) are the analytic lens that turns the taxonomy from a static list into a diagnosis of what each architecture can and cannot do. The neuroscience mapping supplies the design rationale: multi-timescale memory justifies separate memory stores, selective attention justifies content-based retrieval, and consolidation justifies writing and forgetting rules—so the taxonomy is meant to be explanatory, not merely descriptive.
Load-bearing premise
The framework stands or falls on whether the neuroscience constructs genuinely correspond to the engineering categories and whether the surveyed papers were selected systematically rather than because they fit the story.
What would settle it
Take a defined literature window, for example memory-augmented Transformer papers published 2023–2025, collect the full set with a fixed search query, and try to assign each paper to exactly one cell of the three-axis taxonomy. If more than a small fraction resist placement, or the full set shows no chronological trend toward test-time learning, the central claims fail.
If this is right
- Architectures that currently look incomparable—a long-context cache, a gated state model, an external associative memory—become comparable through their position on the three axes.
- The claimed shift to adaptive, test-time learning sets a concrete design target: future models should write and update memory during inference, not only at training time.
- Scalability and interference are named as the two bottlenecks any new memory design must address, so progress on them can be measured directly.
- Hierarchical buffering and surprise-gated updates are named as emerging mechanisms, giving practitioners concrete starting points for new architectures.
- The neuroscience link provides a vocabulary for generating new designs, such as treating a Transformer's memory as a multi-timescale system with distinct stores.
Where Pith is reading between the lines
- Editorial inference: the taxonomy's usefulness depends on whether it can classify papers independently; a natural test is to take a held-out set of recent memory-augmented Transformer papers and see whether two annotators place them in the same taxonomy cells.
- Editorial inference: if the shift-to-test-time-learning claim is right, evaluation practice should change—benchmarks that only test static retrieval will miss the capabilities that differentiate the newer architectures.
- Editorial inference: the paper asserts rather than demonstrates its 'systematic' coverage; a reader should look for a reproducible search protocol, because without one the trend claim could reflect selection bias.
- Editorial inference: the memory-operations framework likely generalizes beyond Transformers to any architecture with an explicit memory component, so the roadmap could also apply to state-space models or recurrent networks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript (arXiv:2508.10824) claims to provide a unified framework bridging neuroscience principles with engineering advances in Memory-Augmented Transformers, organized along three taxonomic dimensions (functional objectives, memory representations, integration mechanisms), and to identify a shift from static caches toward adaptive, test-time learning systems. The abstract is self-consistent and the claims are appropriately modest for a review. However, the supplied full text is heavily corrupted (mojibake) and unreadable; it also contains an inserted fragment from a different paper (arXiv:2508.10821v3 [q-bio.QM]). Consequently, none of the substantive content—sections, tables, equations, references, or the survey corpus—can be verified. The 'systematic' nature of the review and the empirical trend claim cannot be audited from the provided text.
Significance. If the claims hold, the three-axis taxonomy could provide a useful common vocabulary for memory-augmented transformers and a roadmap toward continual-learning architectures. The synthesis of neuroscience constructs and engineering designs could be valuable if the mapping is substantive. However, because the body is unreadable, the significance is conditional. The paper does not derive new mathematical results or ship code; its contribution, if any, is organizational and interpretive. Credit is due for a self-consistent abstract and clearly stated dimensions, but no machine-checkable evidence is available.
major comments (4)
- [Full text (after abstract)] The entire body of the manuscript is encoded mojibake; no section, table, equation, or reference can be read or verified. For a systematic review, the central claims—the taxonomy and the trend toward adaptive test-time learning—rest on the ability to check which papers were surveyed and how they were mapped. This is load-bearing and prevents any audit. I cannot determine whether the summaries are accurate or whether the conclusions follow from the cited literature.
- [Visible fragment in full text] The text includes the line 'arXiv:2508.10821v3 [q-bio.QM] 20 May 2026', which is an identifier for a different paper in quantitative biology. This indicates that the PDF/source has been corrupted or misassembled. As supplied, the document is not self-consistent; it contains content that does not belong to the claimed review. This is a material issue because it makes it impossible to attribute any of the surrounding text to the authors' intended manuscript.
- [Search protocol/inclusion criteria] The abstract calls the review 'systematic,' but no search strategy, inclusion/exclusion criteria, database sources, or PRISMA-style flow is visible in the supplied text. The trend claim ('a shift from static caches toward adaptive, test-time learning systems') is an empirical statement about the surveyed literature; without a reproducible corpus and coding protocol, it may be a cherry-picked narrative. This is an audibility problem, not necessarily an error, but it must be fixed for the review to be evaluable.
- [Neuroscience-to-engineering mapping] The 'unified framework' is the paper's main contribution, but the body that would justify the correspondence between biological constructs (multi-timescale memory, selective attention, consolidation) and architectural families (parameter-encoded, state-based, explicit, hybrid) is unreadable. This raises a correctness risk: the mapping could be substantive constraints or post-hoc labels. I cannot adjudicate without a legible exposition of how each architecture family instantiates each biological principle, including any criteria for correspondence.
minor comments (1)
- [Abstract] The abstract is readable and clear; the title matches the abstract. No further presentation issues can be assessed because the remainder is illegible. If the corruption is a rendering artifact, please resubmit a clean version.
Circularity Check
No significant circularity: the paper is a survey/taxonomy with no fitted parameters, derived equations, or load-bearing self-citations, so there is no derivation chain that reduces to its inputs.
full rationale
The manuscript is a systematic review and taxonomy paper. Its central claims are the proposed three-axis taxonomy (functional objectives, memory representations, integration mechanisms) and an interpretive trend statement that 'analysis ... reveals a shift from static caches toward adaptive, test-time learning systems.' Neither claim is derived from equations or fitted constants; the taxonomy is a classification scheme, and the trend is an inductive reading of the surveyed literature. There is no self-definitional step, no fitted input renamed as prediction, and no self-citation chain invoked to force a conclusion. The abstract and the legible fragments contain no derivation machinery for the trend claim, and the full text is largely mojibake, which makes the 'systematic' methodology unauditable; however, an inability to verify the survey corpus is an evidence/auditability concern, not a circularity of the kind defined here. The visible arXiv identifier (2508.10821v3 [q-bio.QM]) embedded in the text further indicates document corruption, but this does not constitute a logical loop. Accordingly, no specific circular step can be quoted with the required reduction, and the appropriate finding is no significant circularity.
Axiom & Free-Parameter Ledger
axioms (2)
- domain assumption Cited papers report their results accurately
- domain assumption The neuroscience-to-engineering correspondence is substantive
Cite this review
Pith. "Pith review of Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures." pith.science (2026). https://pith.science/paper/VDNDJEAA
@misc{pith2026250810824,
author = {Pith},
title = {Pith review of: Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures},
year = {2026},
howpublished = {\url{https://pith.science/paper/VDNDJEAA}},
note = {Machine review of arXiv:2508.10824}
}
read the original abstract
Memory is fundamental to intelligence, enabling learning, reasoning, and adaptability across biological and artificial systems. While Transformer architectures excel at sequence modeling, they face critical limitations in long-range context retention, continual learning, and knowledge integration. This review presents a unified framework bridging neuroscience principles, including dynamic multi-timescale memory, selective attention, and consolidation, with engineering advances in Memory-Augmented Transformers. We organize recent progress through three taxonomic dimensions: functional objectives (context extension, reasoning, knowledge integration, adaptation), memory representations (parameter-encoded, state-based, explicit, hybrid), and integration mechanisms (attention fusion, gated control, associative retrieval). Our analysis of core memory operations (reading, writing, forgetting, and capacity management) reveals a shift from static caches toward adaptive, test-time learning systems. We identify persistent challenges in scalability and interference, alongside emerging solutions including hierarchical buffering and surprise-gated updates. This synthesis provides a roadmap toward cognitively-inspired, lifelong-learning Transformer architectures.
Forward citations
Cited by 2 Pith papers
-
Neural Procedural Memory: Empowering LLM Agents with Implicit Activation Steering
NPM distills contrastive experiences into implicit activation steering vectors that guide LLM agent execution comparably to explicit RAG instructions, with complementary gains when combined.
-
Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance
The proposed CAMVR framework is not supported by verifiable evidence, and the manuscript itself labels its experimental results as fabricated.
Reference graph
Works this paper leans on
-
[1]
Adult hippocampal neurogenesis and cognitive flexibility—linking memory and mood
Christoph Anacker and Ren \'e Hen. Adult hippocampal neurogenesis and cognitive flexibility—linking memory and mood. Nature Reviews Neuroscience, 18 0 (6): 0 335--346, 2017
2017
-
[2]
Global workspace theory (gwt) and prefrontal cortex: Recent developments
Bernard J Baars, Natalie Geld, and Robert Kozma. Global workspace theory (gwt) and prefrontal cortex: Recent developments. Frontiers in psychology, 12: 0 749868, 2021
2021
-
[3]
Working memory: looking back and looking forward
Alan Baddeley. Working memory: looking back and looking forward. Nature reviews neuroscience, 4 0 (10): 0 829--839, 2003
2003
-
[4]
Fast adaptation to rule switching using neuronal surprise
Martin LLR Barry and Wulfram Gerstner. Fast adaptation to rule switching using neuronal surprise. PLoS computational biology, 20 0 (2): 0 e1011839, 2024
2024
-
[5]
Neuromodulators and long-term synaptic plasticity in learning and memory: A steered-glutamatergic perspective
Amjad H Bazzari and H Rheinallt Parri. Neuromodulators and long-term synaptic plasticity in learning and memory: A steered-glutamatergic perspective. Brain sciences, 9 0 (11): 0 300, 2019
2019
-
[6]
Titans: Learning to memorize at test time
Ali Behrouz, Peilin Zhong, and Vahab Mirrokni. Titans: Learning to memorize at test time. arXiv preprint arXiv:2501.00663, 2024
Pith/arXiv arXiv 2024
-
[7]
Atlas: Learning to optimally memorize the context at test time
Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, and Vahab Mirrokni. Atlas: Learning to optimally memorize the context at test time. arXiv preprint arXiv:2505.23735, 2025
Pith/arXiv arXiv 2025
-
[8]
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020
Pith/arXiv arXiv 2004
-
[9]
Vincent-Pierre Berges, Barlas O g uz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, and Gargi Ghosh. Memory layers at scale. arXiv preprint arXiv:2412.09764, 2024
Pith/arXiv arXiv 2024
-
[10]
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pp.\ 2206--2240. PMLR, 2022
2022
-
[11]
The neuroanatomical, neurophysiological and psychological basis of memory: Current models and their origins
Eduardo Camina and Francisco G \"u ell. The neuroanatomical, neurophysiological and psychological basis of memory: Current models and their origins. Frontiers in pharmacology, 8: 0 438, 2017
2017
-
[12]
An evolved universal transformer memory
Edoardo Cetin, Qi Sun, Tianyu Zhao, and Yujin Tang. An evolved universal transformer memory. In The Thirteenth International Conference on Learning Representations, 2025
2025
-
[13]
Walking down the memory maze: Beyond context limit through interactive reading
Howard Chen, Ramakanth Pasunuru, Jason Weston, and Asli Celikyilmaz. Walking down the memory maze: Beyond context limit through interactive reading. arXiv preprint arXiv:2310.05029, 2023
Pith/arXiv arXiv 2023
-
[14]
Mem0: Building production-ready ai agents with scalable long-term memory
Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready ai agents with scalable long-term memory. arXiv preprint arXiv:2504.19413, 2025
Pith/arXiv arXiv 2025
-
[15]
Interactions between attention and memory
Marvin M Chun and Nicholas B Turk-Browne. Interactions between attention and memory. Current opinion in neurobiology, 17 0 (2): 0 177--184, 2007
2007
-
[16]
Attention-dependent coupling with forebrain and brainstem neuromodulatory nuclei differs across the lifespan
Nicholas G Cicero, Elizabeth Riley, Khena M Swallow, Eve De Rosa, and Adam Anderson. Attention-dependent coupling with forebrain and brainstem neuromodulatory nuclei differs across the lifespan. GeroScience, pp.\ 1--20, 2025
2025
-
[17]
What are the differences between long-term, short-term, and working memory? Progress in brain research, 169: 0 323--338, 2008
Nelson Cowan. What are the differences between long-term, short-term, and working memory? Progress in brain research, 169: 0 323--338, 2008
2008
-
[18]
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019
Pith/arXiv arXiv 1901
-
[19]
The global neuronal workspace model of conscious access: from neuronal architectures to clinical applications
Stanislas Dehaene, Jean-Pierre Changeux, and Lionel Naccache. The global neuronal workspace model of conscious access: from neuronal architectures to clinical applications. Characterizing consciousness: From cognition to the clinic?, pp.\ 55--84, 2011
2011
-
[20]
Pronouns reactivate conceptual representations in human hippocampal neurons
Doris E Dijksterhuis, Matthew W Self, Jessy K Possel, Judith C Peters, ECW van Straaten, Sander Idema, Johannes C Baaijen, Sandra MA van der Salm, Erik J Aarnoutse, Nicole CE van Klink, et al. Pronouns reactivate conceptual representations in human hippocampal neurons. Science, 385 0 (6716): 0 1478--1484, 2024
2024
-
[21]
Interaction between the amygdala and the medial temporal lobe memory system predicts better memory for emotional events
Florin Dolcos, Kevin S LaBar, and Roberto Cabeza. Interaction between the amygdala and the medial temporal lobe memory system predicts better memory for emotional events. Neuron, 42 0 (5): 0 855--863, 2004
2004
-
[22]
Rethinking memory in ai: Taxonomy, operations, topics, and future directions
Yiming Du, Wenyu Huang, Danna Zheng, Zhaowei Wang, Sebastien Montella, Mirella Lapata, Kam-Fai Wong, and Jeff Z Pan. Rethinking memory in ai: Taxonomy, operations, topics, and future directions. arXiv preprint arXiv:2505.00675, 2025
arXiv 2025
-
[23]
Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers
Sanchayan Dutta and Suvrit Sra. Memory-augmented transformers can implement linear first-order optimization methods. arXiv preprint arXiv:2410.07263, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[24]
Efficient llm inference using dynamic input pruning and cache-aware masking
Marco Federici, Davide Belli, Mart Van Baalen, Amir Jalalirad, Andrii Skliar, Bence Major, Markus Nagel, and Paul Whatmough. Efficient llm inference using dynamic input pruning and cache-aware masking. arXiv preprint arXiv:2412.01380, 2024
Pith/arXiv arXiv 2024
-
[25]
Human-like episodic memory for infinite context llms
Zafeirios Fountas, Martin A Benfeghoul, Adnan Oomerjee, Fenia Christopoulou, Gerasimos Lampouras, Haitham Bou-Ammar, and Jun Wang. Human-like episodic memory for infinite context llms. arXiv preprint arXiv:2407.09450, 2024
arXiv 2024
-
[26]
Experiencing surprise: The temporal dynamics of its impact on memory
Darya Frank, Alex Kafkas, and Daniela Montaldi. Experiencing surprise: The temporal dynamics of its impact on memory. Journal of Neuroscience, 42 0 (33): 0 6435--6444, 2022
2022
-
[27]
An efficient context-dependent memory framework for llm-centric agents
Pengyu Gao, Jinming Zhao, Xinyue Chen, and Long Yilin. An efficient context-dependent memory framework for llm-centric agents. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track), pp.\ 1055--1069, 2025
2025
-
[28]
Memotr: Long-term memory-augmented transformer for multi-object tracking
Ruopeng Gao and Limin Wang. Memotr: Long-term memory-augmented transformer for multi-object tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.\ 9901--9910, 2023
2023
-
[29]
Top-down modulation: bridging selective attention and working memory
Adam Gazzaley and Anna C Nobre. Top-down modulation: bridging selective attention and working memory. Trends in cognitive sciences, 16 0 (2): 0 129--135, 2012
2012
-
[30]
zip2zip: Inference-time adaptive vocabularies for language models via token compression
Saibo Geng, Nathan Ranchin, Maxime Peyrard, Chris Wendler, Michael Gastpar, Robert West, et al. zip2zip: Inference-time adaptive vocabularies for language models via token compression. arXiv preprint arXiv:2506.01084, 2025
-
[31]
The role of the ca3 hippocampal subregion in spatial memory: a process oriented behavioral assessment
Paul E Gilbert and Andrea M Brushfield. The role of the ca3 hippocampal subregion in spatial memory: a process oriented behavioral assessment. Progress in Neuro-Psychopharmacology and Biological Psychiatry, 33 0 (5): 0 774--781, 2009
2009
-
[32]
Alex Graves, Greg Wayne, and Ivo Danihelka. Neural turing machines. arXiv preprint arXiv:1410.5401, 2014
Pith/arXiv arXiv 2014
-
[33]
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwi \'n ska, Sergio G \'o mez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al. Hybrid computing using a neural network with dynamic external memory. Nature, 538 0 (7626): 0 471--476, 2016
2016
-
[34]
Hipporag: Neurobiologically inspired long-term memory for large language models
Bernal Jim \'e nez Guti \'e rrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. Hipporag: Neurobiologically inspired long-term memory for large language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024
2024
-
[35]
Hierarchical process memory: memory as an integral component of information processing
Uri Hasson, Janice Chen, and Christopher J Honey. Hierarchical process memory: memory as an integral component of information processing. Trends in cognitive sciences, 19 0 (6): 0 304--313, 2015
2015
-
[36]
Memory matters: The need to improve long-term memory in llm-agents
Kostas Hatalis, Despina Christou, Joshua Myers, Steven Jones, Keith Lambert, Adam Amos-Binks, Zohreh Dannenhauer, and Dustin Dannenhauer. Memory matters: The need to improve long-term memory in llm-agents. In Proceedings of the AAAI Symposium Series, volume 2, pp.\ 277--280, 2023
2023
-
[37]
Hmt: Hierarchical memory transformer for efficient long context language processing
Zifan He, Yingqi Cao, Zongyue Qin, Neha Prakriya, Yizhou Sun, and Jason Cong. Hmt: Hierarchical memory transformer for efficient long context language processing. arXiv preprint arXiv:2405.06067, 2024 a
Pith/arXiv arXiv 2024
-
[38]
Human-inspired perspectives: A survey on ai long-term memory
Zihong He, Weizhe Lin, Hao Zheng, Fan Zhang, Matt W Jones, Laurence Aitchison, Xuhai Xu, Miao Liu, Per Ola Kristensson, and Junxiao Shen. Human-inspired perspectives: A survey on ai long-term memory. arXiv preprint arXiv:2411.00489, 2024 b
Pith/arXiv arXiv 2024
-
[39]
Replay bursts in humans coincide with activation of the default mode and parietal alpha networks
Cameron Higgins, Yunzhe Liu, Diego Vidaurre, Zeb Kurth-Nelson, Ray Dolan, Timothy Behrens, and Mark Woolrich. Replay bursts in humans coincide with activation of the default mode and parietal alpha networks. Neuron, 109 0 (5): 0 882--893, 2021
2021
-
[40]
Transformerfam: Feedback attention is working memory
Dongseong Hwang, Weiran Wang, Zhuoyuan Huo, Khe Chai Sim, and Pedro Moreno Mengibar. Transformerfam: Feedback attention is working memory. arXiv preprint arXiv:2404.09173, 2024
Pith/arXiv arXiv 2024
-
[41]
Jiazheng Kang, Mingming Ji, Zhe Zhao, and Ting Bai. Memory os of ai agent. arXiv preprint arXiv:2506.06326, 2025 a
Pith/arXiv arXiv 2025
-
[42]
Jikun Kang, Wenqi Wu, Filippos Christianos, Alex J Chan, Fraser Greenlee, George Thomas, Marvin Purtorab, and Andy Toulis. Lm2: Large memory models. arXiv preprint arXiv:2502.06049, 2025 b
Pith/arXiv arXiv 2025
-
[43]
Distinguishing examples while building concepts in hippocampal and artificial networks
Louis Kang and Taro Toyoizumi. Distinguishing examples while building concepts in hippocampal and artificial networks. Nature Communications, 15 0 (1): 0 647, 2024
2024
-
[44]
Mechanisms of systems memory consolidation during sleep
Jens G Klinzing, Niels Niethard, and Jan Born. Mechanisms of systems memory consolidation during sleep. Nature neuroscience, 22 0 (10): 0 1598--1610, 2019
2019
-
[45]
Memreasoner: A memory-augmented llm architecture for multi-hop reasoning
Ching-Yun Ko, Sihui Dai, Payel Das, Georgios Kollias, Subhajit Chaudhury, and Aurelie Lozano. Memreasoner: A memory-augmented llm architecture for multi-hop reasoning. In The First Workshop on System-2 Reasoning at Scale, NeurIPS'24, 2024
2024
-
[46]
Flexible working memory through selective gating and attentional tagging
Wouter Kruijne, Sander M Bohte, Pieter R Roelfsema, and Christian NL Olivers. Flexible working memory through selective gating and attentional tagging. Neural Computation, 33 0 (1): 0 1--40, 2021
2021
-
[47]
Semantic memory: A review of methods, models, and current challenges
Abhilasha A Kumar. Semantic memory: A review of methods, models, and current challenges. Psychonomic bulletin & review, 28 0 (1): 0 40--80, 2021
2021
-
[48]
Self-attentive associative memory
Hung Le, Truyen Tran, and Svetha Venkatesh. Self-attentive associative memory. In International conference on machine learning, pp.\ 5682--5691. PMLR, 2020
2020
-
[49]
Reasoning under 1 billion: Memory-augmented reinforcement learning for large language models
Hung Le, Dai Do, Dung Nguyen, and Svetha Venkatesh. Reasoning under 1 billion: Memory-augmented reinforcement learning for large language models. arXiv preprint arXiv:2504.02273, 2025
Pith/arXiv arXiv 2025
-
[50]
MATTER: Memory-Augmented Transformer Using Heterogeneous Knowledge Sources
Dongkyu Lee, Chandana Satya Prakash, Jack FitzGerald, and Jens Lehmann. Matter: Memory-augmented transformer using heterogeneous knowledge sources. arXiv preprint arXiv:2406.04670, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[51]
An update on memory reconsolidation updating
Jonathan LC Lee, Karim Nader, and Daniela Schiller. An update on memory reconsolidation updating. Trends in cognitive sciences, 21 0 (7): 0 531--545, 2017
2017
-
[52]
u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K \"u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \"a schel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33: 0 9459--9474, 2020
2020
-
[53]
FALCON: Feedback-driven Adaptive Long/short-term memory reinforced Coding Optimization system
Zeyuan Li, Yangfan He, Lewei He, Jianhui Wang, Tianyu Shi, Bin Lei, Yuchen Li, and Qiuwu Chen. Falcon: Feedback-driven adaptive long/short-term memory reinforced coding optimization system. arXiv preprint arXiv:2410.21349, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[54]
Self-evolving agents with reflective and memory-augmented abilities
Xuechen Liang, Yangfan He, Yinghui Xia, Xinyuan Song, Jianhui Wang, Meiling Tao, Li Sun, Xinhang Yuan, Jiayi Su, Keqin Li, et al. Self-evolving agents with reflective and memory-augmented abilities. arXiv preprint arXiv:2409.00872, 2024
Pith/arXiv arXiv 2024
-
[55]
Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang, Tanjin He, Sirui Hong, Hongzhang Liu, Shaokun Zhang, Kaitao Song, Kunlun Zhu, et al. Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems. arXiv preprint arXiv:2504.01990, 2025
Pith/arXiv arXiv 2025
-
[56]
Think-in-memory: Recalling and post-thinking enable llms with long-term memory
Lei Liu, Xiaoyan Yang, Yue Shen, Binbin Hu, Zhiqiang Zhang, Jinjie Gu, and Guannan Zhang. Think-in-memory: Recalling and post-thinking enable llms with long-term memory. arXiv preprint arXiv:2311.08719, 2023
Pith/arXiv arXiv 2023
-
[57]
MemLong: Memory-Augmented Retrieval for Long Text Modeling
Weijie Liu, Zecheng Tang, Juntao Li, Kehai Chen, and Min Zhang. Memlong: Memory-augmented retrieval for long text modeling. arXiv preprint arXiv:2408.16967, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[58]
Experience replay is associated with efficient nonlocal learning
Yunzhe Liu, Marcelo G Mattar, Timothy EJ Behrens, Nathaniel D Daw, and Raymond J Dolan. Experience replay is associated with efficient nonlocal learning. Science, 372 0 (6544): 0 eabf1357, 2021
2021
-
[59]
Acquiring new memories in neocortex of hippocampal-lesioned mice
Wenhan Luo, Di Yun, Yi Hu, Miaomiao Tian, Jiajun Yang, Yifan Xu, Yong Tang, Yang Zhan, Hong Xie, and Ji-Song Guan. Acquiring new memories in neocortex of hippocampal-lesioned mice. Nature communications, 13 0 (1): 0 1601, 2022
2022
-
[60]
Memory-augmented graph neural networks: A brain-inspired review
Guixiang Ma, Vy A Vo, Theodore L Willke, and Nesreen K Ahmed. Memory-augmented graph neural networks: A brain-inspired review. IEEE Transactions on Artificial Intelligence, 5 0 (5): 0 2011--2025, 2023
2011
-
[61]
The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects, 2013
Martial Mermillod, Aur \'e lia Bugaiska, and Patrick Bonin. The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects, 2013
2013
-
[62]
The magical number seven, plus or minus two: Some limits on our capacity for processing information
George A Miller. The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological review, 63 0 (2): 0 81, 1956
1956
-
[63]
o ksal, Ayyoob Imani, Mohsen Fayyaz, and Hinrich Sch \
Ali Modarressi, Abdullatif K \"o ksal, Ayyoob Imani, Mohsen Fayyaz, and Hinrich Sch \"u tze. Memllm: Finetuning llms to use an explicit read-write memory. arXiv preprint arXiv:2404.11672, 2024
Pith/arXiv arXiv 2024
-
[64]
Episodic memory and beyond: the hippocampus and neocortex in transformation
Morris Moscovitch, Roberto Cabeza, Gordon Winocur, and Lynn Nadel. Episodic memory and beyond: the hippocampus and neocortex in transformation. Annual review of psychology, 67 0 (1): 0 105--134, 2016
2016
-
[65]
Memory: Neurobiological mechanisms and assessment
Swaleha Mujawar, Jaideep Patil, Bhushan Chaudhari, and Daniel Saldanha. Memory: Neurobiological mechanisms and assessment. Industrial psychiatry journal, 30 0 (Suppl 1): 0 S311--S314, 2021
2021
-
[66]
Most brain activity is ‘background noise’—and that’s upending our understanding of consciousness, 2021
Thomas Nail. Most brain activity is ‘background noise’—and that’s upending our understanding of consciousness, 2021
2021
-
[67]
Dynamic memory compression: Retrofitting llms for accelerated inference
Piotr Nawrot, Adrian a \'n cucki, Marcin Chochowski, David Tarjan, and Edoardo M Ponti. Dynamic memory compression: Retrofitting llms for accelerated inference. arXiv preprint arXiv:2403.09636, 2024
Pith/arXiv arXiv 2024
-
[68]
Functional role of gamma and theta oscillations in episodic memory
Erika Nyhus and Tim Curran. Functional role of gamma and theta oscillations in episodic memory. Neuroscience & Biobehavioral Reviews, 34 0 (7): 0 1023--1035, 2010
2010
-
[69]
Memgpt: Towards llms as operating systems
Charles Packer, Vivian Fang, Shishir\_G Patil, Kevin Lin, Sarah Wooders, and Joseph\_E Gonzalez. Memgpt: Towards llms as operating systems. 2023
2023
-
[70]
ABC: Attention with Bounded-memory Control
Hao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, and Noah A Smith. Abc: Attention with bounded-memory control. arXiv preprint arXiv:2110.02488, 2021
work page internal anchor Pith review Pith/arXiv arXiv 2021
-
[71]
Position: Episodic memory is the missing piece for long-term llm agents
Mathis Pink, Qinyuan Wu, Vy Ai Vo, Javier Turek, Jianing Mu, Alexander Huth, and Mariya Toneva. Position: Episodic memory is the missing piece for long-term llm agents. arXiv preprint arXiv:2502.06975, 2025
Pith/arXiv arXiv 2025
-
[72]
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409, 2021
Pith/arXiv arXiv 2021
-
[73]
Neuromodulation of the feedforward dentate gyrus-ca3 microcircuit
Luke Y Prince, Travis J Bacon, Cezar M Tigaret, and Jack R Mellor. Neuromodulation of the feedforward dentate gyrus-ca3 microcircuit. Frontiers in synaptic neuroscience, 8: 0 32, 2016
2016
-
[74]
Layerwise recurrent router for mixture-of-experts
Zihan Qiu, Zeyu Huang, Shuang Cheng, Yizhi Zhou, Zili Wang, Ivan Titov, and Jie Fu. Layerwise recurrent router for mixture-of-experts. arXiv preprint arXiv:2408.06793, 2024
Pith/arXiv arXiv 2024
-
[75]
A multisensory perspective of working memory
Michel Quak, Raquel Elea London, and Durk Talsma. A multisensory perspective of working memory. Frontiers in human neuroscience, 9: 0 197, 2015
2015
-
[76]
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap. Compressive transformers for long-range sequence modelling. arXiv preprint arXiv:1911.05507, 2019
Pith/arXiv arXiv 1911
-
[77]
A default mode of brain function
Marcus E Raichle, Ann Mary MacLeod, Abraham Z Snyder, William J Powers, Debra A Gusnard, and Gordon L Shulman. A default mode of brain function. Proceedings of the national academy of sciences, 98 0 (2): 0 676--682, 2001
work page 2001
-
[78]
J Ranjith and Santhi Baskaran. Adaptive knowledge consolidation: A dynamic approach to mitigating catastrophic forgetting in text-based neural networks. 2024
work page 2024
-
[79]
Hierarchical dynamics as a macroscopic organizing principle of the human brain
Ryan V Raut, Abraham Z Snyder, and Marcus E Raichle. Hierarchical dynamics as a macroscopic organizing principle of the human brain. Proceedings of the National Academy of Sciences, 117 0 (34): 0 20890--20897, 2020
work page 2020
-
[80]
Daniel Reznik, Robert Trampel, Nikolaus Weiskopf, Menno P Witter, and Christian F Doeller. Dissociating distinct cortical networks associated with subregions of the human medial temporal lobe using precision neuroimaging. Neuron, 111 0 (17): 0 2756--2772, 2023
work page 2023
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.