Pith. sign in

REVIEW 4 major objections 1 minor 2 cited by

A three-axis taxonomy organizes Memory-Augmented Transformers and shows a shift from static caches to adaptive test-time learning.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Memory-augmented Transformer research is organized into a three-axis taxonomy bridging neuroscience memory concepts to network designs, but no new result is produced.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection A coherent survey abstract atop an unreadable body: worth sending back for a clean copy and a methods section, not yet citable. the 4 major comments →

arxiv 2508.10824 v2 pith:VDNDJEAA submitted 2025-08-14 cs.LG cs.CL

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures

classification cs.LG cs.CL
keywords memory-augmented Transformersneuroscience-inspired AIthree-axis taxonomytest-time learninglifelong learningsequence modelingconsolidationsurvey
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review tries to establish that the scattered engineering of memory-augmented Transformers is not a bag of tricks: recent work can be organized along three axes—what the memory is for, how it is represented, and how it is integrated—and those axes, read through neuroscience ideas about multi-timescale memory, selective attention, and consolidation, show a field-level trajectory from static caches toward adaptive, test-time learning. A sympathetic reader would care because if the taxonomy holds, it gives researchers a shared vocabulary for comparing architectures and points the design space toward lifelong-learning Transformers, with hierarchical buffering and surprise-gated updates as explicit next steps. The paper's contribution is organizational and diagnostic, not a new model: it claims to show where the field has been and where it is heading.

Core claim

The paper's central claim is that the diverse memory mechanisms bolted onto Transformers form a coherent design space, not a collection of unrelated patches. It proposes a three-axis taxonomy: functional objectives (context extension, reasoning, knowledge integration, adaptation), memory representations (parameter-encoded, state-based, explicit, hybrid), and integration mechanisms (attention fusion, gated control, associative retrieval). Using reading, writing, forgetting, and capacity management as the core memory operations, it argues that the field is shifting from fixed external caches toward adaptive systems that update their own memory at test time. It further claims that neuroscience

What carries the argument

The organizing device is the taxonomy itself: each memory-augmented Transformer is located by a triple—objective, representation, integration. The four memory operations (reading, writing, forgetting, capacity management) are the analytic lens that turns the taxonomy from a static list into a diagnosis of what each architecture can and cannot do. The neuroscience mapping supplies the design rationale: multi-timescale memory justifies separate memory stores, selective attention justifies content-based retrieval, and consolidation justifies writing and forgetting rules—so the taxonomy is meant to be explanatory, not merely descriptive.

Load-bearing premise

The framework stands or falls on whether the neuroscience constructs genuinely correspond to the engineering categories and whether the surveyed papers were selected systematically rather than because they fit the story.

What would settle it

Take a defined literature window, for example memory-augmented Transformer papers published 2023–2025, collect the full set with a fixed search query, and try to assign each paper to exactly one cell of the three-axis taxonomy. If more than a small fraction resist placement, or the full set shows no chronological trend toward test-time learning, the central claims fail.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Architectures that currently look incomparable—a long-context cache, a gated state model, an external associative memory—become comparable through their position on the three axes.
  • The claimed shift to adaptive, test-time learning sets a concrete design target: future models should write and update memory during inference, not only at training time.
  • Scalability and interference are named as the two bottlenecks any new memory design must address, so progress on them can be measured directly.
  • Hierarchical buffering and surprise-gated updates are named as emerging mechanisms, giving practitioners concrete starting points for new architectures.
  • The neuroscience link provides a vocabulary for generating new designs, such as treating a Transformer's memory as a multi-timescale system with distinct stores.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the taxonomy's usefulness depends on whether it can classify papers independently; a natural test is to take a held-out set of recent memory-augmented Transformer papers and see whether two annotators place them in the same taxonomy cells.
  • Editorial inference: if the shift-to-test-time-learning claim is right, evaluation practice should change—benchmarks that only test static retrieval will miss the capabilities that differentiate the newer architectures.
  • Editorial inference: the paper asserts rather than demonstrates its 'systematic' coverage; a reader should look for a reproducible search protocol, because without one the trend claim could reflect selection bias.
  • Editorial inference: the memory-operations framework likely generalizes beyond Transformers to any architecture with an explicit memory component, so the roadmap could also apply to state-space models or recurrent networks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 1 minor

Summary. The manuscript (arXiv:2508.10824) claims to provide a unified framework bridging neuroscience principles with engineering advances in Memory-Augmented Transformers, organized along three taxonomic dimensions (functional objectives, memory representations, integration mechanisms), and to identify a shift from static caches toward adaptive, test-time learning systems. The abstract is self-consistent and the claims are appropriately modest for a review. However, the supplied full text is heavily corrupted (mojibake) and unreadable; it also contains an inserted fragment from a different paper (arXiv:2508.10821v3 [q-bio.QM]). Consequently, none of the substantive content—sections, tables, equations, references, or the survey corpus—can be verified. The 'systematic' nature of the review and the empirical trend claim cannot be audited from the provided text.

Significance. If the claims hold, the three-axis taxonomy could provide a useful common vocabulary for memory-augmented transformers and a roadmap toward continual-learning architectures. The synthesis of neuroscience constructs and engineering designs could be valuable if the mapping is substantive. However, because the body is unreadable, the significance is conditional. The paper does not derive new mathematical results or ship code; its contribution, if any, is organizational and interpretive. Credit is due for a self-consistent abstract and clearly stated dimensions, but no machine-checkable evidence is available.

major comments (4)
  1. [Full text (after abstract)] The entire body of the manuscript is encoded mojibake; no section, table, equation, or reference can be read or verified. For a systematic review, the central claims—the taxonomy and the trend toward adaptive test-time learning—rest on the ability to check which papers were surveyed and how they were mapped. This is load-bearing and prevents any audit. I cannot determine whether the summaries are accurate or whether the conclusions follow from the cited literature.
  2. [Visible fragment in full text] The text includes the line 'arXiv:2508.10821v3 [q-bio.QM] 20 May 2026', which is an identifier for a different paper in quantitative biology. This indicates that the PDF/source has been corrupted or misassembled. As supplied, the document is not self-consistent; it contains content that does not belong to the claimed review. This is a material issue because it makes it impossible to attribute any of the surrounding text to the authors' intended manuscript.
  3. [Search protocol/inclusion criteria] The abstract calls the review 'systematic,' but no search strategy, inclusion/exclusion criteria, database sources, or PRISMA-style flow is visible in the supplied text. The trend claim ('a shift from static caches toward adaptive, test-time learning systems') is an empirical statement about the surveyed literature; without a reproducible corpus and coding protocol, it may be a cherry-picked narrative. This is an audibility problem, not necessarily an error, but it must be fixed for the review to be evaluable.
  4. [Neuroscience-to-engineering mapping] The 'unified framework' is the paper's main contribution, but the body that would justify the correspondence between biological constructs (multi-timescale memory, selective attention, consolidation) and architectural families (parameter-encoded, state-based, explicit, hybrid) is unreadable. This raises a correctness risk: the mapping could be substantive constraints or post-hoc labels. I cannot adjudicate without a legible exposition of how each architecture family instantiates each biological principle, including any criteria for correspondence.
minor comments (1)
  1. [Abstract] The abstract is readable and clear; the title matches the abstract. No further presentation issues can be assessed because the remainder is illegible. If the corruption is a rendering artifact, please resubmit a clean version.

Circularity Check

0 steps flagged

No significant circularity: the paper is a survey/taxonomy with no fitted parameters, derived equations, or load-bearing self-citations, so there is no derivation chain that reduces to its inputs.

full rationale

The manuscript is a systematic review and taxonomy paper. Its central claims are the proposed three-axis taxonomy (functional objectives, memory representations, integration mechanisms) and an interpretive trend statement that 'analysis ... reveals a shift from static caches toward adaptive, test-time learning systems.' Neither claim is derived from equations or fitted constants; the taxonomy is a classification scheme, and the trend is an inductive reading of the surveyed literature. There is no self-definitional step, no fitted input renamed as prediction, and no self-citation chain invoked to force a conclusion. The abstract and the legible fragments contain no derivation machinery for the trend claim, and the full text is largely mojibake, which makes the 'systematic' methodology unauditable; however, an inability to verify the survey corpus is an evidence/auditability concern, not a circularity of the kind defined here. The visible arXiv identifier (2508.10821v3 [q-bio.QM]) embedded in the text further indicates document corruption, but this does not constitute a logical loop. Accordingly, no specific circular step can be quoted with the required reduction, and the appropriate finding is no significant circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 2 axioms · 0 invented entities

As a review, the paper introduces no fitted numbers and no new physical or architectural entities. Its central claims rest on the accuracy of the primary literature it synthesizes and on the validity of the neuroscience mapping. With the full text unreadable, neither can be audited.

axioms (2)
  • domain assumption Cited papers report their results accurately
    The survey's summaries and trend claim inherit the correctness of the primary sources it synthesizes; no independent replication is possible from the provided text.
  • domain assumption The neuroscience-to-engineering correspondence is substantive
    The 'unified framework' in the abstract depends on multi-timescale memory, selective attention, and consolidation mapping onto the named architectural families as real design counterparts rather than as loose metaphor.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures." pith.science (2026). https://pith.science/paper/VDNDJEAA

@misc{pith2026250810824,
  author       = {Pith},
  title        = {Pith review of: Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VDNDJEAA}},
  note         = {Machine review of arXiv:2508.10824}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Memory is fundamental to intelligence, enabling learning, reasoning, and adaptability across biological and artificial systems. While Transformer architectures excel at sequence modeling, they face critical limitations in long-range context retention, continual learning, and knowledge integration. This review presents a unified framework bridging neuroscience principles, including dynamic multi-timescale memory, selective attention, and consolidation, with engineering advances in Memory-Augmented Transformers. We organize recent progress through three taxonomic dimensions: functional objectives (context extension, reasoning, knowledge integration, adaptation), memory representations (parameter-encoded, state-based, explicit, hybrid), and integration mechanisms (attention fusion, gated control, associative retrieval). Our analysis of core memory operations (reading, writing, forgetting, and capacity management) reveals a shift from static caches toward adaptive, test-time learning systems. We identify persistent challenges in scalability and interference, alongside emerging solutions including hierarchical buffering and surprise-gated updates. This synthesis provides a roadmap toward cognitively-inspired, lifelong-learning Transformer architectures.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Neural Procedural Memory: Empowering LLM Agents with Implicit Activation Steering

    cs.CL 2026-06 unverdicted novelty 6.0

    NPM distills contrastive experiences into implicit activation steering vectors that guide LLM agent execution comparably to explicit RAG instructions, with complementary gains when combined.

  2. Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance

    cs.CV 2025-09 reject novelty 3.0

    The proposed CAMVR framework is not supported by verifiable evidence, and the manuscript itself labels its experimental results as fabricated.

Reference graph

Works this paper leans on

124 extracted references · 33 canonical work pages · cited by 2 Pith papers · 10 internal anchors

  1. [1]

    Adult hippocampal neurogenesis and cognitive flexibility—linking memory and mood

    Christoph Anacker and Ren \'e Hen. Adult hippocampal neurogenesis and cognitive flexibility—linking memory and mood. Nature Reviews Neuroscience, 18 0 (6): 0 335--346, 2017

  2. [2]

    Global workspace theory (gwt) and prefrontal cortex: Recent developments

    Bernard J Baars, Natalie Geld, and Robert Kozma. Global workspace theory (gwt) and prefrontal cortex: Recent developments. Frontiers in psychology, 12: 0 749868, 2021

  3. [3]

    Working memory: looking back and looking forward

    Alan Baddeley. Working memory: looking back and looking forward. Nature reviews neuroscience, 4 0 (10): 0 829--839, 2003

  4. [4]

    Fast adaptation to rule switching using neuronal surprise

    Martin LLR Barry and Wulfram Gerstner. Fast adaptation to rule switching using neuronal surprise. PLoS computational biology, 20 0 (2): 0 e1011839, 2024

  5. [5]

    Neuromodulators and long-term synaptic plasticity in learning and memory: A steered-glutamatergic perspective

    Amjad H Bazzari and H Rheinallt Parri. Neuromodulators and long-term synaptic plasticity in learning and memory: A steered-glutamatergic perspective. Brain sciences, 9 0 (11): 0 300, 2019

  6. [6]

    Titans: Learning to memorize at test time

    Ali Behrouz, Peilin Zhong, and Vahab Mirrokni. Titans: Learning to memorize at test time. arXiv preprint arXiv:2501.00663, 2024

  7. [7]

    Atlas: Learning to optimally memorize the context at test time

    Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, and Vahab Mirrokni. Atlas: Learning to optimally memorize the context at test time. arXiv preprint arXiv:2505.23735, 2025

  8. [8]

    Longformer: The long-document transformer

    Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020

  9. [9]

    Memory layers at scale

    Vincent-Pierre Berges, Barlas O g uz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, and Gargi Ghosh. Memory layers at scale. arXiv preprint arXiv:2412.09764, 2024

  10. [10]

    Improving language models by retrieving from trillions of tokens

    Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pp.\ 2206--2240. PMLR, 2022

  11. [11]

    The neuroanatomical, neurophysiological and psychological basis of memory: Current models and their origins

    Eduardo Camina and Francisco G \"u ell. The neuroanatomical, neurophysiological and psychological basis of memory: Current models and their origins. Frontiers in pharmacology, 8: 0 438, 2017

  12. [12]

    An evolved universal transformer memory

    Edoardo Cetin, Qi Sun, Tianyu Zhao, and Yujin Tang. An evolved universal transformer memory. In The Thirteenth International Conference on Learning Representations, 2025

  13. [13]

    Walking down the memory maze: Beyond context limit through interactive reading

    Howard Chen, Ramakanth Pasunuru, Jason Weston, and Asli Celikyilmaz. Walking down the memory maze: Beyond context limit through interactive reading. arXiv preprint arXiv:2310.05029, 2023

  14. [14]

    Mem0: Building production-ready ai agents with scalable long-term memory

    Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready ai agents with scalable long-term memory. arXiv preprint arXiv:2504.19413, 2025

  15. [15]

    Interactions between attention and memory

    Marvin M Chun and Nicholas B Turk-Browne. Interactions between attention and memory. Current opinion in neurobiology, 17 0 (2): 0 177--184, 2007

  16. [16]

    Attention-dependent coupling with forebrain and brainstem neuromodulatory nuclei differs across the lifespan

    Nicholas G Cicero, Elizabeth Riley, Khena M Swallow, Eve De Rosa, and Adam Anderson. Attention-dependent coupling with forebrain and brainstem neuromodulatory nuclei differs across the lifespan. GeroScience, pp.\ 1--20, 2025

  17. [17]

    What are the differences between long-term, short-term, and working memory? Progress in brain research, 169: 0 323--338, 2008

    Nelson Cowan. What are the differences between long-term, short-term, and working memory? Progress in brain research, 169: 0 323--338, 2008

  18. [18]

    Transformer-xl: Attentive language models beyond a fixed-length context

    Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019

  19. [19]

    The global neuronal workspace model of conscious access: from neuronal architectures to clinical applications

    Stanislas Dehaene, Jean-Pierre Changeux, and Lionel Naccache. The global neuronal workspace model of conscious access: from neuronal architectures to clinical applications. Characterizing consciousness: From cognition to the clinic?, pp.\ 55--84, 2011

  20. [20]

    Pronouns reactivate conceptual representations in human hippocampal neurons

    Doris E Dijksterhuis, Matthew W Self, Jessy K Possel, Judith C Peters, ECW van Straaten, Sander Idema, Johannes C Baaijen, Sandra MA van der Salm, Erik J Aarnoutse, Nicole CE van Klink, et al. Pronouns reactivate conceptual representations in human hippocampal neurons. Science, 385 0 (6716): 0 1478--1484, 2024

  21. [21]

    Interaction between the amygdala and the medial temporal lobe memory system predicts better memory for emotional events

    Florin Dolcos, Kevin S LaBar, and Roberto Cabeza. Interaction between the amygdala and the medial temporal lobe memory system predicts better memory for emotional events. Neuron, 42 0 (5): 0 855--863, 2004

  22. [22]

    Rethinking memory in ai: Taxonomy, operations, topics, and future directions

    Yiming Du, Wenyu Huang, Danna Zheng, Zhaowei Wang, Sebastien Montella, Mirella Lapata, Kam-Fai Wong, and Jeff Z Pan. Rethinking memory in ai: Taxonomy, operations, topics, and future directions. arXiv preprint arXiv:2505.00675, 2025

  23. [23]

    Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers

    Sanchayan Dutta and Suvrit Sra. Memory-augmented transformers can implement linear first-order optimization methods. arXiv preprint arXiv:2410.07263, 2024

  24. [24]

    Efficient llm inference using dynamic input pruning and cache-aware masking

    Marco Federici, Davide Belli, Mart Van Baalen, Amir Jalalirad, Andrii Skliar, Bence Major, Markus Nagel, and Paul Whatmough. Efficient llm inference using dynamic input pruning and cache-aware masking. arXiv preprint arXiv:2412.01380, 2024

  25. [25]

    Human-like episodic memory for infinite context llms

    Zafeirios Fountas, Martin A Benfeghoul, Adnan Oomerjee, Fenia Christopoulou, Gerasimos Lampouras, Haitham Bou-Ammar, and Jun Wang. Human-like episodic memory for infinite context llms. arXiv preprint arXiv:2407.09450, 2024

  26. [26]

    Experiencing surprise: The temporal dynamics of its impact on memory

    Darya Frank, Alex Kafkas, and Daniela Montaldi. Experiencing surprise: The temporal dynamics of its impact on memory. Journal of Neuroscience, 42 0 (33): 0 6435--6444, 2022

  27. [27]

    An efficient context-dependent memory framework for llm-centric agents

    Pengyu Gao, Jinming Zhao, Xinyue Chen, and Long Yilin. An efficient context-dependent memory framework for llm-centric agents. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track), pp.\ 1055--1069, 2025

  28. [28]

    Memotr: Long-term memory-augmented transformer for multi-object tracking

    Ruopeng Gao and Limin Wang. Memotr: Long-term memory-augmented transformer for multi-object tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.\ 9901--9910, 2023

  29. [29]

    Top-down modulation: bridging selective attention and working memory

    Adam Gazzaley and Anna C Nobre. Top-down modulation: bridging selective attention and working memory. Trends in cognitive sciences, 16 0 (2): 0 129--135, 2012

  30. [30]

    zip2zip: Inference-time adaptive vocabularies for language models via token compression

    Saibo Geng, Nathan Ranchin, Maxime Peyrard, Chris Wendler, Michael Gastpar, Robert West, et al. zip2zip: Inference-time adaptive vocabularies for language models via token compression. arXiv preprint arXiv:2506.01084, 2025

  31. [31]

    The role of the ca3 hippocampal subregion in spatial memory: a process oriented behavioral assessment

    Paul E Gilbert and Andrea M Brushfield. The role of the ca3 hippocampal subregion in spatial memory: a process oriented behavioral assessment. Progress in Neuro-Psychopharmacology and Biological Psychiatry, 33 0 (5): 0 774--781, 2009

  32. [32]

    Neural turing machines

    Alex Graves, Greg Wayne, and Ivo Danihelka. Neural turing machines. arXiv preprint arXiv:1410.5401, 2014

  33. [33]

    Hybrid computing using a neural network with dynamic external memory

    Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwi \'n ska, Sergio G \'o mez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al. Hybrid computing using a neural network with dynamic external memory. Nature, 538 0 (7626): 0 471--476, 2016

  34. [34]

    Hipporag: Neurobiologically inspired long-term memory for large language models

    Bernal Jim \'e nez Guti \'e rrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. Hipporag: Neurobiologically inspired long-term memory for large language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  35. [35]

    Hierarchical process memory: memory as an integral component of information processing

    Uri Hasson, Janice Chen, and Christopher J Honey. Hierarchical process memory: memory as an integral component of information processing. Trends in cognitive sciences, 19 0 (6): 0 304--313, 2015

  36. [36]

    Memory matters: The need to improve long-term memory in llm-agents

    Kostas Hatalis, Despina Christou, Joshua Myers, Steven Jones, Keith Lambert, Adam Amos-Binks, Zohreh Dannenhauer, and Dustin Dannenhauer. Memory matters: The need to improve long-term memory in llm-agents. In Proceedings of the AAAI Symposium Series, volume 2, pp.\ 277--280, 2023

  37. [37]

    Hmt: Hierarchical memory transformer for efficient long context language processing

    Zifan He, Yingqi Cao, Zongyue Qin, Neha Prakriya, Yizhou Sun, and Jason Cong. Hmt: Hierarchical memory transformer for efficient long context language processing. arXiv preprint arXiv:2405.06067, 2024 a

  38. [38]

    Human-inspired perspectives: A survey on ai long-term memory

    Zihong He, Weizhe Lin, Hao Zheng, Fan Zhang, Matt W Jones, Laurence Aitchison, Xuhai Xu, Miao Liu, Per Ola Kristensson, and Junxiao Shen. Human-inspired perspectives: A survey on ai long-term memory. arXiv preprint arXiv:2411.00489, 2024 b

  39. [39]

    Replay bursts in humans coincide with activation of the default mode and parietal alpha networks

    Cameron Higgins, Yunzhe Liu, Diego Vidaurre, Zeb Kurth-Nelson, Ray Dolan, Timothy Behrens, and Mark Woolrich. Replay bursts in humans coincide with activation of the default mode and parietal alpha networks. Neuron, 109 0 (5): 0 882--893, 2021

  40. [40]

    Transformerfam: Feedback attention is working memory

    Dongseong Hwang, Weiran Wang, Zhuoyuan Huo, Khe Chai Sim, and Pedro Moreno Mengibar. Transformerfam: Feedback attention is working memory. arXiv preprint arXiv:2404.09173, 2024

  41. [41]

    Memory os of ai agent

    Jiazheng Kang, Mingming Ji, Zhe Zhao, and Ting Bai. Memory os of ai agent. arXiv preprint arXiv:2506.06326, 2025 a

  42. [42]

    Lm2: Large memory models

    Jikun Kang, Wenqi Wu, Filippos Christianos, Alex J Chan, Fraser Greenlee, George Thomas, Marvin Purtorab, and Andy Toulis. Lm2: Large memory models. arXiv preprint arXiv:2502.06049, 2025 b

  43. [43]

    Distinguishing examples while building concepts in hippocampal and artificial networks

    Louis Kang and Taro Toyoizumi. Distinguishing examples while building concepts in hippocampal and artificial networks. Nature Communications, 15 0 (1): 0 647, 2024

  44. [44]

    Mechanisms of systems memory consolidation during sleep

    Jens G Klinzing, Niels Niethard, and Jan Born. Mechanisms of systems memory consolidation during sleep. Nature neuroscience, 22 0 (10): 0 1598--1610, 2019

  45. [45]

    Memreasoner: A memory-augmented llm architecture for multi-hop reasoning

    Ching-Yun Ko, Sihui Dai, Payel Das, Georgios Kollias, Subhajit Chaudhury, and Aurelie Lozano. Memreasoner: A memory-augmented llm architecture for multi-hop reasoning. In The First Workshop on System-2 Reasoning at Scale, NeurIPS'24, 2024

  46. [46]

    Flexible working memory through selective gating and attentional tagging

    Wouter Kruijne, Sander M Bohte, Pieter R Roelfsema, and Christian NL Olivers. Flexible working memory through selective gating and attentional tagging. Neural Computation, 33 0 (1): 0 1--40, 2021

  47. [47]

    Semantic memory: A review of methods, models, and current challenges

    Abhilasha A Kumar. Semantic memory: A review of methods, models, and current challenges. Psychonomic bulletin & review, 28 0 (1): 0 40--80, 2021

  48. [48]

    Self-attentive associative memory

    Hung Le, Truyen Tran, and Svetha Venkatesh. Self-attentive associative memory. In International conference on machine learning, pp.\ 5682--5691. PMLR, 2020

  49. [49]

    Reasoning under 1 billion: Memory-augmented reinforcement learning for large language models

    Hung Le, Dai Do, Dung Nguyen, and Svetha Venkatesh. Reasoning under 1 billion: Memory-augmented reinforcement learning for large language models. arXiv preprint arXiv:2504.02273, 2025

  50. [50]

    MATTER: Memory-Augmented Transformer Using Heterogeneous Knowledge Sources

    Dongkyu Lee, Chandana Satya Prakash, Jack FitzGerald, and Jens Lehmann. Matter: Memory-augmented transformer using heterogeneous knowledge sources. arXiv preprint arXiv:2406.04670, 2024

  51. [51]

    An update on memory reconsolidation updating

    Jonathan LC Lee, Karim Nader, and Daniela Schiller. An update on memory reconsolidation updating. Trends in cognitive sciences, 21 0 (7): 0 531--545, 2017

  52. [52]

    u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K \"u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \"a schel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33: 0 9459--9474, 2020

  53. [53]

    FALCON: Feedback-driven Adaptive Long/short-term memory reinforced Coding Optimization system

    Zeyuan Li, Yangfan He, Lewei He, Jianhui Wang, Tianyu Shi, Bin Lei, Yuchen Li, and Qiuwu Chen. Falcon: Feedback-driven adaptive long/short-term memory reinforced coding optimization system. arXiv preprint arXiv:2410.21349, 2024

  54. [54]

    Self-evolving agents with reflective and memory-augmented abilities

    Xuechen Liang, Yangfan He, Yinghui Xia, Xinyuan Song, Jianhui Wang, Meiling Tao, Li Sun, Xinhang Yuan, Jiayi Su, Keqin Li, et al. Self-evolving agents with reflective and memory-augmented abilities. arXiv preprint arXiv:2409.00872, 2024

  55. [55]

    Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems

    Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang, Tanjin He, Sirui Hong, Hongzhang Liu, Shaokun Zhang, Kaitao Song, Kunlun Zhu, et al. Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems. arXiv preprint arXiv:2504.01990, 2025

  56. [56]

    Think-in-memory: Recalling and post-thinking enable llms with long-term memory

    Lei Liu, Xiaoyan Yang, Yue Shen, Binbin Hu, Zhiqiang Zhang, Jinjie Gu, and Guannan Zhang. Think-in-memory: Recalling and post-thinking enable llms with long-term memory. arXiv preprint arXiv:2311.08719, 2023

  57. [57]

    MemLong: Memory-Augmented Retrieval for Long Text Modeling

    Weijie Liu, Zecheng Tang, Juntao Li, Kehai Chen, and Min Zhang. Memlong: Memory-augmented retrieval for long text modeling. arXiv preprint arXiv:2408.16967, 2024

  58. [58]

    Experience replay is associated with efficient nonlocal learning

    Yunzhe Liu, Marcelo G Mattar, Timothy EJ Behrens, Nathaniel D Daw, and Raymond J Dolan. Experience replay is associated with efficient nonlocal learning. Science, 372 0 (6544): 0 eabf1357, 2021

  59. [59]

    Acquiring new memories in neocortex of hippocampal-lesioned mice

    Wenhan Luo, Di Yun, Yi Hu, Miaomiao Tian, Jiajun Yang, Yifan Xu, Yong Tang, Yang Zhan, Hong Xie, and Ji-Song Guan. Acquiring new memories in neocortex of hippocampal-lesioned mice. Nature communications, 13 0 (1): 0 1601, 2022

  60. [60]

    Memory-augmented graph neural networks: A brain-inspired review

    Guixiang Ma, Vy A Vo, Theodore L Willke, and Nesreen K Ahmed. Memory-augmented graph neural networks: A brain-inspired review. IEEE Transactions on Artificial Intelligence, 5 0 (5): 0 2011--2025, 2023

  61. [61]

    The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects, 2013

    Martial Mermillod, Aur \'e lia Bugaiska, and Patrick Bonin. The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects, 2013

  62. [62]

    The magical number seven, plus or minus two: Some limits on our capacity for processing information

    George A Miller. The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological review, 63 0 (2): 0 81, 1956

  63. [63]

    o ksal, Ayyoob Imani, Mohsen Fayyaz, and Hinrich Sch \

    Ali Modarressi, Abdullatif K \"o ksal, Ayyoob Imani, Mohsen Fayyaz, and Hinrich Sch \"u tze. Memllm: Finetuning llms to use an explicit read-write memory. arXiv preprint arXiv:2404.11672, 2024

  64. [64]

    Episodic memory and beyond: the hippocampus and neocortex in transformation

    Morris Moscovitch, Roberto Cabeza, Gordon Winocur, and Lynn Nadel. Episodic memory and beyond: the hippocampus and neocortex in transformation. Annual review of psychology, 67 0 (1): 0 105--134, 2016

  65. [65]

    Memory: Neurobiological mechanisms and assessment

    Swaleha Mujawar, Jaideep Patil, Bhushan Chaudhari, and Daniel Saldanha. Memory: Neurobiological mechanisms and assessment. Industrial psychiatry journal, 30 0 (Suppl 1): 0 S311--S314, 2021

  66. [66]

    Most brain activity is ‘background noise’—and that’s upending our understanding of consciousness, 2021

    Thomas Nail. Most brain activity is ‘background noise’—and that’s upending our understanding of consciousness, 2021

  67. [67]

    Dynamic memory compression: Retrofitting llms for accelerated inference

    Piotr Nawrot, Adrian a \'n cucki, Marcin Chochowski, David Tarjan, and Edoardo M Ponti. Dynamic memory compression: Retrofitting llms for accelerated inference. arXiv preprint arXiv:2403.09636, 2024

  68. [68]

    Functional role of gamma and theta oscillations in episodic memory

    Erika Nyhus and Tim Curran. Functional role of gamma and theta oscillations in episodic memory. Neuroscience & Biobehavioral Reviews, 34 0 (7): 0 1023--1035, 2010

  69. [69]

    Memgpt: Towards llms as operating systems

    Charles Packer, Vivian Fang, Shishir\_G Patil, Kevin Lin, Sarah Wooders, and Joseph\_E Gonzalez. Memgpt: Towards llms as operating systems. 2023

  70. [70]

    ABC: Attention with Bounded-memory Control

    Hao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, and Noah A Smith. Abc: Attention with bounded-memory control. arXiv preprint arXiv:2110.02488, 2021

  71. [71]

    Position: Episodic memory is the missing piece for long-term llm agents

    Mathis Pink, Qinyuan Wu, Vy Ai Vo, Javier Turek, Jianing Mu, Alexander Huth, and Mariya Toneva. Position: Episodic memory is the missing piece for long-term llm agents. arXiv preprint arXiv:2502.06975, 2025

  72. [72]

    Train short, test long: Attention with linear biases enables input length extrapolation

    Ofir Press, Noah A Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409, 2021

  73. [73]

    Neuromodulation of the feedforward dentate gyrus-ca3 microcircuit

    Luke Y Prince, Travis J Bacon, Cezar M Tigaret, and Jack R Mellor. Neuromodulation of the feedforward dentate gyrus-ca3 microcircuit. Frontiers in synaptic neuroscience, 8: 0 32, 2016

  74. [74]

    Layerwise recurrent router for mixture-of-experts

    Zihan Qiu, Zeyu Huang, Shuang Cheng, Yizhi Zhou, Zili Wang, Ivan Titov, and Jie Fu. Layerwise recurrent router for mixture-of-experts. arXiv preprint arXiv:2408.06793, 2024

  75. [75]

    A multisensory perspective of working memory

    Michel Quak, Raquel Elea London, and Durk Talsma. A multisensory perspective of working memory. Frontiers in human neuroscience, 9: 0 197, 2015

  76. [76]

    Compressive transformers for long-range sequence modelling

    Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap. Compressive transformers for long-range sequence modelling. arXiv preprint arXiv:1911.05507, 2019

  77. [77]

    A default mode of brain function

    Marcus E Raichle, Ann Mary MacLeod, Abraham Z Snyder, William J Powers, Debra A Gusnard, and Gordon L Shulman. A default mode of brain function. Proceedings of the national academy of sciences, 98 0 (2): 0 676--682, 2001

  78. [78]

    Adaptive knowledge consolidation: A dynamic approach to mitigating catastrophic forgetting in text-based neural networks

    J Ranjith and Santhi Baskaran. Adaptive knowledge consolidation: A dynamic approach to mitigating catastrophic forgetting in text-based neural networks. 2024

  79. [79]

    Hierarchical dynamics as a macroscopic organizing principle of the human brain

    Ryan V Raut, Abraham Z Snyder, and Marcus E Raichle. Hierarchical dynamics as a macroscopic organizing principle of the human brain. Proceedings of the National Academy of Sciences, 117 0 (34): 0 20890--20897, 2020

  80. [80]

    Dissociating distinct cortical networks associated with subregions of the human medial temporal lobe using precision neuroimaging

    Daniel Reznik, Robert Trampel, Nikolaus Weiskopf, Menno P Witter, and Christian F Doeller. Dissociating distinct cortical networks associated with subregions of the human medial temporal lobe using precision neuroimaging. Neuron, 111 0 (17): 0 2756--2772, 2023

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.