REVIEW 45 cited by
Memory Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We describe a new class of learning models called memory networks. Memory networks reason with inference components combined with a long-term memory component; they learn how to use these jointly. The long-term memory can be read and written to, with the goal of using it for prediction. We investigate these models in the context of question answering (QA) where the long-term memory effectively acts as a (dynamic) knowledge base, and the output is a textual response. We evaluate them on a large-scale QA task, and a smaller, but more complex, toy task generated from a simulated world. In the latter, we show the reasoning power of such models by chaining multiple supporting sentences to answer questions that require understanding the intension of verbs.
Forward citations
Cited by 45 Pith papers
-
REALM: Retrieval-Augmented Language Model Pre-Training
REALM augments language-model pre-training with an unsupervised retriever over Wikipedia documents and reports 4-16% absolute gains on open-domain QA benchmarks over prior implicit and explicit knowledge methods.
-
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
Per-frame natural-language action prompts enable simultaneous multi-entity control and cross-entity action transfer in interactive video world models, outperforming discrete action-index interfaces.
-
Beyond Detection: A Structure-Aware Framework for Scene Text Tracking
SymTrack is the first systematic detection-free framework for scene text tracking that constructs benchmarks from video text spotting datasets and reports up to 11.97% AUC gains over prior trackers.
-
Learning Compositional Functions with Transformers from Easy-to-Hard Data
A transformer with O(log k) layers provably learns the k-fold permutation composition task in poly(N,k) samples with curriculum or mixed easy-to-hard data, despite an SQ lower bound requiring N^{Omega(k)} samples on h...
-
Graph Retention Networks for Dynamic Graphs
Graph Retention Networks extend retention to dynamic graphs to enable parallelizable training, O(1) inference, and chunkwise long-term training while delivering competitive performance with major efficiency gains.
-
Semantic Role Labeling with Associated Memory Network
A neural SRL model that attends to labels of similar training sentences via an associated memory network reaches 89.6 F1 on CoNLL-2009 English in-domain, 79.7 on Brown, and 83.8 on Chinese, with small, consistent gain...
-
Metis: Memory Foundation Model
Metis puts a trainable fixed-size memory matrix inside a frozen LLM backbone and learns to remember, update, forget, and reflect across turns without replaying original context.
-
MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning
A typed, editable memory built from egocentric video improves memory-grounded question answering and out-of-distribution robot planning over flat-text and graph baselines.
-
BiDeMem: Bidirectional Degradation Memory for Explainable Image Restoration
BiDeMem retrieves compact memory slots via a query from restoration features to jointly improve restoration quality and provide a falsifiable degradation explanation path in a controlled NAFNet multi-degradation setting.
-
Repeated Shared Access Enables Grokking, but Edit Propagation Depends on an Addressable Memory
A 2x2 ablation shows repeated shared access enables grokking while addressable memory (not recurrence) enables edit propagation in transformer variants on synthetic KG QA.
-
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning
ThoughtFold applies introspective redundancy detection within correct CoT trajectories to create sub-trajectory spectra, then uses masked preference optimization to penalize redundant explorations, yielding 56% token ...
-
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
In a cellular automata rule-inference task designed to block memorization, neural models achieve high next-step accuracy but accuracy falls sharply with longer reasoning chains; depth, recurrence, memory, and test-tim...
-
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
CEM-Net stores cross-emotion expression displacements in a memory bank so a generated talking face matches the emotion in the audio even when the reference image emotion conflicts.
-
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
A two-stage key-value memory model generates personalized 3D facial animation from audio alone, topping prior methods on FVE, LVE, FID, LDTW, and lip-max on VOCASET and BIWI.
-
Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection
Sketch-based support can replace photo support in few-shot keypoint detection, beating an adapted photo-based baseline on novel classes.
-
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
MemAgent uses multi-conversation RL to train a memory agent that reads text in segments and overwrites memory, extrapolating from 8K training to 3.5M token QA with under 5% loss and 95%+ on 512K RULER.
-
Ella: Embodied Social Agents with Lifelong Memory
Ella, an embodied social agent with a name-centric semantic memory and a spatiotemporal episodic memory, outperformed two re-implemented baselines in social influence and leadership tasks in a 3D simulation.
-
A New Perspective On AI Safety Through Control Theory Methodologies
This paper outlines a new conceptual paradigm, data control, which transfers control-theoretic system analysis and properties to AI systems to support generic AI safety assurance.
-
Low-Rank Head Avatar Personalization with Registers
A Register Module, a learnable 3D feature space rigged to a 3DMM mesh, improves LoRA-based personalization of head avatars by teaching the model to focus on identity-specific DINOv2 features during adaptation.
-
Lego Sketch: A Scalable Memory-augmented Neural Network for Sketching Data Streams
The Lego sketch is a scalable neural sketch that partitions a data stream into multiple memory bricks via hashing, with a Deep Sets-based scanning module and self-guided loss to improve frequency estimation.
-
AI for the Open-World: the Learning Principles
Open-world AI requires rich features, disentangled representations, and inference-time learning; the thesis presents techniques and large-scale experiments supporting these principles.
-
Titans: Learning to Memorize at Test Time
Titans combine attention for current context with a learnable neural memory for long-term history, achieving better performance and scaling to over 2M-token contexts on language, reasoning, genomics, and time-series tasks.
-
Memory Layers at Scale
A scaled-up, shared, gated memory layer improves factual recall in language models and can match or beat denser models trained with much more compute.
-
3D Reconstruction with Spatial Memory
Spann3R uses a learned spatial memory to regress per-image pointmaps directly in a shared global coordinate system, removing the need for optimization-based alignment after per-pair predictions.
-
Cognitive Architectures for Language Agents
CoALA is a modular cognitive architecture for language agents that organizes memory components, action spaces for internal and external interaction, and a generalized decision-making loop to support more systematic de...
-
Compressive Transformers for Long-Range Sequence Modelling
Compressive Transformer sets new records on WikiText-103 (17.1 ppl) and Enwik8 (0.97 bpc) via memory compression and introduces the PG-19 long-range language benchmark.
-
Parallel Causal Associative Fields: Gated Sparse Memory for Long-Context Language Modeling
PCAF uses parallel hash-based associative memory over causal records plus a learned gate to reach 36.31 perplexity on WikiText-103 and 52.45 on PG-19 at 303M parameters and 2048 context while running at 0.61M tokens/s...
-
RowNet: A Memory Transformer for Tabular Regression
RowNet uses a memory bank of labeled properties, two retrieval layers with attention, and a mixture-of-experts module to predict real estate price per square meter.
-
BatteryMFormer: Multi-level Learning for Battery Degradation Trajectory Forecasting
BatteryMFormer is a multi-level Transformer that adds an aging-condition-aware decoder, meta degradation pattern memory, and dual-view encoder to forecast battery state-of-health trajectories from early operational da...
-
PRGCN: A Graph Memory Network for Cross-Sequence Pattern Reuse in 3D Human Pose Estimation
PRGCN, a graph-memory network with a Mamba-attention dual stream, reports state-of-the-art MPJPE of 37.1 mm on Human3.6M and 13.4 mm on MPI-INF-3DHP.
-
ST-Hyper: Learning High-Order Dependencies Across Multiple Spatial-Temporal Scales for Multivariate Time Series Forecasting
ST-Hyper combines spatial-temporal pyramid feature extraction with adaptive sparse hypergraph learning and tri-phase propagation to achieve state-of-the-art results on six multivariate time series forecasting benchmarks.
-
Beyond Attention: Toward Machines with Intrinsic Higher Mental States
The Co4 mechanism adds triadic Q-K-V modulation loops before attention, claiming O(N) complexity and much faster learning than standard Transformers on small benchmarks.
-
Logarithmic Memory Networks (LMNs): Efficient Long-Range Sequence Modeling for Resource-Constrained Environments
Logarithmic Memory Networks are a new hierarchical-memory architecture claiming O(log n) attention, but the complexity math is internally inconsistent and the empirical evidence is a single small run.
-
A Novel Aspect-Guided Deep Transition Model for Aspect Based Sentiment Analysis
AGDT, an aspect-guided deep transition GRU with an aspect-reconstruction loss, improves non-BERT state-of-the-art accuracy on four SemEval aspect-based sentiment analysis datasets.
-
ASNets: Deep Learning for Generalised Planning
ASNets learn generalized planning policies from small instances and solve all 18,300 large Blocksworld test instances after training on 50 small ones.
-
Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance
A deterministic 46-block runtime wrapper around LLMs with pre-response safety gates and impact-weighted cache eviction reports reproducible control signals and modest efficiency gains in scoped single-instance tests.
-
Contextual Memory Intelligence -- A Foundational Paradigm for Human-AI Collaboration and Reflective Generative AI Systems
Contextual Memory Intelligence reframes memory as dynamic infrastructure and proposes the Insight Layer to preserve decision rationale, detect semantic drift, and support human-in-the-loop reflection.
-
Unstructured Text Enhanced Open-domain Dialogue System: A Systematic Survey
A structured survey of dialogue systems that use unstructured text as external knowledge, organizing datasets, retrieval and generative model components, evaluation metrics, and future directions.
-
Ensemble approach for natural language question answering problem
A class-weighted voting ensemble of BiDAF, QANet, and Mnemonic Reader reports F1 81.96 and EM 73.77 on the SQuAD dev set, marginally above Mnemonic Reader's 81.57 and 73.25.
-
Knowledge Query Network: How Knowledge Interacts with Skills
KQN predicts student correctness as the dot product of a student knowledge vector and a skill vector, and claims the distances between learned skill vectors reveal how related skills are.
-
Asset Pricing in Pre-trained Transformer
A new encoder-only Transformer variant with autoencoder pre-training is reported to achieve high out-of-sample R2 for US stock returns, but the headline numbers come from test-set model selection.
-
PersonaAI: Leveraging Retrieval-Augmented Generation and Personalized Context for AI-Driven Digital Avatars
PersonaAI uses standard RAG with LLAMA to answer as a user based on retrieved personal text, but the evaluation is anecdotal and the 'accurate personality mimicry' claim is unproven.
-
Evolutionary Algorithm for Sinhala to English Translation
An evolutionary algorithm identifies meanings in Sinhala sentences to produce English translations that are then grammatically corrected, reported to yield accurate results.
-
Machine Reading Comprehension: a Literature Review
A 2019 survey of machine reading comprehension corpora and methods.
-
Survey on Deep Neural Networks in Speech and Vision Systems
A broad survey of deep learning architectures and systems for vision and speech, with an emphasis on mobile deployment and emerging applications, containing no new results.
Discussion (0). Continue with ORCID to comment.