REVIEW 11 cited by
Neural Map: Structured Memory for Deep Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
A critical component to enabling intelligent reasoning in partially observable environments is memory. Despite this importance, Deep Reinforcement Learning (DRL) agents have so far used relatively simple memory architectures, with the main methods to overcome partial observability being either a temporal convolution over the past k frames or an LSTM layer. More recent work (Oh et al., 2016) has went beyond these architectures by using memory networks which can allow more sophisticated addressing schemes over the past k frames. But even these architectures are unsatisfactory due to the reason that they are limited to only remembering information from the last k frames. In this paper, we develop a memory system with an adaptable write operator that is customized to the sorts of 3D environments that DRL agents typically interact with. This architecture, called the Neural Map, uses a spatially structured 2D memory image to learn to store arbitrary information about the environment over long time lags. We demonstrate empirically that the Neural Map surpasses previous DRL memories on a set of challenging 2D and 3D maze environments and show that it is capable of generalizing to environments that were not seen during training.
Forward citations
Cited by 11 Pith papers
-
CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence
CoCoSI is a training-free multi-agent system for collaborative cognitive map construction that improves spatial understanding in arbitrary pretrained MLLMs.
-
The Sword, Shield, and Achilles' Heel: Characterizing the Linguistic Inductive Bias of Large Language Models for Spatial Reasoning in Navigation Planning
Experiments reveal that topological cues robustly support LLM navigation planning while incorrect semantic cues derail it, with linguistic format effects varying by model size and compression.
-
Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers
Tensor Memory augments Transformers with a constant-size 3D voxel grid using differentiable soft writes at predicted locations, local interaction, and gated recurrent dynamics to decouple memory capacity from sequence length.
-
BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning
BrainMem equips LLM-based embodied planners with working, episodic, and semantic memory that evolves interaction histories into retrievable knowledge graphs and guidelines, raising success rates on long-horizon 3D benchmarks.
-
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments
Flow equivariant world models use a latent memory that shifts with the agent and with inferred object motion, giving stable long-horizon prediction under partial observability.
-
Situational Fusion of Visual Representation for Visual Navigation
Combining 25 frozen visual representations at the action level, with a task-affinity regularizer, roughly doubles success rate on unseen indoor navigation scenes compared with an ImageNet-pretrained ResNet baseline.
-
Cooperative Multi-UAV Navigation in Complex Environments via Systematic Multi-Agent Deep Reinforcement Learning
A modular MARL framework with local-optima override, hierarchical self-imitation cloning, dual-condition curriculum, and a structure-aware mixture-of-experts achieves 75% cooperative success on an unseen mixed maze in...
-
STMA: A Spatio-Temporal Memory Agent for Long-Horizon Embodied Task Planning
A spatio-temporal memory agent combining a textual history summarizer, a spatial knowledge graph, and a planner-critic loop outperforms ReAct, Reflexion, and AdaPlanner on TextWorld cooking tasks.
-
Inverse Reinforcement Learning with Switching Rewards and History Dependency for Characterizing Animal Behaviors
SWIRL adds state-dependent mode switching and reward functions that depend on recent state history to inverse RL, improving test likelihood on simulated and labyrinth mouse behavior, but not on spontaneous behavior syllables.
-
To Learn or Not to Learn: Analyzing the Role of Learning for Navigation in Virtual Environments
Classical agents outperform learning-based ones on MINOS and Stanford 3D Indoor Spaces, with learned agents weaker at collision avoidance and memory but stronger at handling ambiguity and noise.
-
Mobile Robots through Task-Based Human Instructions using Incremental Curriculum Learning
A simulated mobile robot learns multi-step household instructions better when training is staged from short sub-goals to full instructions, but the supporting experiments lack quantitative comparison.
Discussion (0). Continue with ORCID to comment.