Pith. sign in

REVIEW 11 cited by

Neural Map: Structured Memory for Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1702.08360 v1 pith:4H5XE4S3 submitted 2017-02-27 cs.LG

classification cs.LG
keywords memoryenvironmentsarchitecturesframesneuralagentsdeepinformation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A critical component to enabling intelligent reasoning in partially observable environments is memory. Despite this importance, Deep Reinforcement Learning (DRL) agents have so far used relatively simple memory architectures, with the main methods to overcome partial observability being either a temporal convolution over the past k frames or an LSTM layer. More recent work (Oh et al., 2016) has went beyond these architectures by using memory networks which can allow more sophisticated addressing schemes over the past k frames. But even these architectures are unsatisfactory due to the reason that they are limited to only remembering information from the last k frames. In this paper, we develop a memory system with an adaptable write operator that is customized to the sorts of 3D environments that DRL agents typically interact with. This architecture, called the Neural Map, uses a spatially structured 2D memory image to learn to store arbitrary information about the environment over long time lags. We demonstrate empirically that the Neural Map surpasses previous DRL memories on a set of challenging 2D and 3D maze environments and show that it is capable of generalizing to environments that were not seen during training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    CoCoSI is a training-free multi-agent system for collaborative cognitive map construction that improves spatial understanding in arbitrary pretrained MLLMs.

  2. The Sword, Shield, and Achilles' Heel: Characterizing the Linguistic Inductive Bias of Large Language Models for Spatial Reasoning in Navigation Planning

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    Experiments reveal that topological cues robustly support LLM navigation planning while incorrect semantic cues derail it, with linguistic format effects varying by model size and compression.

  3. Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Tensor Memory augments Transformers with a constant-size 3D voxel grid using differentiable soft writes at predicted locations, local interaction, and gated recurrent dynamics to decouple memory capacity from sequence length.

  4. BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning

    cs.RO 2026-03 unverdicted novelty 6.0 of 10

    BrainMem equips LLM-based embodied planners with working, episodic, and semantic memory that evolves interaction histories into retrievable knowledge graphs and guidelines, raising success rates on long-horizon 3D benchmarks.

  5. Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Flow equivariant world models use a latent memory that shifts with the agent and with inferred object motion, giving stable long-horizon prediction under partial observability.

  6. Situational Fusion of Visual Representation for Visual Navigation

    cs.CV 2019-08 conditional novelty 6.0 of 10

    Combining 25 frozen visual representations at the action level, with a task-affinity regularizer, roughly doubles success rate on unseen indoor navigation scenes compared with an ImageNet-pretrained ResNet baseline.

  7. Cooperative Multi-UAV Navigation in Complex Environments via Systematic Multi-Agent Deep Reinforcement Learning

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A modular MARL framework with local-optima override, hierarchical self-imitation cloning, dual-condition curriculum, and a structure-aware mixture-of-experts achieves 75% cooperative success on an unseen mixed maze in...

  8. STMA: A Spatio-Temporal Memory Agent for Long-Horizon Embodied Task Planning

    cs.AI 2025-02 conditional novelty 5.0 of 10

    A spatio-temporal memory agent combining a textual history summarizer, a spatial knowledge graph, and a planner-critic loop outperforms ReAct, Reflexion, and AdaPlanner on TextWorld cooking tasks.

  9. Inverse Reinforcement Learning with Switching Rewards and History Dependency for Characterizing Animal Behaviors

    cs.LG 2025-01 conditional novelty 5.0 of 10

    SWIRL adds state-dependent mode switching and reward functions that depend on recent state history to inverse RL, improving test likelihood on simulated and labyrinth mouse behavior, but not on spontaneous behavior syllables.

  10. To Learn or Not to Learn: Analyzing the Role of Learning for Navigation in Virtual Environments

    cs.CV 2019-07 unverdicted novelty 4.0 of 10

    Classical agents outperform learning-based ones on MINOS and Stanford 3D Indoor Spaces, with learned agents weaker at collision avoidance and memory but stronger at handling ambiguity and noise.

  11. Mobile Robots through Task-Based Human Instructions using Incremental Curriculum Learning

    cs.RO 2024-12 conditional novelty 3.0 of 10

    A simulated mobile robot learns multi-step household instructions better when training is staged from short sub-goals to full instructions, but the supporting experiments lack quantitative comparison.

Pith tools