Pith. sign in

REVIEW 16 cited by

Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.18313 v5 pith:4OL6D7GA submitted 2024-09-26 cs.RO cs.AIcs.LG

Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation

classification cs.RO cs.AIcs.LG
keywords embodied-ragembodiednon-parametricacrossgenerationknowledgelanguagememory
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

There is no limit to how much a robot might explore and learn, but all of that knowledge needs to be searchable and actionable. Within language research, retrieval augmented generation (RAG) has become the workhorse of large-scale non-parametric knowledge; however, existing techniques do not directly transfer to the embodied domain, which is multimodal, where data is highly correlated, and perception requires abstraction. To address these challenges, we introduce Embodied-RAG, a framework that enhances the foundational model of an embodied agent with a non-parametric memory system capable of autonomously constructing hierarchical knowledge for both navigation and language generation. Embodied-RAG handles a full range of spatial and semantic resolutions across diverse environments and query types, whether for a specific object or a holistic description of ambiance. At its core, Embodied-RAG's memory is structured as a semantic forest, storing language descriptions at varying levels of detail. This hierarchical organization allows the system to efficiently generate context-sensitive outputs across different robotic platforms. We demonstrate that Embodied-RAG effectively bridges RAG to the robotics domain, successfully handling over 250 explanation and navigation queries across kilometer-level environments, highlighting its promise as a general-purpose non-parametric system for embodied agents.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Remember with Confidence: Uncertainty Quantification for Spatio-temporal Memory with Probabilistic Guarantees

    cs.CV 2026-06 unverdicted novelty 7.0

    Introduces object-level semantic uncertainty for VLM memory, the UQ-DAAAM refinement system, and probabilistic guarantees that selected high-quality views reduce uncertainty more effectively.

  2. Worth Remembering: Surprise-Gated Robot Episodic Memory

    cs.RO 2026-06 unverdicted novelty 7.0

    Surprise-gated episodic memory using V-JEPA-2 improves robot QA by ≥12% over prior memory methods and outperforms supervised baselines on event segmentation.

  3. ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models

    cs.RO 2026-05 unverdicted novelty 7.0

    ECHO organizes VLA experiences into a hierarchical memory tree in hyperbolic space via autoencoder and entailment constraints, delivering a 12.8% success-rate gain on LIBERO-Long over the pi0 baseline.

  4. A Survey of Spatial Memory Representations for Efficient Robot Navigation

    cs.CV 2026-04 unverdicted novelty 7.0

    The survey reviews spatial memory methods across 88 references, defines α as peak runtime memory over map size, profiles neural methods showing α from 2.3 to 215 on A100 GPU, and proposes a standardized evaluation pro...

  5. NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning

    cs.AI 2026-06 conditional novelty 6.5

    MLLM embodied planners hit explicit goals ~67% of the time but hidden social norms only ~26%; scene-grounded cues, not generic knowledge, close much of the gap.

  6. NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning

    cs.AI 2026-06 unverdicted novelty 6.0

    NormAct shows MLLMs reach explicit goals in 67.3% of cases but comply with hidden norms in only 26.4%, with NormPerceptor raising task success from 24.2% to 46.7%.

  7. SAGE-Nav: Leveraging LLM Planning and Alignment Fusion for Hierarchical Scene Graph-Guided Navigation

    cs.RO 2026-06 unverdicted novelty 6.0

    SAGE-Nav decouples LLM global planning from reactive control via hierarchical scene graphs and alignment fusion, reporting SOTA results on i-THOR and RoboTHOR with improved efficiency and zero-shot generalization.

  8. PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraft

    cs.AI 2026-05 unverdicted novelty 6.0

    PEAM is a parametric memory framework for Minecraft agents that internalizes experiences into a multimodal MoE-LoRA module using contrastive objectives on failures and a scale-free self-triggered consolidation mechanism.

  9. FOUND-IT: Foundation-model-first Task-driven 3D Scene Graphs with Granularity on Demand

    cs.RO 2026-05 unverdicted novelty 6.0

    FOUND-IT constructs evolving task-driven 3D scene graphs with on-demand granularity from monocular cameras by augmenting foundation models, reporting 79% higher accuracy on a grounding benchmark and real-time Jetson d...

  10. Memento: Towards Proactive Visualization of Everyday Memories with Personal Wearable AR Assistant

    cs.HC 2026-01 unverdicted novelty 6.0

    Memento captures verbal queries with spatiotemporal contexts, discovers recurring interest patterns, and proactively delivers updated AR responses when matching situations are detected.

  11. CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

    cs.CV 2026-06 unverdicted novelty 5.0

    CLOSER-VLN combines hierarchical reasoning, multidimensional verification, and triggered retrieval to reach 32.01% SR and 21.28% SPL on CityNav test-unseen without task-specific training.

  12. G-DRAGON: Geospatial Reasoning and Dynamic Planning for Retrieval-Augmented Outdoor Navigation

    cs.RO 2026-05 unverdicted novelty 5.0

    G-DRAGON framework maps language commands to OSM coordinates via lightweight LLM for global planning and uses frontier exploration for local targets, outperforming baselines in simulation and completing real UGV perso...

  13. From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

    cs.CV 2026-05 unverdicted novelty 5.0

    BehaviorVLA introduces a symmetric encoder-decoder architecture with causal Mamba and phase conditioning to learn unified long-horizon behavioral representations for improved generalization in VLA models.

  14. TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation

    cs.CV 2026-05 unverdicted novelty 5.0

    TrajRAG uses a topological-polar trajectory representation and hierarchical retrieval to accumulate and reuse geometric-semantic navigation experiences, improving zero-shot ObjectNav on MP3D and HM3D benchmarks.

  15. From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

    cs.CV 2026-05 unverdicted novelty 4.0

    BehaviorVLA learns long-horizon behavioral representations via causal Mamba encoder and phase-conditioned decoder, reporting SOTA results of 58% on RoboTwin 2.0, 98% on LIBERO, 4.36 on CALVIN, and matching OpenVLA-OFT...

  16. 3D Scene Graphs: Open Challenges and Future Directions

    cs.RO 2026-06 unverdicted novelty 2.0

    A survey that formalizes 3D Scene Graphs under a common definition, analyzes modeling choices, reviews construction from sensory data, examines applications and evaluations, and highlights open challenges with a suppo...