REVIEW 15 cited by
RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents
read the original abstract
Owing to recent advancements, Large Language Models (LLMs) can now be deployed as agents for increasingly complex decision-making applications in areas including robotics, gaming, and API integration. However, reflecting past experiences in current decision-making processes, an innate human behavior, continues to pose significant challenges. Addressing this, we propose Retrieval-Augmented Planning (RAP) framework, designed to dynamically leverage past experiences corresponding to the current situation and context, thereby enhancing agents' planning capabilities. RAP distinguishes itself by being versatile: it excels in both text-only and multimodal environments, making it suitable for a wide range of tasks. Empirical evaluations demonstrate RAP's effectiveness, where it achieves SOTA performance in textual scenarios and notably enhances multimodal LLM agents' performance for embodied tasks. These results highlight RAP's potential in advancing the functionality and applicability of LLM agents in complex, real-world applications.
Forward citations
Cited by 15 Pith papers
-
The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory
Conflicting memory enters early and propagates with weak recovery, producing similar compliance rates across models but larger absolute damage for stronger agents.
-
Your Agent's Memories Are Not Its Own: Forged Reasoning Attacks on LLM Agent Memory and Defenses
FARMA forges and self-amplifies an agent's reasoning history with evasive language to induce unsafe skips; SENTINEL's Reasoning Guard reduces ASR to 0% across tested agents and models.
-
When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents
GhostWriter poisons tool-using personal agents' long-term memory via untrusted emails/calendar invites (~98% injection, ~60% activation); AM-Sentry policies and retrieval screens sharply reduce success while preservin...
-
Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
Modeling agent trajectories as action-centric probabilistic graphs lets a GNN warn LLM agents of likely step-level errors before execution, improving pass ratio ~14.7% across four benchmarks.
-
Object-Centric Environment Modeling for Agentic Tasks
Object-Centric Environment Modeling (OCM) builds an online executable object-and-procedure code model that improves average rank and cuts invalid actions on ScienceWorld, ALFWorld, and PlanCraft.
-
Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents
FCGraft synthesizes code policies for embodied agents by grafting KV caches from a library of validated functions, claiming 18.31% higher success rate and 2.3x faster synthesis than prompt-level caching.
-
State Contamination in Memory-Augmented LLM Agents
Toxic context can be laundered into memory summaries that stay below toxicity thresholds while still driving higher downstream toxicity in LLM agents compared to neutral baselines.
-
ScreenSearch: Uncertainty-Aware OS Exploration
ScreenSearch combines structural screen retrieval and deduplication with an ambiguity-aware PUCT graph-bandit to collect over 1M screenshots and 30K deduplicated states across 11 desktop applications, showing a novelt...
-
KGLAMP: Knowledge Graph-guided Language model for Adaptive Multi-robot Planning and Replanning
KGLAMP uses a dynamically updated knowledge graph to guide LLMs in creating and replanning PDDL specifications for heterogeneous multi-robot teams, reporting at least 25.3% better performance than LLM-only or classica...
-
Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents
By reusing validated function-level KV caches (stitching) and regenerating only localized error spans (patching), FCGraft makes CodeLLM policies for embodied agents faster and more robust than prompt-level caching.
-
WISE: A Long-Horizon Agent in Minecraft with Why-Which Reasoning
WISE augments Minecraft agents with causal memory graphs and opportunistic scheduling to raise success rates on long-horizon sparse-reward tasks.
-
CLASP: Language-Driven Robot Skill Selection and Composition using Task-Parameterized Learning
CLASP combines TP-KMPs with VLMs for language-guided skill selection, covariance-weighted composition, and active learning requests, reporting 73.3-100% success on a 7-DoF manipulator.
-
Inference-Time Distillation: Cost-Efficient Agents Without Fine-Tuning or Manual Prompt Engineering
Inference-time distillation combines dynamic in-context learning from teacher demonstrations with self-consistency cascades to cut LLM agent costs 2.5-3.5x while recovering most accuracy, without training or manual prompts.
-
Efficient Test-time Inference for Generative Planning Models with OCL Search
Modified OCL search integrates generative rollouts and learned heuristics for efficient inference in planning models across combinatorial domains.
-
Hierarchical Prompting with Dual LLM Modules for Robotic Task and Motion Planning
A dual-LLM hierarchical framework for robotic task and motion planning, integrating object detection, achieves 86% success across 24 test scenarios ranging from simple spatial commands to infeasible requests.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.