Toxic context can be laundered into memory summaries that stay below toxicity thresholds while still driving higher downstream toxicity in LLM agents compared to neutral baselines.
Rap: Retrieval-augmented planning with contextual memory for multimodal llm agents.arXiv preprint arXiv:2402.03610
9 Pith papers cite this work, alongside 5 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
ScreenSearch combines structural screen retrieval and deduplication with an ambiguity-aware PUCT graph-bandit to collect over 1M screenshots and 30K deduplicated states across 11 desktop applications, showing a novelty-ambiguity trade-off in exploration policies.
KGLAMP uses a dynamically updated knowledge graph to guide LLMs in creating and replanning PDDL specifications for heterogeneous multi-robot teams, reporting at least 25.3% better performance than LLM-only or classical PDDL baselines on the MAT-THOR benchmark.
By reusing validated function-level KV caches (stitching) and regenerating only localized error spans (patching), FCGraft makes CodeLLM policies for embodied agents faster and more robust than prompt-level caching.
WISE augments Minecraft agents with causal memory graphs and opportunistic scheduling to raise success rates on long-horizon sparse-reward tasks.
CLASP combines TP-KMPs with VLMs for language-guided skill selection, covariance-weighted composition, and active learning requests, reporting 73.3-100% success on a 7-DoF manipulator.
Inference-time distillation combines dynamic in-context learning from teacher demonstrations with self-consistency cascades to cut LLM agent costs 2.5-3.5x while recovering most accuracy, without training or manual prompts.
Modified OCL search integrates generative rollouts and learned heuristics for efficient inference in planning models across combinatorial domains.
A dual-LLM hierarchical framework for robotic task and motion planning, integrating object detection, achieves 86% success across 24 test scenarios ranging from simple spatial commands to infeasible requests.
citing papers explorer
-
State Contamination in Memory-Augmented LLM Agents
Toxic context can be laundered into memory summaries that stay below toxicity thresholds while still driving higher downstream toxicity in LLM agents compared to neutral baselines.
-
ScreenSearch: Uncertainty-Aware OS Exploration
ScreenSearch combines structural screen retrieval and deduplication with an ambiguity-aware PUCT graph-bandit to collect over 1M screenshots and 30K deduplicated states across 11 desktop applications, showing a novelty-ambiguity trade-off in exploration policies.
-
KGLAMP: Knowledge Graph-guided Language model for Adaptive Multi-robot Planning and Replanning
KGLAMP uses a dynamically updated knowledge graph to guide LLMs in creating and replanning PDDL specifications for heterogeneous multi-robot teams, reporting at least 25.3% better performance than LLM-only or classical PDDL baselines on the MAT-THOR benchmark.
-
Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents
By reusing validated function-level KV caches (stitching) and regenerating only localized error spans (patching), FCGraft makes CodeLLM policies for embodied agents faster and more robust than prompt-level caching.
-
WISE: A Long-Horizon Agent in Minecraft with Why-Which Reasoning
WISE augments Minecraft agents with causal memory graphs and opportunistic scheduling to raise success rates on long-horizon sparse-reward tasks.
-
CLASP: Language-Driven Robot Skill Selection and Composition using Task-Parameterized Learning
CLASP combines TP-KMPs with VLMs for language-guided skill selection, covariance-weighted composition, and active learning requests, reporting 73.3-100% success on a 7-DoF manipulator.
-
Inference-Time Distillation: Cost-Efficient Agents Without Fine-Tuning or Manual Prompt Engineering
Inference-time distillation combines dynamic in-context learning from teacher demonstrations with self-consistency cascades to cut LLM agent costs 2.5-3.5x while recovering most accuracy, without training or manual prompts.
-
Efficient Test-time Inference for Generative Planning Models with OCL Search
Modified OCL search integrates generative rollouts and learned heuristics for efficient inference in planning models across combinatorial domains.
-
Hierarchical Prompting with Dual LLM Modules for Robotic Task and Motion Planning
A dual-LLM hierarchical framework for robotic task and motion planning, integrating object detection, achieves 86% success across 24 test scenarios ranging from simple spatial commands to infeasible requests.