Pith. sign in

REVIEW 3 cited by

Empowering Large Language Models on Robotic Manipulation with Affordance Prompting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.11027 v1 pith:AMV43JZ7 submitted 2024-04-17 cs.AI

classification cs.AI
keywords controlplanstasksaffordancelanguagellmsmanipulationphysical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While large language models (LLMs) are successful in completing various language processing tasks, they easily fail to interact with the physical world by generating control sequences properly. We find that the main reason is that LLMs are not grounded in the physical world. Existing LLM-based approaches circumvent this problem by relying on additional pre-defined skills or pre-trained sub-policies, making it hard to adapt to new tasks. In contrast, we aim to address this problem and explore the possibility to prompt pre-trained LLMs to accomplish a series of robotic manipulation tasks in a training-free paradigm. Accordingly, we propose a framework called LLM+A(ffordance) where the LLM serves as both the sub-task planner (that generates high-level plans) and the motion controller (that generates low-level control sequences). To ground these plans and control sequences on the physical world, we develop the affordance prompting technique that stimulates the LLM to 1) predict the consequences of generated plans and 2) generate affordance values for relevant objects. Empirically, we evaluate the effectiveness of LLM+A in various language-conditioned robotic manipulation tasks, which show that our approach substantially improves performance by enhancing the feasibility of generated plans and control and can easily generalize to different environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DipLLM: Fine-Tuning LLM for Strategic Decision-making in Diplomacy

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A fine-tuned LLM with a unit-by-unit action decomposition outperforms prior Diplomacy agents while using far less training data.

  2. Conversational Interfaces for Parametric Conceptual Architectural Design: Integrating Mixed Reality with LLM-driven Interaction

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A conversational MR interface with three LLM agents converts speech and gestures into parametric modeling code; a 27-person user study reports lower barriers and successful compilation, though quantitative evidence is...

  3. RoboNav-Arm: Agentic AI-Driven Navigation and Obstacle Avoidance for Robotic Manipulator in Cluttered Environments

    cs.RO 2026-06 conditional novelty 5.0 of 10

    An LLM-orchestrated stack of open-vocabulary 3D scene understanding, memory retrieval, and adaptive RRT-family planning reaches 83–100% collision-free success for a 7-DOF arm in Gazebo clutter scenarios.

Pith tools