Pith. sign in

REVIEW 4 cited by

Affordance-Guided Reinforcement Learning via Visual Prompting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.10341 v6 pith:MHNC2GI6 submitted 2024-07-14 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords autonomouskagilearningmanipulationrewardtasksdemonstrationsdense
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Robots equipped with reinforcement learning (RL) have the potential to learn a wide range of skills solely from a reward signal. However, obtaining a robust and dense reward signal for general manipulation tasks remains a challenge. Existing learning-based approaches require significant data, such as human demonstrations of success and failure, to learn task-specific reward functions. Recently, there is also a growing adoption of large multi-modal foundation models for robotics that can perform visual reasoning in physical contexts and generate coarse robot motions for manipulation tasks. Motivated by this range of capability, in this work, we present Keypoint-based Affordance Guidance for Improvements (KAGI), a method leveraging rewards shaped by vision-language models (VLMs) for autonomous RL. State-of-the-art VLMs have demonstrated impressive zero-shot reasoning about affordances through keypoints, and we use these to define dense rewards that guide autonomous robotic learning. On diverse real-world manipulation tasks specified by natural language descriptions, KAGI improves the sample efficiency of autonomous RL and enables successful task completion in 30K online fine-tuning steps. Additionally, we demonstrate the robustness of KAGI to reductions in the number of in-domain demonstrations used for pre-training, reaching similar performance in 45K online fine-tuning steps. Project website: https://sites.google.com/view/affordance-guided-rl

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A one-shot training regime with DINOv2-enriched point clouds and joint cross-attention predicts 3D object-to-object affordance maps that guide optimization-based robotic manipulation.

  2. DEMONSTRATE: Zero-shot Language to Robotic Control via Multi-task Demonstration Learning

    cs.RO 2025-07 conditional novelty 6.0 of 10

    DEMONSTRATE learns a zero-shot mapping from natural-language embeddings to MPC cost parameters from demonstrations, achieving tabletop manipulation success rates comparable to prior LLM-based pipelines.

  3. VLM-TDP: VLM-guided Trajectory-conditioned Diffusion Policy for Robust Long-Horizon Manipulation

    cs.RO 2025-07 conditional novelty 6.0 of 10

    VLM-TDP guides a diffusion-based robot policy with VLM-generated voxel trajectories, improving success rates by roughly 30-44% and adding robustness to noise and scene changes.

  4. DORA: Object Affordance-Guided Reinforcement Learning for Dexterous Robotic Manipulation

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Using object affordance maps as priors and constraints improves success rates of dexterous manipulation RL policies by an average of 15.4% in simulation.

Pith tools