Pith. sign in

REVIEW 2 cited by

Practice Makes Perfect: Planning to Learn Skill Parameter Policies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15025 v2 pith:UXXFMW7Z submitted 2024-02-22 cs.RO cs.LG

classification cs.ROcs.LG
keywords robotskillapproachpracticeskillscompetenceimproveparameter
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

One promising approach towards effective robot decision making in complex, long-horizon tasks is to sequence together parameterized skills. We consider a setting where a robot is initially equipped with (1) a library of parameterized skills, (2) an AI planner for sequencing together the skills given a goal, and (3) a very general prior distribution for selecting skill parameters. Once deployed, the robot should rapidly and autonomously learn to improve its performance by specializing its skill parameter selection policy to the particular objects, goals, and constraints in its environment. In this work, we focus on the active learning problem of choosing which skills to practice to maximize expected future task success. We propose that the robot should estimate the competence of each skill, extrapolate the competence (asking: "how much would the competence improve through practice?"), and situate the skill in the task distribution through competence-aware planning. This approach is implemented within a fully autonomous system where the robot repeatedly plans, practices, and learns without any environment resets. Through experiments in simulation, we find that our approach learns effective parameter policies more sample-efficiently than several baselines. Experiments in the real-world demonstrate our approach's ability to handle noise from perception and control and improve the robot's ability to solve two long-horizon mobile-manipulation tasks after a few hours of autonomous practice. Project website: http://ees.csail.mit.edu

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MuTRAP: Multi-trigger Trojans Attacking Robot Task Planning Systems

    cs.RO 2025-04 conditional novelty 6.0 of 10

    A two-step soft-prompt backdoor attack called Robo-Troj (listed as MuTRAP on arXiv) makes LLM-based robot planners emit malicious plans when hidden trigger words are present, with near-perfect attack success.

  2. Learning to Navigate in Mazes with Novel Layouts using Abstract Top-down Maps

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A map-conditioned hypermodel plus MuZero-style planning lets an agent navigate novel maze layouts in zero-shot from an abstract top-down map.

Pith tools