Pith. sign in

REVIEW 1 cited by

Guided Policy Search for Parameterized Skills using Adverbs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.15799 v1 pith:AEFU63F4 submitted 2021-10-23 cs.AI cs.CLcs.HCcs.RO

classification cs.AIcs.CLcs.HCcs.RO
keywords policysearchmethodsadverbfeedbackgroundingshumanmethod
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present a method for using adverb phrases to adjust skill parameters via learned adverb-skill groundings. These groundings allow an agent to use adverb feedback provided by a human to directly update a skill policy, in a manner similar to traditional local policy search methods. We show that our method can be used as a drop-in replacement for these policy search methods when dense reward from the environment is not available but human language feedback is. We demonstrate improved sample efficiency over modern policy search methods in two experiments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reinforcement Learning of Flexible Policies for Symbolic Instructions with Adjustable Mapping Specifications

    cs.RO 2025-01 conditional novelty 6.0 of 10

    A policy that takes both an LTL instruction and a mapping specification as inputs can satisfy symbols under varied criteria, outperforming context-aware multi-task RL baselines in navigation and inspection simulations.

Pith tools