A policy that takes both an LTL instruction and a mapping specification as inputs can satisfy symbols under varied criteria, outperforming context-aware multi-task RL baselines in navigation and inspection simulations.
Guided Policy Search for Parameterized Skills using Adverbs
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We present a method for using adverb phrases to adjust skill parameters via learned adverb-skill groundings. These groundings allow an agent to use adverb feedback provided by a human to directly update a skill policy, in a manner similar to traditional local policy search methods. We show that our method can be used as a drop-in replacement for these policy search methods when dense reward from the environment is not available but human language feedback is. We demonstrate improved sample efficiency over modern policy search methods in two experiments.
citation-role summary
citation-polarity summary
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Reinforcement Learning of Flexible Policies for Symbolic Instructions with Adjustable Mapping Specifications
A policy that takes both an LTL instruction and a mapping specification as inputs can satisfy symbols under varied criteria, outperforming context-aware multi-task RL baselines in navigation and inspection simulations.