Pith. sign in

REVIEW 3 cited by

Guiding Reinforcement Learning Exploration Using Natural Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1707.08616 v2 pith:M5YZAVN2 submitted 2017-07-26 cs.AI cs.CLcs.LGstat.ML

classification cs.AIcs.CLcs.LGstat.ML
keywords languagelearningnaturalpolicyshapingtechniqueagentenvironments
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this work we present a technique to use natural language to help reinforcement learning generalize to unseen environments. This technique uses neural machine translation, specifically the use of encoder-decoder networks, to learn associations between natural language behavior descriptions and state-action information. We then use this learned model to guide agent exploration using a modified version of policy shaping to make it more effective at learning in unseen environments. We evaluate this technique using the popular arcade game, Frogger, under ideal and non-ideal conditions. This evaluation shows that our modified policy shaping algorithm improves over a Q-learning agent as well as a baseline version of policy shaping.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Formalizing Learning from Language Feedback with Provable Guarantees

    cs.LG 2025-06 conditional novelty 7.0 of 10

    Introduces a formal framework for learning from language feedback, a transfer eluder dimension complexity measure, and HELiX, a no-regret algorithm whose regret scales with this dimension.

  2. Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework

    cs.LG 2024-11 conditional novelty 6.0 of 10

    A conceptual framework classifies human feedback to RL agents along nine dimensions and seven quality criteria, unifying human-centered, interface-centered, and model-centered design perspectives.

  3. Guiding Reinforcement Learning Using Uncertainty-Aware Large Language Models

    cs.LG 2024-11 reject novelty 3.0 of 10

    Using Monte Carlo Dropout to detect LLM uncertainty and weighting LLM advice by entropy gave a small increase in training reward area in a Minigrid task.

Pith tools