Pith. sign in

REVIEW 4 cited by

Beating Atari with Natural Language Guided Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1704.05539 v1 pith:R5BJS6ZF submitted 2017-04-18 cs.AI

classification cs.AI
keywords agentatariinstructionslanguagenaturalagentsdeepenvironment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce the first deep reinforcement learning agent that learns to beat Atari games with the aid of natural language instructions. The agent uses a multimodal embedding between environment observations and natural language to self-monitor progress through a list of English instructions, granting itself reward for completing instructions in addition to increasing the game score. Our agent significantly outperforms Deep Q-Networks (DQNs), Asynchronous Advantage Actor-Critic (A3C) agents, and the best agents posted to OpenAI Gym on what is often considered the hardest Atari 2600 environment: Montezuma's Revenge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 35 citations worldwide. Full citation record

  1. Reinforcement Learning with Physics-Informed Symbolic Program Priors for Zero-Shot Wireless Indoor Navigation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Physics priors written as symbolic programs constrain a PPO agent's action choices, yielding better zero-shot wireless indoor navigation and 26%+ training-time savings on Gibson maps.

  2. From Text to Trajectory: Exploring Complex Constraint Representation and Decomposition in Safe Reinforcement Learning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A text-trajectory contrastive model with per-step cost assignment reduces safety violations in reinforcement learning agents under natural language constraints.

  3. Effective Reward Specification in Deep Reinforcement Learning

    cs.LG 2024-12 conditional novelty 4.0 of 10

    A thesis presenting four methods (ASAF, TeamReg, CoachReg, constrained RL, goal-conditioned GFlowNets) that improve reward specification for deep RL through demonstrations, policy regularization, behavior constraints,...

  4. A Survey On Enhancing Reinforcement Learning in Complex Environments: Insights from Human and LLM Feedback

    cs.LG 2024-11 conditional novelty 2.0 of 10

    A survey of prior work on using human and LLM feedback to improve reinforcement learning, plus attention-based methods for large state spaces.

Pith tools