Pith. sign in

REVIEW 4 cited by

Counting to Explore and Generalize in Text-based Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.11525 v2 pith:KRVL4OZO submitted 2018-06-29 cs.CL cs.LG

classification cs.CLcs.LG
keywords text-basedgamesagentdifficultygeneralizepoliciesapproacheschain
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose a recurrent RL agent with an episodic exploration mechanism that helps discovering good policies in text-based game environments. We show promising results on a set of generated text-based games of varying difficulty where the goal is to collect a coin located at the end of a chain of rooms. In contrast to previous text-based RL approaches, we observe that our agent learns policies that generalize to unseen games of greater difficulty.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large Language Models Think Too Fast To Explore Effectively

    cs.AI 2025-01 conditional novelty 6.0 of 10

    In Little Alchemy 2, most LLMs discover fewer elements than humans and rely on uncertainty rather than empowerment; reasoning models o1 and DeepSeek-R1 explore more effectively.

  2. Interactive Language Learning by Question Answering

    cs.CL 2019-08 conditional novelty 6.0 of 10

    QAit turns question answering into an interactive text-game task, and the paper's baselines show current agents cannot generalize beyond memorized games, while humans can.

  3. Learn How to Cook a New Recipe in a New House: Using Map Familiarization, Curriculum Learning, and Bandit Feedback to Learn Families of Text-Based Adventure Games

    cs.CL 2019-08 conditional novelty 6.0 of 10

    Curriculum learning, room-aware action labels, and LinUCB exploration improve zero-shot performance on TextWorld cooking game families, reaching 72% and 68% of achievable points.

  4. Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A thesis proposal repurposing two prior papers on LM agents for text games, framed as a path to theory-of-mind AI, with no new theory-of-mind evidence.

Pith tools