Agents that accumulate their own successful trajectories as retrieval examples gain up to 20 points on ALFWorld, Wordcraft, and InterCode-SQL, with two curation methods pushing gains further.
WordCraft: An Environment for Benchmarking Commonsense Agents
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The ability to quickly solve a wide range of real-world tasks requires a commonsense understanding of the world. Yet, how to best extract such knowledge from natural language corpora and integrate it with reinforcement learning (RL) agents remains an open challenge. This is partly due to the lack of lightweight simulation environments that sufficiently reflect the semantics of the real world and provide knowledge sources grounded with respect to observations in an RL environment. To better enable research on agents making use of commonsense knowledge, we propose WordCraft, an RL environment based on Little Alchemy 2. This lightweight environment is fast to run and built upon entities and relations inspired by real-world semantics. We evaluate several representation learning methods on this new benchmark and propose a new method for integrating knowledge graphs with an RL agent.
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
Agents that accumulate their own successful trajectories as retrieval examples gain up to 20 points on ALFWorld, Wordcraft, and InterCode-SQL, with two curation methods pushing gains further.