Randomly explored environment states, used as exact-match state-reaching RL objectives, improve LLM agent performance on ALFWorld, ScienceWorld, and a mobile GUI benchmark.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
State2State: Environment-Derived Mid-Training for LLM Agents
Randomly explored environment states, used as exact-match state-reaching RL objectives, improve LLM agent performance on ALFWorld, ScienceWorld, and a mobile GUI benchmark.