Gated-BEPO derives step-level Bellman advantages from empirical rollout graphs and uses a confidence gate to mix them with episode-level credit, improving LLM agent success on WebShop, ALFWorld, and Sokoban.
Proceedings of the 36th International Conference on Neural Information Processing Systems , articleno =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents
Gated-BEPO derives step-level Bellman advantages from empirical rollout graphs and uses a confidence gate to mix them with episode-level credit, improving LLM agent success on WebShop, ALFWorld, and Sokoban.