An LLM classifies game states into situations and selects the RL agent with the best historical average reward for each situation, outperforming static ensemble baselines on Atari.
Hierarchical reinforcement learning for scarce medical resource allocation with imperfect information
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One
An LLM classifies game states into situations and selects the RL agent with the best historical average reward for each situation, outperforming static ensemble baselines on Atari.