An RL-trained meta-policy that uses ensemble uncertainty to choose between a cheap reactive policy and costly planning reaches goals faster than fixed baselines and adapts as the reactive policy improves.
Title resolution pending
1 Pith paper cite this work, alongside 50 external citations. Polarity classification is still indexing.
1
Pith paper citing it
50
external citations · OpenAlex
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
When to Plan: Learning to Select Between Reactive Control and Deliberative Planning
An RL-trained meta-policy that uses ensemble uncertainty to choose between a cheap reactive policy and costly planning reaches goals faster than fixed baselines and adapts as the reactive policy improves.