REVIEW 1 cited by
Informed POMDP: Leveraging Additional Information in Model-Based RL
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this work, we generalize the problem of learning through interaction in a POMDP by accounting for eventual additional information available at training time. First, we introduce the informed POMDP, a new learning paradigm offering a clear distinction between the information at training and the observation at execution. Next, we propose an objective that leverages this information for learning a sufficient statistic of the history for the optimal control. We then adapt this informed objective to learn a world model able to sample latent trajectories. Finally, we empirically show a learning speed improvement in several environments using this informed world model in the Dreamer algorithm. These results and the simplicity of the proposed adaptation advocate for a systematic consideration of eventual additional information when learning in a POMDP using model-based RL.
Forward citations
Cited by 1 Pith paper
-
GoStop: Reinforcement Learning for Adaptive Temporal Aggregation in Event-Based Feature Tracking
An RL agent that adaptively decides when to accumulate events and when to run tracking inference improves event-based feature tracking on a new dynamic benchmark, but the gains are less consistent on an existing benchmark.
Discussion (0). Continue with ORCID to comment.