REVIEW 9 cited by
Model-based Reinforcement Learning: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Sequential decision making, commonly formalized as Markov Decision Process (MDP) optimization, is a important challenge in artificial intelligence. Two key approaches to this problem are reinforcement learning (RL) and planning. This paper presents a survey of the integration of both fields, better known as model-based reinforcement learning. Model-based RL has two main steps. First, we systematically cover approaches to dynamics model learning, including challenges like dealing with stochasticity, uncertainty, partial observability, and temporal abstraction. Second, we present a systematic categorization of planning-learning integration, including aspects like: where to start planning, what budgets to allocate to planning and real data collection, how to plan, and how to integrate planning in the learning and acting loop. After these two sections, we also discuss implicit model-based RL as an end-to-end alternative for model learning and planning, and we cover the potential benefits of model-based RL. Along the way, the survey also draws connections to several related RL fields, like hierarchical RL and transfer learning. Altogether, the survey presents a broad conceptual overview of the combination of planning and learning for MDP optimization.
Forward citations
Cited by 9 Pith papers
-
CaLiSym: Learning Symplectic Dynamics of Real-World Systems through Structured Canonical Lifts
Lifting non-conservative, actuated, and contact-constrained robot dynamics into an exactly symplectic phase-space map yields state-of-the-art out-of-distribution autoregressive rollout error at low parameter and FLOP cost.
-
PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
A file-based operating-system layer with a session verifier and persistent memory improves embodied-agent task completion on game, simulated, and real-robot platforms without retraining policies.
-
Temporal Basis Function Models for Closed-Loop Neural Stimulation
Temporal basis function models predict the spatiotemporal LFP response to optogenetic stimulation with test-set R2 around 0.46, beating linear state-space and LSTM baselines while training 30 to 100 times faster.
-
LLM-Guided Probabilistic Program Induction for POMDP Model Estimation
LLM-guided probabilistic program induction can learn low-complexity POMDP models from ten demonstrations and outperform tabular learning, behavior cloning, and direct LLM planning in simulated and real robot domains.
-
Improving Transformer World Models for Data-Efficient RL
A transformer world model agent using a static patch tokenizer, warmup before imagination training, and block teacher forcing reaches 69.66% reward on Craftax-classic, beating DreamerV3 and the human expert figure.
-
Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review
A survey that organizes combinations of Bayesian inference and reinforcement learning, rates them on four properties, and raises ten open questions.
-
Model-free Reinforcement Learning for Model-based Control: Towards Safe, Interpretable and Sample-efficient Agents
A perspective paper argues that model predictive control can be used as a learned policy in model-free reinforcement learning and reviews the methods and open problems.
-
Survey on safe robot control via learning
A survey of safe robot control that reviews classical and learning-based approaches, illustrated with a toy oven example.
-
A Review of Cooperative Multi-Agent Deep Reinforcement Learning
A review that categorizes cooperative multi-agent deep RL into independent learners, observable critics, value factorization, consensus, and communication, with errors in the taxonomy and references.
Discussion (0). Continue with ORCID to comment.