Pith. sign in

REVIEW 1 cited by

A Unified View on Solving Objective Mismatch in Model-Based Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.06253 v2 pith:U7LMDM5B submitted 2023-10-10 cs.LG

classification cs.LG
keywords modellearningmbrlmismatchobjectiveresearchaccurateagents
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model-based Reinforcement Learning (MBRL) aims to make agents more sample-efficient, adaptive, and explainable by learning an explicit model of the environment. While the capabilities of MBRL agents have significantly improved in recent years, how to best learn the model is still an unresolved question. The majority of MBRL algorithms aim at training the model to make accurate predictions about the environment and subsequently using the model to determine the most rewarding actions. However, recent research has shown that model predictive accuracy is often not correlated with action quality, tracing the root cause to the objective mismatch between accurate dynamics model learning and policy optimization of rewards. A number of interrelated solution categories to the objective mismatch problem have emerged as MBRL continues to mature as a research area. In this work, we provide an in-depth survey of these solution categories and propose a taxonomy to foster future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Safe Planning and Policy Optimization via World Model Learning

    cs.AI 2025-06 conditional novelty 6.0 of 10

    SPOWL is a model-based safe RL method that uses a value-equivalent world model, a Lagrangian-trained safe policy, and adaptive planning thresholds to achieve low-cost, high-reward control on SafetyGymnasium tasks.

Pith tools