Pith. sign in

REVIEW 9 cited by

When to Trust Your Model: Model-Based Policy Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.08253 v3 pith:F65TGKWH submitted 2019-06-19 cs.LG cs.AIstat.ML

When to Trust Your Model: Model-Based Policy Optimization

classification cs.LG cs.AIstat.ML
keywords model-baseddatamodelalgorithmsanalysismodel-generatedlearningmethods
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Designing effective model-based reinforcement learning algorithms is difficult because the ease of data generation must be weighed against the bias of model-generated data. In this paper, we study the role of model usage in policy optimization both theoretically and empirically. We first formulate and analyze a model-based reinforcement learning algorithm with a guarantee of monotonic improvement at each step. In practice, this analysis is overly pessimistic and suggests that real off-policy data is always preferable to model-generated on-policy data, but we show that an empirical estimate of model generalization can be incorporated into such analysis to justify model usage. Motivated by this analysis, we then demonstrate that a simple procedure of using short model-generated rollouts branched from real data has the benefits of more complicated model-based algorithms without the usual pitfalls. In particular, this approach surpasses the sample efficiency of prior model-based methods, matches the asymptotic performance of the best model-free algorithms, and scales to horizons that cause other model-based methods to fail entirely.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Is a Learned Command Adapter Worth It? Closed-Loop Identification and Counterfactual Auditing of Frozen Locomotion Policies

    cs.AI 2026-07 conditional novelty 7.0

    A counterfactual audit separates same-state headroom from recoverable state-allocation gain, returning NO-GO or ABSTAIN for learned command adapters on frozen Go2 and H1 locomotion policies at 1% thresholds.

  2. MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

    cs.AI 2026-05 unverdicted novelty 7.0

    MiraBench defines action-conditioned reliability via three levels (physics adherence, action-following fidelity, optimism bias detection) and applies it to 12 model configurations using a 16,000-judgment human corpus,...

  3. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  4. Hallucination in World Models is Predictable and Preventable

    cs.LG 2026-06 unverdicted novelty 6.0

    Hallucination in world models is a data coverage issue predictable by three signals and preventable through targeted training sampling and online data collection.

  5. CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization

    cs.RO 2026-06 unverdicted novelty 6.0

    CLAW is an end-to-end self-supervised method that learns semantically meaningful continuous latent actions and predictive world models from action-free videos to support imitation learning and goal-directed planning.

  6. How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position

    cs.LG 2026-06 unverdicted novelty 5.0

    The paper proposes an L0-L7 evidential ladder for evaluating world models in embodied decision-making, prioritizing interventional action fidelity and policy optimization utility over visual plausibility.

  7. Can Predicted Dynamics Exist in the Physical World?

    cs.RO 2026-05 unverdicted novelty 4.0

    Physical admissibility is defined as a prediction-control interface using kinematic, dynamic, and composed-horizon conditions to reject invalid dynamics proposals, with AUC 0.957 on LeRobot PushT and 87-89% prevention...

  8. A Comprehensive Review of Reinforcement Learning for Autonomous Driving in the CARLA Simulator

    cs.RO 2025-09 conditional novelty 4.0

    A survey of roughly 100 CARLA reinforcement learning papers, mapping algorithm families, representations, rewards, evaluation metrics, towns, and open challenges.

  9. Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

    cs.AI 2026-05 unverdicted novelty 2.0

    A survey that maps risks along the agent workflow and consolidates metrics and benchmarks for safety, robustness, privacy, and security in agentic AI.