Pith. sign in

REVIEW 8 cited by

Investigating Compounding Prediction Errors in Learned Dynamics Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.09637 v1 pith:BHNLPCNA submitted 2022-03-17 cs.LG

classification cs.LG
keywords predictionerrorcompoundingcontroldynamicsmbrlproblemactions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Accurately predicting the consequences of agents' actions is a key prerequisite for planning in robotic control. Model-based reinforcement learning (MBRL) is one paradigm which relies on the iterative learning and prediction of state-action transitions to solve a task. Deep MBRL has become a popular candidate, using a neural network to learn a dynamics model that predicts with each pass from high-dimensional states to actions. These "one-step" predictions are known to become inaccurate over longer horizons of composed prediction - called the compounding error problem. Given the prevalence of the compounding error problem in MBRL and related fields of data-driven control, we set out to understand the properties of and conditions causing these long-horizon errors. In this paper, we explore the effects of subcomponents of a control problem on long term prediction error: including choosing a system, collecting data, and training a model. These detailed quantitative studies on simulated and real-world data show that the underlying dynamics of a system are the strongest factor determining the shape and magnitude of prediction error. Given a clearer understanding of compounding prediction error, researchers can implement new types of models beyond "one-step" that are more useful for control.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Time-Aware World Model for Adaptive Prediction and Control

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A time-conditioned world model trained on mixed time steps matches or beats a fixed-time-step baseline at its native rate and far exceeds it at slower observation rates, using the same sample count.

  2. Echoes of Discord: Forecasting Hater Reactions to Counterspeech

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A three-way classifier trained on Reddit hate speech/counterspeech pairs predicts hater reentry and reentry type more accurately than a two-stage predictor, with linguistic features of counterspeech signaling differen...

  3. GenPlan: Generative Sequence Models as Adaptive Planners

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A discrete-flow sequence model with energy and entropy guidance enables adaptive planning that beats LEAP and Decision Transformer in BabyAI adaptive benchmarks.

  4. A Definition and Roadmap for World Models

    cs.AI 2026-07 conditional novelty 5.0 of 10

    A perspective article defining world models as finite-resource compression of physical state transitions and outlining a roadmap toward physical AGI via unified representations and interactive simulators.

  5. Coupled Local and Global World Models for Efficient First Order RL

    cs.RO 2026-02 conditional novelty 5.0 of 10

    Coupled local/global world models let first-order RL train image-space robot policies inside a learned diffusion simulator, outperforming PPO and a DreamerV3-only ablation on two tasks.

  6. First Order Model-Based RL through Decoupled Backpropagation

    cs.RO 2025-08 conditional novelty 5.0 of 10

    By computing gradients through a learned dynamics model while unrolling trajectories in the real simulator, DMO achieves SHAC-level sample efficiency with standard simulators and deploys on a real quadruped.

  7. Improving Transformer World Models for Data-Efficient RL

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A transformer world model agent using a static patch tokenizer, warmup before imagination training, and block teacher forcing reaches 69.66% reward on Craftax-classic, beating DreamerV3 and the human expert figure.

  8. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools