Pith. sign in

REVIEW 4 cited by

PWM: Policy Learning with Multi-Task World Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02466 v3 pith:MQ2C7ZZ4 submitted 2024-07-02 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords worldlearningmethodsoptimizationmodelsmulti-taskpolicydynamics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement Learning (RL) has made significant strides in complex tasks but struggles in multi-task settings with different embodiments. World model methods offer scalability by learning a simulation of the environment but often rely on inefficient gradient-free optimization methods for policy extraction. In contrast, gradient-based methods exhibit lower variance but fail to handle discontinuities. Our work reveals that well-regularized world models can generate smoother optimization landscapes than the actual dynamics, facilitating more effective first-order optimization. We introduce Policy learning with multi-task World Models (PWM), a novel model-based RL algorithm for continuous control. Initially, the world model is pre-trained on offline data, and then policies are extracted from it using first-order optimization in less than 10 minutes per task. PWM effectively solves tasks with up to 152 action dimensions and outperforms methods that use ground-truth dynamics. Additionally, PWM scales to an 80-task setting, achieving up to 27% higher rewards than existing baselines without relying on costly online planning. Visualizations and code are available at https://www.imgeorgiev.com/pwm/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The SIGReg Objective as Variational Free Energy: A Theoretical Active-Inference Account of JEPA World Models

    cs.LG 2026-07 conditional novelty 7.0 of 10 partial

    Under constant noise and successful SIGReg enforcement, the JEPA objective is an exact variational free-energy/information bottleneck, while VICReg leaves an irreducible anisotropic gap.

  2. Coupled Local and Global World Models for Efficient First Order RL

    cs.RO 2026-02 conditional novelty 5.0 of 10

    Coupled local/global world models let first-order RL train image-space robot policies inside a learned diffusion simulator, outperforming PPO and a DreamerV3-only ablation on two tasks.

  3. First Order Model-Based RL through Decoupled Backpropagation

    cs.RO 2025-08 conditional novelty 5.0 of 10

    By computing gradients through a learned dynamics model while unrolling trajectories in the real simulator, DMO achieves SHAC-level sample efficiency with standard simulators and deploys on a real quadruped.

  4. TD-MPC-Opt: Distilling Model-Based Multi-Task Reinforcement Learning Agents

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Reward-level distillation and FP16 quantization compress a 317M-parameter TD-MPC2 agent to 1M parameters, reaching 28.45 normalized score on MT30, though most of the gap over the original 18.93 comes from a longer tra...

Pith tools