Pith. sign in

REVIEW 3 cited by

Learning Multimodal Behaviors from Scratch with Diffusion Policy Gradient

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.00681 v1 pith:5SLDCMDS submitted 2024-06-02 cs.LG

classification cs.LG
keywords policydiffusionmultimodalddiffpglearningbehaviorsdeepmodes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep reinforcement learning (RL) algorithms typically parameterize the policy as a deep network that outputs either a deterministic action or a stochastic one modeled as a Gaussian distribution, hence restricting learning to a single behavioral mode. Meanwhile, diffusion models emerged as a powerful framework for multimodal learning. However, the use of diffusion policies in online RL is hindered by the intractability of policy likelihood approximation, as well as the greedy objective of RL methods that can easily skew the policy to a single mode. This paper presents Deep Diffusion Policy Gradient (DDiffPG), a novel actor-critic algorithm that learns from scratch multimodal policies parameterized as diffusion models while discovering and maintaining versatile behaviors. DDiffPG explores and discovers multiple modes through off-the-shelf unsupervised clustering combined with novelty-based intrinsic motivation. DDiffPG forms a multimodal training batch and utilizes mode-specific Q-learning to mitigate the inherent greediness of the RL objective, ensuring the improvement of the diffusion policy across all modes. Our approach further allows the policy to be conditioned on mode-specific embeddings to explicitly control the learned modes. Empirical studies validate DDiffPG's capability to master multimodal behaviors in complex, high-dimensional continuous control tasks with sparse rewards, also showcasing proof-of-concept dynamic online replanning when navigating mazes with unseen obstacles.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training Diffusion Policies via Prior-Mapping Co-Evolution

    cs.LG 2025-12 conditional novelty 6.0 of 10

    GoRL outperforms Gaussian and generative RL baselines on continuous control by optimizing a Gaussian latent policy while a separately trained diffusion or flow decoder maps fixed noise to actions.

  2. DIME:Diffusion-Based Maximum Entropy Reinforcement Learning

    cs.LG 2025-02 conditional novelty 6.0 of 10

    DIME derives a variational lower bound on the maximum entropy RL objective for diffusion policies and shows strong continuous-control benchmark results.

  3. Exploratory Diffusion Model for Unsupervised Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A diffusion-model denoising loss serves as an intrinsic reward to guide unsupervised RL exploration, plus an alternating fine-tuning scheme for diffusion policies.

Pith tools