Pith. sign in

REVIEW 2 cited by

DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14790 v1 pith:I7XAUBTU submitted 2024-05-23 cs.LG

DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation

classification cs.LG
keywords dididiverseofflinedatadiffusion-guideddiversitygenerationskill
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In this paper, we propose a novel approach called DIffusion-guided DIversity (DIDI) for offline behavioral generation. The goal of DIDI is to learn a diverse set of skills from a mixture of label-free offline data. We achieve this by leveraging diffusion probabilistic models as priors to guide the learning process and regularize the policy. By optimizing a joint objective that incorporates diversity and diffusion-guided regularization, we encourage the emergence of diverse behaviors while maintaining the similarity to the offline data. Experimental results in four decision-making domains (Push, Kitchen, Humanoid, and D4RL tasks) show that DIDI is effective in discovering diverse and discriminative skills. We also introduce skill stitching and skill interpolation, which highlight the generalist nature of the learned skill space. Further, by incorporating an extrinsic reward function, DIDI enables reward-guided behavior generation, facilitating the learning of diverse and optimal behaviors from sub-optimal data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning

    cs.LG 2025-09 conditional novelty 6.0

    A wavelet-Fourier conditioning scheme for trajectory diffusion improves offline RL returns on most D4RL tasks by modeling low- and high-frequency components separately.

  2. Expert Behavior Prior Reinforcement Learning

    cs.AI 2026-07 conditional novelty 5.0

    An online RL method that learns a generative behavior prior from the replay buffer via a Q-guided CVAE and uses adaptive gradient correction to combine Q-guidance with expert-action supervision.