Pith. sign in

REVIEW 9 cited by

Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.20109 v2 pith:IO22FBZE submitted 2024-07-29 cs.LG cs.AI

classification cs.LGcs.AI
keywords diffusion-dicedistributionoptimalpolicymethodsactionsdiffusionfunction
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

One important property of DIstribution Correction Estimation (DICE) methods is that the solution is the optimal stationary distribution ratio between the optimized and data collection policy. In this work, we show that DICE-based methods can be viewed as a transformation from the behavior distribution to the optimal policy distribution. Based on this, we propose a novel approach, Diffusion-DICE, that directly performs this transformation using diffusion models. We find that the optimal policy's score function can be decomposed into two terms: the behavior policy's score function and the gradient of a guidance term which depends on the optimal distribution ratio. The first term can be obtained from a diffusion model trained on the dataset and we propose an in-sample learning objective to learn the second term. Due to the multi-modality contained in the optimal policy distribution, the transformation in Diffusion-DICE may guide towards those local-optimal modes. We thus generate a few candidate actions and carefully select from them to approach global-optimum. Different from all other diffusion-based offline RL methods, the guide-then-select paradigm in Diffusion-DICE only uses in-sample actions for training and brings minimal error exploitation in the value function. We use a didatic toycase example to show how previous diffusion-based methods fail to generate optimal actions due to leveraging these errors and how Diffusion-DICE successfully avoids that. We then conduct extensive experiments on benchmark datasets to show the strong performance of Diffusion-DICE. Project page at https://ryanxhr.github.io/Diffusion-DICE/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Bilinear contrastive critics remain good compatibility rankers but are unsafe to maximize for action selection; cosine bounding does not fix value decalibration, while Bellman TD-Q does.

  2. Semi-gradient DICE for Offline Constrained Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Semi-gradient DICE outputs a policy correction instead of a stationary distribution correction, and CORSDICE recovers the latter to enable accurate cost estimation and safe offline constrained RL.

  3. Decision Flow Policy Optimization

    cs.LG 2025-05 reject novelty 6.0 of 10

    Decision Flow frames the gradual action generation of flow-based policies as a flow MDP and updates the flow policy with flow-level value functions, reporting state-of-the-art results on several D4RL tasks.

  4. Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Guidance-first diffusion training, which trains and freezes the value guidance before the policy, improves sample efficiency and enables cross-algorithm reuse of guidance modules in offline RL.

  5. Offline Multi-agent Reinforcement Learning via Sequential Score Decomposition

    cs.LG 2025-05 conditional novelty 6.0 of 10

    OMSD uses diffusion models to estimate per-agent conditional score functions of the joint behavior policy, replacing the standard product-factorization assumption and improving offline cooperative MARL performance on ...

  6. Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    AEPO approximates the log-expectation in energy-guided diffusion policy sampling using Taylor expansion and the Gaussian moment-generating function, and reports state-of-the-art average scores on D4RL offline RL benchmarks.

  7. Flow-Based Policy for Online Reinforcement Learning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    FlowRL learns online RL policies as flow-matching models regularized by a Wasserstein-2 constraint toward behavior-optimal replay-buffer actions.

  8. Unpacking the Individual Components of Diffusion Policy

    cs.LG 2024-11 conditional novelty 4.0 of 10

    An ablation study shows that observation sequences, action sequences, receding horizon control, U-Net backbones, and FiLM conditioning each help Diffusion Policy in task-dependent ways, with absolute-control and hard ...

  9. Spatial-Temporal Aware Visuomotor Diffusion Policy Learning

    cs.RO 2025-07 conditional novelty 3.0 of 10

    A diffusion-based visuomotor policy gains 3D and 4D scene awareness from a dynamic Gaussian world model, improving simulated and real robot manipulation success rates.

Pith tools