Pith. sign in

REVIEW 1 cited by

Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.04080 v3 pith:EQ46P5IJ submitted 2024-02-06 cs.LG cs.SYeess.SY

classification cs.LGcs.SYeess.SY
keywords policydiffusionofflineq-ensemblesentropy-regularizedlearningreinforcementachieves
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents advanced techniques of training diffusion policies for offline reinforcement learning (RL). At the core is a mean-reverting stochastic differential equation (SDE) that transfers a complex action distribution into a standard Gaussian and then samples actions conditioned on the environment state with a corresponding reverse-time SDE, like a typical diffusion policy. We show that such an SDE has a solution that we can use to calculate the log probability of the policy, yielding an entropy regularizer that improves the exploration of offline datasets. To mitigate the impact of inaccurate value functions from out-of-distribution data points, we further propose to learn the lower confidence bound of Q-ensembles for more robust policy improvement. By combining the entropy-regularized diffusion policy with Q-ensembles in offline RL, our method achieves state-of-the-art performance on most tasks in D4RL benchmarks. Code is available at https://github.com/ruoqizzz/Entropy-Regularized-Diffusion-Policy-with-QEnsemble.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning

    cs.LG 2025-02 conditional novelty 6.0 of 10

    BDPO computes the behavior-regularization penalty for diffusion policies as a sum of per-denoising-step KL divergences and optimizes with a two-time-scale actor-critic, achieving strong D4RL performance.

Pith tools