Pith. sign in

REVIEW 2 cited by

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.13368 v1 pith:H26CUI4I submitted 2025-04-17 cs.LG cs.AI

classification cs.LGcs.AI
keywords distributionidrloptimalvisitationdatasetdatasetsdiscriminatordual-rl
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Iterative Dual Reinforcement Learning (IDRL), a new method that takes an optimal discriminator-weighted imitation view of solving RL. Our method is motivated by a simple experiment in which we find training a discriminator using the offline dataset plus an additional expert dataset and then performing discriminator-weighted behavior cloning gives strong results on various types of datasets. That optimal discriminator weight is quite similar to the learned visitation distribution ratio in Dual-RL, however, we find that current Dual-RL methods do not correctly estimate that ratio. In IDRL, we propose a correction method to iteratively approach the optimal visitation distribution ratio in the offline dataset given no addtional expert dataset. During each iteration, IDRL removes zero-weight suboptimal transitions using the learned ratio from the previous iteration and runs Dual-RL on the remaining subdataset. This can be seen as replacing the behavior visitation distribution with the optimized visitation distribution from the previous iteration, which theoretically gives a curriculum of improved visitation distribution ratios that are closer to the optimal discriminator weight. We verify the effectiveness of IDRL on various kinds of offline datasets, including D4RL datasets and more realistic corrupted demonstrations. IDRL beats strong Primal-RL and Dual-RL baselines in terms of both performance and stability, on all datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semi-gradient DICE for Offline Constrained Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Semi-gradient DICE outputs a policy correction instead of a stationary distribution correction, and CORSDICE recovers the latter to enable accurate cost estimation and safe offline constrained RL.

  2. Dichotomous Diffusion Policy Optimization

    cs.LG 2025-12 conditional novelty 5.0 of 10

    DIPOLE decomposes a KL-regularized RL objective into a pair of sigmoid-weighted diffusion policies whose score combination (CFG-like) yields stable and controllable policy improvement.

Pith tools