Pith. sign in

REVIEW 2 cited by

A Behavior Regularized Implicit Policy for Offline Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.09673 v2 pith:M2NIBJ2T submitted 2022-02-19 stat.ML cs.LG

classification stat.MLcs.LG
keywords learningdatasetframeworkpolicytrainingcorrectnesseffectivefurther
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Offline reinforcement learning enables learning from a fixed dataset, without further interactions with the environment. The lack of environmental interactions makes the policy training vulnerable to state-action pairs far from the training dataset and prone to missing rewarding actions. For training more effective agents, we propose a framework that supports learning a flexible yet well-regularized fully-implicit policy. We further propose a simple modification to the classical policy-matching methods for regularizing with respect to the dual form of the Jensen--Shannon divergence and the integral probability metrics. We theoretically show the correctness of the policy-matching approach, and the correctness and a good finite-sample property of our modification. An effective instantiation of our framework through the GAN structure is provided, together with techniques to explicitly smooth the state-action mapping for robust generalization beyond the static dataset. Extensive experiments and ablation study on the D4RL benchmark validate our framework and the effectiveness of our algorithmic designs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Drift Q-Learning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    DriftQL is a single-pass offline RL algorithm using drift regularization that outperforms diffusion and flow policies on standard benchmarks.

  2. Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A wavelet-Fourier conditioning scheme for trajectory diffusion improves offline RL returns on most D4RL tasks by modeling low- and high-frequency components separately.

Pith tools