Pith. sign in

REVIEW 3 cited by

Learning from Random Demonstrations: Offline Reinforcement Learning with Importance-Sampled Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19878 v1 pith:MX3EBN6Q submitted 2024-05-30 cs.LG cs.GT

classification cs.LGcs.GT
keywords learningofflinepolicyworlddiffusionmodelmodelsreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Generative models such as diffusion have been employed as world models in offline reinforcement learning to generate synthetic data for more effective learning. Existing work either generates diffusion models one-time prior to training or requires additional interaction data to update it. In this paper, we propose a novel approach for offline reinforcement learning with closed-loop policy evaluation and world-model adaptation. It iteratively leverages a guided diffusion world model to directly evaluate the offline target policy with actions drawn from it, and then performs an importance-sampled world model update to adaptively align the world model with the updated policy. We analyzed the performance of the proposed method and provided an upper bound on the return gap between our method and the real environment under an optimal policy. The result sheds light on various factors affecting learning performance. Evaluations in the D4RL environment show significant improvement over state-of-the-art baselines, especially when only random or medium-expertise demonstrations are available -- thus requiring improved alignment between the world model and offline policy evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees

    cs.LG 2025-08 unverdicted novelty 5.0 of 10

    Two new Monte Carlo tree search algorithms, CVaR-MCTS and W-MCTS, give provable PAC-level tail-risk controls and regret bounds for worst-case outcome scenarios.

  2. Residual Reward Models for Preference-based Reinforcement Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Combining a hand-designed or learned prior reward with a preference-trained residual improves sample efficiency and final performance in preference-based reinforcement learning.

  3. Perception Graph for Cognitive Attack Reasoning in Augmented Reality

    cs.AI 2025-08 reject novelty 3.0 of 10

    The Perception Graph paper proposes detecting cognitive attacks in AR by measuring cosine distance between vision-language descriptions of scenes, demonstrated on three attacks in one scene.

Pith tools