REVIEW 5 cited by
Diffusion Reward: Learning Rewards via Conditional Video Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Learning rewards from expert videos offers an affordable and effective solution to specify the intended behaviors for reinforcement learning (RL) tasks. In this work, we propose Diffusion Reward, a novel framework that learns rewards from expert videos via conditional video diffusion models for solving complex visual RL problems. Our key insight is that lower generative diversity is exhibited when conditioning diffusion on expert trajectories. Diffusion Reward is accordingly formalized by the negative of conditional entropy that encourages productive exploration of expert behaviors. We show the efficacy of our method over robotic manipulation tasks in both simulation platforms and the real world with visual input. Moreover, Diffusion Reward can even solve unseen tasks successfully and effectively, largely surpassing baseline methods. Project page and code: https://diffusion-reward.github.io.
Forward citations
Cited by 5 Pith papers
-
AMPLIFY: Actionless Motion Priors for Robot Learning from Videos
A three-stage pipeline that turns keypoint tracks into discrete motion tokens, predicts them from action-free video, and decodes them into actions yields large few-shot and zero-shot policy improvements in robot manipulation.
-
Solving New Tasks by Adapting Internet Video Knowledge
Inverse Probabilistic Adaptation, a score-composition variant that keeps the large video model as the base and consults a small in-domain model, achieves 68.3% average success on MetaWorld policy supervision and stays...
-
MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation
MJ-VIDEO, a 2B mixture-of-experts reward model trained on a new 28-criteria video preference benchmark, predicts human video preferences more accurately than existing judges and improves text-to-video alignment when u...
-
PROGRESSOR: A Perceptually Guided Reward Estimator with Self-Supervised Online Refinement
A video-trained progress-estimation model, refined online with a push-back objective, provides dense rewards that enable goal-conditioned robot learning without manual reward design or action labels.
-
4D Visual Pre-training for Robot Learning
A next-frame point-cloud diffusion pre-training method (FVP) improves DP3 and RDT-1B manipulation success rates on the paper's own tasks.
Discussion (0). Continue with ORCID to comment.