REVIEW 34 cited by
Diffusion Guidance Is a Controllable Policy Improvement Operator
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Diffusion Guidance Is a Controllable Policy Improvement Operator
read the original abstract
At the core of reinforcement learning is the idea of learning beyond the performance in the data. However, scaling such systems has proven notoriously tricky. In contrast, techniques from generative modeling have proven remarkably scalable and are simple to train. In this work, we combine these strengths, by deriving a direct relation between policy improvement and guidance of diffusion models. The resulting framework, CFGRL, is trained with the simplicity of supervised learning, yet can further improve on the policies in the data. On offline RL tasks, we observe a reliable trend -- increased guidance weighting leads to increased performance. Of particular importance, CFGRL can operate without explicitly learning a value function, allowing us to generalize simple supervised methods (e.g., goal-conditioned behavioral cloning) to further prioritize optimality, gaining performance for "free" across the board.
Forward citations
Cited by 34 Pith papers
-
VINE: Taming Generative Control Policies for Reinforcement Learning
Reconstructing a fresh noisy interpolation state at every denoising step stabilizes end-to-end value-gradient training of multi-step flow-matching policies and yields state-of-the-art offline and real-robot results.
-
Improving Robotic Generalist Policies via Flow Reversal Steering
Flow Reversal Steering steers flow matching generalist policies by reversing suboptimal actions to nearby better modes, enabling improved zero-shot control, quick distillation, and RL bootstrapping in robotic manipulation.
-
Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning
QGF performs test-time policy optimization for flow models in RL by guiding a behavior-cloned reference policy with value-function gradients, achieving strong results on high-dimensional offline RL benchmarks without ...
-
JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning
JEDI is the first online end-to-end latent diffusion world model that trains latents from denoising loss rather than reconstruction, achieving competitive Atari100k results with 43% less VRAM and over 3x faster sampli...
-
Reinforcement Learning via Value Gradient Flow
VGF solves behavior-regularized RL by transporting particles from a reference distribution to the value-induced optimal policy via discrete value-guided gradient flow.
-
ViVa: A Video-Generative Value Model for Robot Reinforcement Learning
ViVa turns a video generator into a value model for robot RL that jointly forecasts future states and task value, yielding better performance on real-world box assembly when integrated with RECAP.
-
DiffusionNFT: Online Diffusion Reinforcement with Forward Process
DiffusionNFT performs online RL for diffusion models on the forward process via flow matching and positive-negative contrasts, delivering up to 25x efficiency gains and rapid benchmark improvements over prior reverse-...
-
GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning
GeMPO unifies diffusion RL reweighting as measure matching to a regularized target, enabling flexible and negative weights that improve exploration and performance.
-
CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning
Closed-loop fine-tuning through a supervised-data API alone — no weights, gradients, or losses — lifts a closed-weight humanoid VLA to near-perfect success on three contact-rich tasks after two self-improvement cycles.
-
RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy
RedFlow improves flow-matching VLA policies offline by labeling individual actions from failed rollouts with progress signals and pushing those actions toward successful alternatives found in similar contexts.
-
TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training
A tactile-aware world model recognizes failure-adjacent contact states, imagines local visuo-tactile corrections, and post-trains VLAs with knowledge insulation, raising average success by 44% over the base policy.
-
Controllable Sim Agents with Behavior Latents
CNeVA combines variational behavior latents with rectified-flow generators and soft eligibility to deliver controllable yet realistic traffic simulation on Waymo data.
-
STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning
STEAM learns advantages from expert trajectories via self-supervised temporal ensemble modeling to improve policy learning on real robot tasks like bimanual folding and pick-and-place.
-
ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies
ReGuide is a self-improving framework that uses phase-conditioned guidance to generate corrective rollouts and absorbs successful ones back into diffusion policy training, yielding 1.3-7.7x success gains on Robomimic tasks.
-
FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning
FlowR2A learns reward-conditioned action distributions via flow-matching decoder to unify dense reward supervision with dynamic proposal generation for multimodal driving planning.
-
Reversal Q-Learning
Reversal Q-Learning (RQL) proposes reversing flows for virtual trajectories and bias-variance reduction in an expanded MDP to train flow policies, reporting best average performance on 50 simulated robotic tasks versu...
-
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy
Q-Flow enables stable optimization of expressive flow-based policies in RL by propagating terminal values along deterministic flow dynamics to intermediate states for gradient updates without solver unrolling.
-
Refining Compositional Diffusion for Reliable Long-Horizon Planning
RCD steers compositional diffusion sampling toward high-density coherent plans by combining reconstruction-error guidance with overlap consistency, outperforming prior methods on locomotion, manipulation, and pixel-ba...
-
Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models
Reward-weighted classifier-free guidance approximates Q-function policy improvement in autoregressive models, enabling test-time reward optimization and faster RL convergence via distillation.
-
Value-Guidance MeanFlow for Offline Multi-Agent Reinforcement Learning
VGM²P achieves SOTA-comparable performance in offline MARL via value-guided conditional behavior cloning with MeanFlow, enabling efficient single-step action generation insensitive to regularization coefficients.
-
Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence
Projection-aware activation steering using logistic regression recovers honesty and compassion under malicious prompts while preserving coherence and benchmark performance better than uniform steering.
-
Update-Free On-Policy Steering via Verifiers
Lightweight verifiers trained on a diffusion policy’s own evaluation rollouts raise real-robot success rates ~49% on average via Best-of-N or classifier guidance, without changing base parameters.
-
From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning
Residual off-policy RL with selective BC regularization and value-guided sampling contracts a pretrained generative robot policy around successful actions, reaching high success on hard long-horizon tasks from pixels ...
-
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
Steerable VLAs trained on rich synthetic commands at subtask, motion, and pixel levels enable VLMs to steer robot behavior more effectively, outperforming prior hierarchical baselines on real-world manipulation and ge...
-
RISE: Self-Improving Robot Policy with Compositional World Model
RISE combines a controllable dynamics model and progress value model into a closed-loop self-improving pipeline that updates robot policies entirely in imagination, reporting over 35% absolute gains on three real-world tasks.
-
Latent Reasoning in TRMs is Secretly a Policy Improvement Operator
Recursive reasoning in TRMs is reinterpreted as policy improvement, and a stepwise denoising supervision scheme (DIS) cuts forward passes 18× while improving small-model ARC scores.
-
$\pi^{*}_{0.6}$: a VLA That Learns From Experience
RECAP enables a generalist VLA to self-improve via advantage-conditioned RL on mixed real-world data, more than doubling throughput and halving failure rates on hard manipulation tasks.
-
DexPIE: Stable Dexterous Policy Improvement from Real-World Experience
DexPIE improves dexterous manipulation success rates by 37% over demo policies via real-world experience collection with adapted intervention, multi-stage DAgger, asynchronous relative-action inference, and optimality...
-
Scaling by Diversified Experience for Vision-Language-Action Models
SyVLA uses Intention Decoupling and similar-sample guided RL on diversified experiences to improve VLA model task success and out-of-distribution generalization while keeping vision-language abilities.
-
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy
Q-Flow bridges stability and expressivity in flow-based RL policies by propagating terminal trajectory values to intermediate states for gradient-based optimization.
-
Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence
Projection-aware activation steering recovers alignment under dishonesty and dismissiveness threat models while preserving capabilities better than fixed-coefficient steering, and generalizes to several OOD tests.
-
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
ALOE uses chunked TD bootstrapping with a pessimistic Q-ensemble to enable action-level off-policy value estimation for advantage-weighted post-training of flow-based VLA policies, reporting consistent success-rate ga...
-
Dichotomous Diffusion Policy Optimization
DIPOLE decomposes a KL-regularized RL objective into a pair of sigmoid-weighted diffusion policies whose score combination (CFG-like) yields stable and controllable policy improvement.
-
Robot Self-Improvement via Human-Video Dynamics Models
Human-video dynamics models enable cross-embodiment robot self-improvement via training-free Dynamics-Guided Action Correction, raising success rates from 40% to 81% on seven real-world tasks.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.