Pith. sign in

REVIEW 34 cited by

Diffusion Guidance Is a Controllable Policy Improvement Operator

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.23458 v1 pith:E7YY3HPC submitted 2025-05-29 cs.LG

Diffusion Guidance Is a Controllable Policy Improvement Operator

classification cs.LG
keywords learningguidanceperformancecfgrldatadiffusionfurtherimprovement
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

At the core of reinforcement learning is the idea of learning beyond the performance in the data. However, scaling such systems has proven notoriously tricky. In contrast, techniques from generative modeling have proven remarkably scalable and are simple to train. In this work, we combine these strengths, by deriving a direct relation between policy improvement and guidance of diffusion models. The resulting framework, CFGRL, is trained with the simplicity of supervised learning, yet can further improve on the policies in the data. On offline RL tasks, we observe a reliable trend -- increased guidance weighting leads to increased performance. Of particular importance, CFGRL can operate without explicitly learning a value function, allowing us to generalize simple supervised methods (e.g., goal-conditioned behavioral cloning) to further prioritize optimality, gaining performance for "free" across the board.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 34 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VINE: Taming Generative Control Policies for Reinforcement Learning

    cs.RO 2026-07 conditional novelty 7.0

    Reconstructing a fresh noisy interpolation state at every denoising step stabilizes end-to-end value-gradient training of multi-step flow-matching policies and yields state-of-the-art offline and real-robot results.

  2. Improving Robotic Generalist Policies via Flow Reversal Steering

    cs.RO 2026-06 unverdicted novelty 7.0

    Flow Reversal Steering steers flow matching generalist policies by reversing suboptimal actions to nearby better modes, enabling improved zero-shot control, quick distillation, and RL bootstrapping in robotic manipulation.

  3. Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning

    cs.LG 2026-06 unverdicted novelty 7.0

    QGF performs test-time policy optimization for flow models in RL by guiding a behavior-cloned reference policy with value-function gradients, achieving strong results on high-dimensional offline RL benchmarks without ...

  4. JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning

    cs.LG 2026-05 unverdicted novelty 7.0

    JEDI is the first online end-to-end latent diffusion world model that trains latents from denoising loss rather than reconstruction, achieving competitive Atari100k results with 43% less VRAM and over 3x faster sampli...

  5. Reinforcement Learning via Value Gradient Flow

    cs.LG 2026-04 unverdicted novelty 7.0

    VGF solves behavior-regularized RL by transporting particles from a reference distribution to the value-induced optimal policy via discrete value-guided gradient flow.

  6. ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

    cs.RO 2026-04 unverdicted novelty 7.0

    ViVa turns a video generator into a value model for robot RL that jointly forecasts future states and task value, yielding better performance on real-world box assembly when integrated with RECAP.

  7. DiffusionNFT: Online Diffusion Reinforcement with Forward Process

    cs.LG 2025-09 unverdicted novelty 7.0

    DiffusionNFT performs online RL for diffusion models on the forward process via flow matching and positive-negative contrasts, delivering up to 25x efficiency gains and rapid benchmark improvements over prior reverse-...

  8. GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning

    cs.LG 2026-03 conditional novelty 6.5

    GeMPO unifies diffusion RL reweighting as measure matching to a regularized target, enabling flexible and negative weights that improve exploration and performance.

  9. CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

    cs.RO 2026-07 conditional novelty 6.0

    Closed-loop fine-tuning through a supervised-data API alone — no weights, gradients, or losses — lifts a closed-weight humanoid VLA to near-perfect success on three contact-rich tasks after two self-improvement cycles.

  10. RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy

    cs.RO 2026-07 conditional novelty 6.0

    RedFlow improves flow-matching VLA policies offline by labeling individual actions from failed rollouts with progress signals and pushing those actions toward successful alternatives found in similar contexts.

  11. TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training

    cs.RO 2026-07 conditional novelty 6.0

    A tactile-aware world model recognizes failure-adjacent contact states, imagines local visuo-tactile corrections, and post-trains VLAs with knowledge insulation, raising average success by 44% over the base policy.

  12. Controllable Sim Agents with Behavior Latents

    cs.RO 2026-07 unverdicted novelty 6.0

    CNeVA combines variational behavior latents with rectified-flow generators and soft eligibility to deliver controllable yet realistic traffic simulation on Waymo data.

  13. STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning

    cs.RO 2026-06 unverdicted novelty 6.0

    STEAM learns advantages from expert trajectories via self-supervised temporal ensemble modeling to improve policy learning on real robot tasks like bimanual folding and pick-and-place.

  14. ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies

    cs.LG 2026-06 unverdicted novelty 6.0

    ReGuide is a self-improving framework that uses phase-conditioned guidance to generate corrective rollouts and absorbs successful ones back into diffusion policy training, yielding 1.3-7.7x success gains on Robomimic tasks.

  15. FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning

    cs.AI 2026-06 unverdicted novelty 6.0

    FlowR2A learns reward-conditioned action distributions via flow-matching decoder to unify dense reward supervision with dynamic proposal generation for multimodal driving planning.

  16. Reversal Q-Learning

    cs.LG 2026-06 unverdicted novelty 6.0

    Reversal Q-Learning (RQL) proposes reversing flows for virtual trajectories and bias-variance reduction in an expanded MDP to train flow policies, reporting best average performance on 50 simulated robotic tasks versu...

  17. Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy

    cs.LG 2026-05 unverdicted novelty 6.0

    Q-Flow enables stable optimization of expressive flow-based policies in RL by propagating terminal values along deterministic flow dynamics to intermediate states for gradient updates without solver unrolling.

  18. Refining Compositional Diffusion for Reliable Long-Horizon Planning

    cs.RO 2026-05 unverdicted novelty 6.0

    RCD steers compositional diffusion sampling toward high-density coherent plans by combining reconstruction-error guidance with overlap consistency, outperforming prior methods on locomotion, manipulation, and pixel-ba...

  19. Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models

    cs.LG 2026-04 unverdicted novelty 6.0

    Reward-weighted classifier-free guidance approximates Q-function policy improvement in autoregressive models, enabling test-time reward optimization and faster RL convergence via distillation.

  20. Value-Guidance MeanFlow for Offline Multi-Agent Reinforcement Learning

    cs.LG 2026-04 unverdicted novelty 6.0

    VGM²P achieves SOTA-comparable performance in offline MARL via value-guided conditional behavior cloning with MeanFlow, enabling efficient single-step action generation insensitive to regularization coefficients.

  21. Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence

    cs.AI 2026-04 unverdicted novelty 6.0

    Projection-aware activation steering using logistic regression recovers honesty and compassion under malicious prompts while preserving coherence and benchmark performance better than uniform steering.

  22. Update-Free On-Policy Steering via Verifiers

    cs.RO 2026-03 conditional novelty 6.0

    Lightweight verifiers trained on a diffusion policy’s own evaluation rollouts raise real-robot success rates ~49% on average via Best-of-N or classifier guidance, without changing base parameters.

  23. From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning

    cs.RO 2026-03 accept novelty 6.0

    Residual off-policy RL with selective BC regularization and value-guided sampling contracts a pretrained generative robot policy around successful actions, reaching high success on hard long-horizon tasks from pixels ...

  24. Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control

    cs.RO 2026-02 unverdicted novelty 6.0

    Steerable VLAs trained on rich synthetic commands at subtask, motion, and pixel levels enable VLMs to steer robot behavior more effectively, outperforming prior hierarchical baselines on real-world manipulation and ge...

  25. RISE: Self-Improving Robot Policy with Compositional World Model

    cs.RO 2026-02 unverdicted novelty 6.0

    RISE combines a controllable dynamics model and progress value model into a closed-loop self-improving pipeline that updates robot policies entirely in imagination, reporting over 35% absolute gains on three real-world tasks.

  26. Latent Reasoning in TRMs is Secretly a Policy Improvement Operator

    cs.CL 2025-11 reject novelty 6.0

    Recursive reasoning in TRMs is reinterpreted as policy improvement, and a stepwise denoising supervision scheme (DIS) cuts forward passes 18× while improving small-model ARC scores.

  27. $\pi^{*}_{0.6}$: a VLA That Learns From Experience

    cs.LG 2025-11 unverdicted novelty 6.0

    RECAP enables a generalist VLA to self-improve via advantage-conditioned RL on mixed real-world data, more than doubling throughput and halving failure rates on hard manipulation tasks.

  28. DexPIE: Stable Dexterous Policy Improvement from Real-World Experience

    cs.RO 2026-06 unverdicted novelty 5.0

    DexPIE improves dexterous manipulation success rates by 37% over demo policies via real-world experience collection with adapted intervention, multi-stage DAgger, asynchronous relative-action inference, and optimality...

  29. Scaling by Diversified Experience for Vision-Language-Action Models

    cs.CV 2026-06 unverdicted novelty 5.0

    SyVLA uses Intention Decoupling and similar-sample guided RL on diversified experiences to improve VLA model task success and out-of-distribution generalization while keeping vision-language abilities.

  30. Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy

    cs.LG 2026-05 unverdicted novelty 5.0

    Q-Flow bridges stability and expressivity in flow-based RL policies by propagating terminal trajectory values to intermediate states for gradient-based optimization.

  31. Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence

    cs.AI 2026-04 unverdicted novelty 5.0

    Projection-aware activation steering recovers alignment under dishonesty and dismissiveness threat models while preserving capabilities better than fixed-coefficient steering, and generalizes to several OOD tests.

  32. ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training

    cs.RO 2026-02 conditional novelty 5.0

    ALOE uses chunked TD bootstrapping with a pessimistic Q-ensemble to enable action-level off-policy value estimation for advantage-weighted post-training of flow-based VLA policies, reporting consistent success-rate ga...

  33. Dichotomous Diffusion Policy Optimization

    cs.LG 2025-12 conditional novelty 5.0

    DIPOLE decomposes a KL-regularized RL objective into a pair of sigmoid-weighted diffusion policies whose score combination (CFG-like) yields stable and controllable policy improvement.

  34. Robot Self-Improvement via Human-Video Dynamics Models

    cs.RO 2026-06 unverdicted novelty 4.0

    Human-video dynamics models enable cross-embodiment robot self-improvement via training-free Dynamics-Guided Action Correction, raising success rates from 40% to 81% on seven real-world tasks.