Pith. sign in

REVIEW 23 cited by

Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.01392 v4 pith:2LABCZ5J submitted 2024-07-01 cs.LG cs.CVcs.RO

classification cs.LGcs.CVcs.RO
keywords diffusionforcingtokensnext-tokenpredictiontrainingcausalfull-sequence
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper presents Diffusion Forcing, a new training paradigm where a diffusion model is trained to denoise a set of tokens with independent per-token noise levels. We apply Diffusion Forcing to sequence generative modeling by training a causal next-token prediction model to generate one or several future tokens without fully diffusing past ones. Our approach is shown to combine the strengths of next-token prediction models, such as variable-length generation, with the strengths of full-sequence diffusion models, such as the ability to guide sampling to desirable trajectories. Our method offers a range of additional capabilities, such as (1) rolling-out sequences of continuous tokens, such as video, with lengths past the training horizon, where baselines diverge and (2) new sampling and guiding schemes that uniquely profit from Diffusion Forcing's variable-horizon and causal architecture, and which lead to marked performance gains in decision-making and planning tasks. In addition to its empirical success, our method is proven to optimize a variational lower bound on the likelihoods of all subsequences of tokens drawn from the true joint distribution. Project website: https://boyuan.space/diffusion-forcing

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RoboTTT: Context Scaling for Robot Policies

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A robot policy that updates its own weights during deployment can use 8,000 steps of history, steadily improving as context grows and enabling one-shot imitation from human videos.

  2. MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

    cs.CV 2026-07 conditional novelty 7.0 of 10

    MV-Forcing composes temporal and view-sequential autoregression in a single diffusion model, using a recurrent 3D reconstruction model as a geometric bridge to generate arbitrarily long, multi-view consistent videos.

  3. Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Vorch-Director reduces drift in long audio-visual generation by injecting noise-level-matched prediction residuals into the conditioning history during training.

  4. Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Vorch-IR replaces one or two subjects' identities, with optional background replacement, in a driving video using indexed reference images and a textual instruction, trained on automatically synthesized pairs.

  5. LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    LeapTalk distills a multi-step diffusion teacher into a one-step Brownian-bridge student and reports stable streaming talking-head generation at up to 200 FPS.

  6. Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping

    cs.CV 2026-07 conditional novelty 6.0 of 10

    For video face swapping, adaptively adding swapped anchor frames at the moments of worst identity drift should make synthetic training pairs more faithful than the current first-and-last-frame-only scheme.

  7. OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Post-training few-step autoregressive video generators with on-policy self-distillation using real long-video context as teacher cache reduces long-horizon degradation at zero inference cost.

  8. WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

    cs.RO 2026-02 conditional novelty 6.0 of 10

    WoVR shows that reinforcement learning can improve VLA robot policies through imagined rollouts in a video world model, reporting +29.3 points on LIBERO and +30.0 points on real Franka tasks.

  9. Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Flow equivariant world models use a latent memory that shifts with the agent and with inferred object motion, giving stable long-horizon prediction under partial observability.

  10. Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Context-as-Memory conditions video generation on selected historical frames chosen by camera FOV overlap, improving scene consistency in long generated videos.

  11. Unlocking the Power of Diffusion Models in Sequential Recommendation: A Simple and Effective Approach

    cs.IR 2025-05 conditional novelty 6.0 of 10

    ADRec applies token-level, per-token diffusion with causal attention to sequential recommendation, reducing embedding collapse and outperforming ten baselines on six datasets.

  12. Scalable Autoregressive 3D Molecule Generation

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Quetzal is an autoregressive 3D molecule generator that matches diffusion-model sample quality on QM9 and GEOM while sampling much faster and enabling exact likelihood computation.

  13. Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A three-stage pipeline combining sparse 'tile' attention with multi-step consistency distillation makes Open-Sora-Plan video generation up to 7.8x faster while keeping the aggregate VBench final score within 1%.

  14. CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A token-pruning cache method cuts diffusion model computation by roughly half while keeping image quality, using noise magnitude, spatial clustering, and selection balance.

  15. Taming Teacher Forcing for Masked Autoregressive Video Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Complete Teacher Forcing, conditioning masked frames on complete previous frames instead of masked ones, substantially improves frame-level autoregressive video generation quality and temporal coherence.

  16. Playable Game Generation

    cs.AI 2024-12 conditional novelty 6.0 of 10

    An autoregressive latent diffusion system, PlayGen, generates real-time playable Super Mario Bros and Doom sessions on an RTX 2060, with accuracy of game mechanics measured by action-recognition metrics.

  17. One Diffusion to Generate Them All

    cs.CV 2024-11 conditional novelty 6.0 of 10

    OneDiffusion shows that a single 2.8B-parameter diffusion model, trained by treating all tasks as frame sequences with varying noise scales, can handle image generation and image understanding tasks bidirectionally.

  18. VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

    cs.CV 2026-03 conditional novelty 5.5 of 10

    A VGGT-based pointwise reprojection reward with geometry-aware sampling improves video geometric consistency via SFT/DPO and causal test-time search.

  19. OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Object-based environment inference via kernel density estimates on learned object features achieves zero-shot room retrieval and beats scene-based CLIP.

  20. Video-GPT via Next Clip Diffusion

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A 3.8B transformer pretrained on 70M unlabeled videos with next clip diffusion (autoregressive clean-history conditioning plus in-clip denoising) reports state-of-the-art Physics-IQ and Kinetics-600 video prediction.

  21. Exploratory Diffusion Model for Unsupervised Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A diffusion-model denoising loss serves as an intrinsic reward to guide unsupervised RL exploration, plus an alternating fine-tuning scheme for diffusion policies.

  22. MSC: Multi-Scale Spatio-Temporal Causal Attention for Autoregressive Video Diffusion

    cs.CV 2024-12 reject novelty 5.0 of 10

    A description of a multi-scale causal attention framework for autoregressive video diffusion, with no empirical validation.

  23. Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

    cs.RO 2025-08 unverdicted novelty 4.0 of 10

    OC-VLA re-labels robot action targets from the robot base frame to the camera frame using the camera's extrinsic calibration, improving cross-view generalization of VLA policies.

Pith tools