REVIEW 23 cited by
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper presents Diffusion Forcing, a new training paradigm where a diffusion model is trained to denoise a set of tokens with independent per-token noise levels. We apply Diffusion Forcing to sequence generative modeling by training a causal next-token prediction model to generate one or several future tokens without fully diffusing past ones. Our approach is shown to combine the strengths of next-token prediction models, such as variable-length generation, with the strengths of full-sequence diffusion models, such as the ability to guide sampling to desirable trajectories. Our method offers a range of additional capabilities, such as (1) rolling-out sequences of continuous tokens, such as video, with lengths past the training horizon, where baselines diverge and (2) new sampling and guiding schemes that uniquely profit from Diffusion Forcing's variable-horizon and causal architecture, and which lead to marked performance gains in decision-making and planning tasks. In addition to its empirical success, our method is proven to optimize a variational lower bound on the likelihoods of all subsequences of tokens drawn from the true joint distribution. Project website: https://boyuan.space/diffusion-forcing
Forward citations
Cited by 23 Pith papers
-
RoboTTT: Context Scaling for Robot Policies
A robot policy that updates its own weights during deployment can use 8,000 steps of history, steadily improving as context grows and enabling one-shot imitation from human videos.
-
MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
MV-Forcing composes temporal and view-sequential autoregression in a single diffusion model, using a recurrent 3D reconstruction model as a geometric bridge to generate arbitrarily long, multi-view consistent videos.
-
Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
Vorch-Director reduces drift in long audio-visual generation by injecting noise-level-matched prediction residuals into the conditioning history during training.
-
Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
Vorch-IR replaces one or two subjects' identities, with optional background replacement, in a driving video using indexed reference images and a textual instruction, trained on automatically synthesized pairs.
-
LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
LeapTalk distills a multi-step diffusion teacher into a one-step Brownian-bridge student and reports stable streaming talking-head generation at up to 200 FPS.
-
Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping
For video face swapping, adaptively adding swapped anchor frames at the moments of worst identity drift should make synthetic training pairs more faithful than the current first-and-last-frame-only scheme.
-
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
Post-training few-step autoregressive video generators with on-policy self-distillation using real long-video context as teacher cache reduces long-horizon degradation at zero inference cost.
-
WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
WoVR shows that reinforcement learning can improve VLA robot policies through imagined rollouts in a video world model, reporting +29.3 points on LIBERO and +30.0 points on real Franka tasks.
-
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments
Flow equivariant world models use a latent memory that shifts with the agent and with inferred object motion, giving stable long-horizon prediction under partial observability.
-
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
Context-as-Memory conditions video generation on selected historical frames chosen by camera FOV overlap, improving scene consistency in long generated videos.
-
Unlocking the Power of Diffusion Models in Sequential Recommendation: A Simple and Effective Approach
ADRec applies token-level, per-token diffusion with causal attention to sequential recommendation, reducing embedding collapse and outperforming ten baselines on six datasets.
-
Scalable Autoregressive 3D Molecule Generation
Quetzal is an autoregressive 3D molecule generator that matches diffusion-model sample quality on QM9 and GEOM while sampling much faster and enabling exact likelihood computation.
-
Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile
A three-stage pipeline combining sparse 'tile' attention with multi-step consistency distillation makes Open-Sora-Plan video generation up to 7.8x faster while keeping the aggregate VBench final score within 1%.
-
CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models
A token-pruning cache method cuts diffusion model computation by roughly half while keeping image quality, using noise magnitude, spatial clustering, and selection balance.
-
Taming Teacher Forcing for Masked Autoregressive Video Generation
Complete Teacher Forcing, conditioning masked frames on complete previous frames instead of masked ones, substantially improves frame-level autoregressive video generation quality and temporal coherence.
-
Playable Game Generation
An autoregressive latent diffusion system, PlayGen, generates real-time playable Super Mario Bros and Doom sessions on an RTX 2060, with accuracy of game mechanics measured by action-recognition metrics.
-
One Diffusion to Generate Them All
OneDiffusion shows that a single 2.8B-parameter diffusion model, trained by treating all tasks as frame sequences with varying noise scales, can handle image generation and image understanding tasks bidirectionally.
-
VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment
A VGGT-based pointwise reprojection reward with geometry-aware sampling improves video geometric consistency via SFT/DPO and causal test-time search.
-
OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference
Object-based environment inference via kernel density estimates on learned object features achieves zero-shot room retrieval and beats scene-based CLIP.
-
Video-GPT via Next Clip Diffusion
A 3.8B transformer pretrained on 70M unlabeled videos with next clip diffusion (autoregressive clean-history conditioning plus in-clip denoising) reports state-of-the-art Physics-IQ and Kinetics-600 video prediction.
-
Exploratory Diffusion Model for Unsupervised Reinforcement Learning
A diffusion-model denoising loss serves as an intrinsic reward to guide unsupervised RL exploration, plus an alternating fine-tuning scheme for diffusion policies.
-
MSC: Multi-Scale Spatio-Temporal Causal Attention for Autoregressive Video Diffusion
A description of a multi-scale causal attention framework for autoregressive video diffusion, with no empirical validation.
-
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
OC-VLA re-labels robot action targets from the robot base frame to the camera frame using the camera's extrinsic calibration, improving cross-view generalization of VLA policies.
Discussion (0). Continue with ORCID to comment.