REVIEW 21 cited by
Rolling Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Rolling Diffusion Models
read the original abstract
Diffusion models have recently been increasingly applied to temporal data such as video, fluid mechanics simulations, or climate data. These methods generally treat subsequent frames equally regarding the amount of noise in the diffusion process. This paper explores Rolling Diffusion: a new approach that uses a sliding window denoising process. It ensures that the diffusion process progressively corrupts through time by assigning more noise to frames that appear later in a sequence, reflecting greater uncertainty about the future as the generation process unfolds. Empirically, we show that when the temporal dynamics are complex, Rolling Diffusion is superior to standard diffusion. In particular, this result is demonstrated in a video prediction task using the Kinetics-600 video dataset and in a chaotic fluid dynamics forecasting experiment.
Forward citations
Cited by 21 Pith papers
-
DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
DVG-WM disentangles dynamics learning and visual synthesis in video world models using flow matching and latent degradation to achieve faster inference up to 3.97 times with improved quality on LIBERO and real-world r...
-
AsyncPatch Diffusion: spatially-flexible image generation
AsyncPatch Diffusion introduces asynchronous per-region noise levels in diffusion models, proves a valid ELBO, and uses a controlled sampler to support spatially adaptive generation and native inpainting.
-
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...
-
FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring
FreqForcing stabilizes autoregressive video generation by fusing high-frequency local attention with low-frequency anchor attention, extending a 5s-trained model to 120s.
-
FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring
Spectral Self-Anchoring fuses low-frequency anchor attention with high-frequency local attention to stop autoregressive video collapse, enabling 24× length extrapolation without retraining.
-
LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
LeapTalk distills a multi-step diffusion teacher into a one-step Brownian-bridge student and reports stable streaming talking-head generation at up to 200 FPS.
-
Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction
Selective re-noising (ReRoll) during diffusion decoding improves long-horizon robot planning and policy success on maze, LIBERO-10, and unified video-action benchmarks.
-
Surprise Forcing: What to Remember, When to Skip in Long Video Generation
A training-free 'surprise' controller decides which old frames to keep in memory and which chunks need fewer denoising steps, improving long-video consistency at real-time speed.
-
Point as Skeleton: Accumulated Point Cloud Enhanced Autoregressive Generation for Closed-Loop Autonomous Driving Simulation
Point-cloud skeleton conditions and a Reset-and-Roll inference scheme enable stable frame-wise autoregressive driving video generation for closed-loop autonomous driving simulation.
-
FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars
FacePlex introduces a unified streaming model with Rolling Flow Matching and Rolling Cross-Attention to enable full-duplex joint real-time generation of speech and facial motion tokens.
-
AR Forcing: Towards Long-Horizon Robot Navigation World Model
AR Forcing trains diffusion world models by integrating standard noise prediction loss into an autoregressive loop that uses self-generated predictions as context, reducing train-inference mismatch for improved long-h...
-
Autoregressive One-Step Generative Modeling for Dynamical System Forecasting
MeLISA delivers one-step blockwise generative forecasting for dynamical systems that improves short-term accuracy and long-horizon statistical fidelity over neural operators while matching or exceeding their inference speed.
-
Autoregressive One-Step Generative Modeling for Dynamical System Forecasting
MeLISA extends pixel-space MeanFlow to one-step window-conditioned autoregressive forecasting, improving long-horizon turbulence statistics over neural-operator baselines.
-
Flow Learners for PDEs: Toward a Physics-to-Physics Paradigm for Scientific Computing
Flow learners parameterize transport vector fields to generate PDE trajectories through integration, offering a physics-to-physics organizing principle for learned solvers.
-
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
Resampling Forcing trains autoregressive video diffusion models on self-resampled degraded histories with a causal mask, achieving stable long-horizon generation without a teacher or discriminator.
-
Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
Self-Forcing++ scales autoregressive video diffusion to over 4 minutes by using self-generated segments for guidance, reducing error accumulation and outperforming baselines in fidelity and consistency.
-
DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
DVG-WM disentangles dynamics learning from visual synthesis via flow matching and latent degradation to deliver faster, higher-quality video predictions for robotic manipulation.
-
Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions
Ultra Flash introduces a cascaded streaming super-resolution framework with specialized training, upsampling, and optimization to enable real-time high-resolution video generation from low-res diffusion models.
-
Recursive Flow Matching
RecFM uses recursive self-consistency in flow matching to enable high-fidelity one- and few-step (2-4 step) generation of scientific dynamics, claiming 20x speedup over diffusion emulators and 15% lower MSE than vanil...
-
Flow Learners for PDEs: Toward a Physics-to-Physics Paradigm for Scientific Computing
Learned PDE solving should target transport over admissible futures via flow learners, not snapshot state regression.
-
Accelerating Redshift-Conditioned Galaxy Image Synthesis with One-step Generative Modeling
One-step pixel-MeanFlow models recover key galaxy morphology statistics at orders-of-magnitude lower computational cost than standard DDPM sampling while remaining weaker on fine-grained structure.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.