Pith. sign in

REVIEW 21 cited by

Rolling Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09470 v3 pith:GJAQEQPG submitted 2024-02-12 cs.LG stat.ML

Rolling Diffusion Models

classification cs.LG stat.ML
keywords diffusionprocessrollingvideodatadynamicsfluidframes
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Diffusion models have recently been increasingly applied to temporal data such as video, fluid mechanics simulations, or climate data. These methods generally treat subsequent frames equally regarding the amount of noise in the diffusion process. This paper explores Rolling Diffusion: a new approach that uses a sliding window denoising process. It ensures that the diffusion process progressively corrupts through time by assigning more noise to frames that appear later in a sequence, reflecting greater uncertainty about the future as the generation process unfolds. Empirically, we show that when the temporal dynamics are complex, Rolling Diffusion is superior to standard diffusion. In particular, this result is demonstrated in a video prediction task using the Kinetics-600 video dataset and in a chaotic fluid dynamics forecasting experiment.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 7.0

    DVG-WM disentangles dynamics learning and visual synthesis in video world models using flow matching and latent degradation to achieve faster inference up to 3.97 times with improved quality on LIBERO and real-world r...

  2. AsyncPatch Diffusion: spatially-flexible image generation

    cs.CV 2026-06 unverdicted novelty 7.0

    AsyncPatch Diffusion introduces asynchronous per-region noise levels in diffusion models, proves a valid ELBO, and uses a controlled sampler to support spatially adaptive generation and native inpainting.

  3. Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

    eess.AS 2026-07 conditional novelty 6.0

    Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...

  4. FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring

    cs.CV 2026-07 conditional novelty 6.0

    FreqForcing stabilizes autoregressive video generation by fusing high-frequency local attention with low-frequency anchor attention, extending a 5s-trained model to 120s.

  5. FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring

    cs.CV 2026-07 conditional novelty 6.0

    Spectral Self-Anchoring fuses low-frequency anchor attention with high-frequency local attention to stop autoregressive video collapse, enabling 24× length extrapolation without retraining.

  6. LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

    cs.CV 2026-07 conditional novelty 6.0

    LeapTalk distills a multi-step diffusion teacher into a one-step Brownian-bridge student and reports stable streaming talking-head generation at up to 200 FPS.

  7. Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction

    cs.RO 2026-07 conditional novelty 6.0

    Selective re-noising (ReRoll) during diffusion decoding improves long-horizon robot planning and policy success on maze, LIBERO-10, and unified video-action benchmarks.

  8. Surprise Forcing: What to Remember, When to Skip in Long Video Generation

    cs.CV 2026-07 conditional novelty 6.0

    A training-free 'surprise' controller decides which old frames to keep in memory and which chunks need fewer denoising steps, improving long-video consistency at real-time speed.

  9. Point as Skeleton: Accumulated Point Cloud Enhanced Autoregressive Generation for Closed-Loop Autonomous Driving Simulation

    cs.CV 2026-07 conditional novelty 6.0

    Point-cloud skeleton conditions and a Reset-and-Roll inference scheme enable stable frame-wise autoregressive driving video generation for closed-loop autonomous driving simulation.

  10. FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars

    cs.AI 2026-06 unverdicted novelty 6.0

    FacePlex introduces a unified streaming model with Rolling Flow Matching and Rolling Cross-Attention to enable full-duplex joint real-time generation of speech and facial motion tokens.

  11. AR Forcing: Towards Long-Horizon Robot Navigation World Model

    cs.RO 2026-05 unverdicted novelty 6.0

    AR Forcing trains diffusion world models by integrating standard noise prediction loss into an autoregressive loop that uses self-generated predictions as context, reducing train-inference mismatch for improved long-h...

  12. Autoregressive One-Step Generative Modeling for Dynamical System Forecasting

    cs.LG 2026-05 unverdicted novelty 6.0

    MeLISA delivers one-step blockwise generative forecasting for dynamical systems that improves short-term accuracy and long-horizon statistical fidelity over neural operators while matching or exceeding their inference speed.

  13. Autoregressive One-Step Generative Modeling for Dynamical System Forecasting

    cs.LG 2026-05 conditional novelty 6.0

    MeLISA extends pixel-space MeanFlow to one-step window-conditioned autoregressive forecasting, improving long-horizon turbulence statistics over neural-operator baselines.

  14. Flow Learners for PDEs: Toward a Physics-to-Physics Paradigm for Scientific Computing

    cs.LG 2026-04 unverdicted novelty 6.0

    Flow learners parameterize transport vector fields to generate PDE trajectories through integration, offering a physics-to-physics organizing principle for learned solvers.

  15. End-to-End Training for Autoregressive Video Diffusion via Self-Resampling

    cs.CV 2025-12 conditional novelty 6.0

    Resampling Forcing trains autoregressive video diffusion models on self-resampled degraded histories with a causal mask, achieving stable long-horizon generation without a teacher or discriminator.

  16. Self-Forcing++: Towards Minute-Scale High-Quality Video Generation

    cs.CV 2025-10 conditional novelty 6.0

    Self-Forcing++ scales autoregressive video diffusion to over 4 minutes by using self-generated segments for guidance, reducing error accumulation and outperforming baselines in fidelity and consistency.

  17. DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 5.0

    DVG-WM disentangles dynamics learning from visual synthesis via flow matching and latent degradation to deliver faster, higher-quality video predictions for robotic manipulation.

  18. Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions

    cs.CV 2026-06 unverdicted novelty 5.0

    Ultra Flash introduces a cascaded streaming super-resolution framework with specialized training, upsampling, and optimization to enable real-time high-resolution video generation from low-res diffusion models.

  19. Recursive Flow Matching

    cs.LG 2026-05 unverdicted novelty 5.0

    RecFM uses recursive self-consistency in flow matching to enable high-fidelity one- and few-step (2-4 step) generation of scientific dynamics, claiming 20x speedup over diffusion emulators and 15% lower MSE than vanil...

  20. Flow Learners for PDEs: Toward a Physics-to-Physics Paradigm for Scientific Computing

    cs.LG 2026-04 conditional novelty 5.0

    Learned PDE solving should target transport over admissible futures via flow learners, not snapshot state regression.

  21. Accelerating Redshift-Conditioned Galaxy Image Synthesis with One-step Generative Modeling

    astro-ph.IM 2026-05 unverdicted novelty 4.0

    One-step pixel-MeanFlow models recover key galaxy morphology statistics at orders-of-magnitude lower computational cost than standard DDPM sampling while remaining weaker on fine-grained structure.