Pith. sign in

REVIEW 3 major objections 29 references

SAGA stabilizes autoregressive video diffusion by suppressing high-frequency latent acceleration at inference time, without retraining.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-10 13:36 UTC pith:2FFW65TC

load-bearing objection Clean training-free stabilizer for chunk-wise AR video diffusion: real multi-backbone gains, modest size, and a load-bearing premise that is only partly stress-tested. the 3 major comments →

arxiv 2607.08020 v1 pith:2FFW65TC submitted 2026-07-09 cs.CV

SAGA: Stable Acceleration Guidance for Autoregressive Video Generation

classification cs.CV
keywords autoregressive video generationvideo diffusion modelstemporal consistencyacceleration guidanceSlepian transformtraining-free inferencespectral regularizationchunk-wise rollout
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Autoregressive video diffusion generates long or streaming video by repeatedly feeding its own past latents back as causal context. That reuse amplifies small temporal errors into flicker, motion jitter, and structural drift. This paper argues that those unstable errors appear most clearly as high-frequency energy in the discrete acceleration of the latent trajectory, because acceleration acts as a second-order operator that boosts high frequencies while attenuating smooth inertial motion. SAGA is a training-free remedy that pairs two complementary steps: a structured noise initialization that cancels short-range temporal correlations while keeping longer-range structure, and spectral guidance that projects latent acceleration onto a finite-window Slepian basis and takes a gradient step away from the high-frequency modes. Applied only at inference to existing chunk-wise backbones, the method raises temporal quality and human preference while holding or improving image quality. A reader who wants reliable streaming generation without re-training large models has a concrete, plug-in reason to care.

Core claim

The paper establishes that discrete latent acceleration is an effective signal for exposing unstable high-frequency temporal perturbations in autoregressive video diffusion, and that a training-free combination of Slepian-based acceleration spectral guidance plus structured opposite-correlation AR noise initialization consistently improves temporal quality across chunk-wise backbones while preserving visual fidelity.

What carries the argument

SAGA: structured AR noise initialization (SAN) that superposes two variance-preserving AR(1) streams with opposite correlations so lag-1 autocorrelation is zero, plus acceleration-domain spectral guidance (SG) that projects short-window discrete second differences onto a band-limited Slepian/DPSS basis and descends the high-frequency kinematic energy.

Load-bearing premise

High-frequency energy in short-window latent acceleration is mostly non-physical instability rather than legitimate motion, so suppressing it stabilizes rollout without systematically harming the pretrained generative content.

What would settle it

On matched prompts and seeds, if applying SAGA fails to reduce high-frequency acceleration power while temporal metrics and human preference stay flat or reverse, or if per-frame image quality drops, the central claim does not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper proposes SAGA, a training-free inference-time method for stabilizing chunk-wise autoregressive video diffusion. It attributes rollout failures (flicker, jitter, drift) to amplification of high-frequency temporal perturbations in discrete latent acceleration, and counters them with two components: Structured Autoregressive Noise (SAN), a variance-preserving mixture of opposite-correlation AR(1) streams that zeros lag-1 noise correlation while retaining longer-range structure, and Spectral Guidance (SG), which projects short-window latent acceleration onto a DPSS/Slepian basis and takes a gradient step to suppress high-frequency kinematic energy (Eqs. 7–8). Applied without retraining to CausVid, Self-Forcing, and Causal-Forcing, SAGA improves VBench Temporal Quality (e.g., Self-Forcing TQ 97.30→97.91, IQ 69.60→70.51), shows reduced acceleration RMS/HF power on matched videos, and is preferred in a 500-judgment human study. Ablations isolate SAN and SG, compare Slepian vs FFT, and test chunk-wise vs frame-wise rollout and longer horizons.

Significance. If the result holds, SAGA is a practical, plug-in stabilizer for the dominant high-quality AR-diffusion setting (chunk-wise causal DiT rollouts). Training-free applicability across three backbones, multi-metric VBench gains that do not systematically sacrifice image quality when both components are used, paired spectral analysis with bootstrap CIs, and a human preference study are concrete strengths. The acceleration-domain framing and finite-window Slepian implementation are a clear, reusable design pattern for short-context temporal regularization. Gains are modest in absolute terms, but the method is immediately usable and the experimental package is stronger than typical inference-time video guidance papers.

major comments (3)
  1. Sec. 4.1 and Eqs. 7–8 rest on the load-bearing premise that high-frequency discrete latent acceleration over short AR windows is predominantly non-physical instability rather than legitimate motion. The paper’s own evidence only partially tests this: SG alone raises TQ (97.30→97.58) but lowers AQ/IQ (Table 2); frame-wise rollout with insufficient temporal support shows no gain or a slight TQ drop (Table 4); and Fig. 3 reports reductions in the exact quantity being minimized, so it is a weak independent diagnostic of artifact vs. content. A stronger test is needed—e.g., motion-content controls (high-frequency legitimate motion prompts), or a comparison against a non-acceleration high-frequency regularizer—to show that the suppressed energy is not systematically over-smoothing the generative prior.
  2. Table 3 shows that FFT-based guidance nearly matches Slepian (TQ 97.89 vs 97.91). The manuscript already softens the claim that DPSS is the primary novelty, but the abstract and contributions still foreground “finite-window Slepian projections” as a core element. The central claim should be reframed more explicitly around the acceleration-domain objective itself, with Slepian retained as a principled default rather than a decisive empirical driver, unless additional evidence (e.g., leakage-sensitive windows or qualitative failure modes of FFT) is provided.
  3. Inference configuration (Sec. 5.1) uses a single global hyperparameter set (η=3, ρ=0.9, NW=1.5, Kc=3) across prompts and three backbones. Free parameters are numerous relative to the reported sensitivity analysis. At minimum, a compact sensitivity or transfer plot for η and Kc (and confirmation that the same set was not tuned on the evaluation prompts) is needed to support the “no per-model retuning” claim that underpins the training-free multi-backbone result.

Circularity Check

1 steps flagged

No definitional circularity in the method or claims; the only mild circularity is using reduction of the optimized quantity (acceleration energy) as mechanism evidence.

specific steps
  1. fitted input called prediction [Sec. 5.2 Acceleration-Spectrum Analysis / Fig. 3]
    "SAGA reduces acceleration RMS by 16.7% and total acceleration power by 30.2%... For f>0.375, SAGA reduces global and local high-frequency acceleration power by 25.7% and 28.4%, respectively. This spectral response aligns with the guidance objective in Sec. 4.3 and supports our hypothesis that autoregressive rollout instability is associated with high-frequency acceleration components"

    The guidance step (Eq. 8) is exactly a gradient descent step on E_kin (Eq. 7), the high-frequency Slepian energy of discrete acceleration. Therefore reductions in acceleration RMS, total power, and HF power are expected by construction of the optimizer, not an independent diagnostic that the suppressed energy was non-physical instability. This does not force the VBench or human-preference claims, but it weakens Fig. 3 as mechanism evidence.

full rationale

SAGA is an inference-time guidance method whose central claims are empirical improvements on external VBench temporal metrics (SC/BC/TF/MS), aesthetic/image quality, and human preference, none of which are definitionally equal to the guidance objective E_kin. The kinematic prior (Eqs. 3, 7–8) and structured AR noise (Eqs. 5a–5b) are design choices motivated by a frequency-response analysis (Eq. 4), not fitted parameters that force the reported TQ/IQ gains. Hyperparameters are fixed globally (η=3, ρ=0.9, NW=1.5, Kc=3) rather than re-fit per claim. The sole mild circularity is that Fig. 3’s reductions in acceleration RMS and high-frequency power are near-direct consequences of minimizing E_kin, so that plot is weak as independent mechanism proof; it does not make the main performance claims circular. No self-citation uniqueness theorems, ansatz smuggling, or renaming of known results load-bear the derivation. Score 2 reflects only this minor diagnostic circularity.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

The central claim rests on standard finite-difference/spectral math, domain assumptions about AR error accumulation and inertial motion, several hand-chosen free hyperparameters, and the invented SAGA objective/entities. No new physical particles or forces; the invented pieces are algorithmic constructs whose only external handle is improved VBench/human scores.

free parameters (5)
  • guidance strength η = 3
    Controls gradient step size on E_kin; fixed globally at 3 with no derivation from first principles.
  • AR coefficient ρ = 0.9
    Sets opposite AR(1) correlations for SAN; chosen as 0.9 to zero lag-1 correlation while keeping lag-2 structure.
  • DPSS time-bandwidth product NW = 1.5
    Sets Slepian concentration bandwidth; fixed at 1.5.
  • spectral cutoff index Kc = 3
    Separates low- vs high-frequency DPSS modes retained/suppressed; set to 3 from ~2NW rule of thumb.
  • SAN mixing angle θ = π/4
    Balances low- and high-frequency AR streams; set to π/4 for equal mix and R(1)=0.
axioms (5)
  • standard math Discrete second-difference acceleration attenuates low-frequency motion and amplifies high-frequency variations with gain 4(1−cosω)².
    Sec. 4.1 harmonic analysis; standard finite-difference frequency response.
  • domain assumption Recursive reuse of generated latents as causal context accumulates temporal errors into flicker/jitter/drift in AR video diffusion.
    Stated in Introduction and Problem Formulation; supported by cited AR-diffusion literature but not re-derived here.
  • domain assumption Natural/inertial motion concentrates acceleration energy in low frequencies; high-frequency acceleration is a useful proxy for non-physical instability.
    Physics-informed prior in Sec. 4.1 and Related Work; load-bearing for why suppressing HF acceleration is desirable.
  • standard math DPSS/Slepian sequences maximize in-band energy concentration under finite support and reduce spectral leakage vs plain FFT on short windows.
    Sec. 4.3 citing Slepian/Thomson; used to justify the projection basis.
  • ad hoc to paper A single global hyperparameter set (η, ρ, NW, Kc) transfers across prompts and three AR backbones without per-model retuning.
    Inference Configuration Sec. 5.1; experimental convenience treated as general practice.
invented entities (2)
  • SAGA acceleration-domain kinematic energy E_kin via Slepian coefficients independent evidence
    purpose: Inference-time objective whose gradient steers predicted clean latents away from high-frequency acceleration modes.
    Defined in Eqs. 3 and 7–8; not a previously standard loss in AR video diffusion.
  • Structured Autoregressive Noise (SAN) opposite-correlation AR(1) initialization independent evidence
    purpose: Construct kinematically neutral noise with R(1)=0 while preserving longer-range correlation and unit variance.
    Eqs. 5a–5b; algorithmic construct specific to this paper’s design.

pith-pipeline@v1.1.0-grok45 · 18549 in / 3796 out tokens · 41347 ms · 2026-07-10T13:36:24.948554+00:00 · methodology

0 comments
read the original abstract

Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context can amplify temporal errors, resulting in flickering, motion jitter, and structural drift. In this paper, we investigate this failure mode from a spectral kinematic perspective and identify discrete latent acceleration as an effective signal for revealing unstable high-frequency temporal perturbations. To this end, we propose SAGA, a training-free \textbf{\textit{s}}table \textbf{\textit{a}}cceleration \textbf{\textit{g}}uidance approach for \textbf{\textit{a}}utoregressive video generation. SAGA integrates an acceleration domain spectral guidance objective based on finite-window Slepian projections with a structured autoregressive noise initialization strategy that suppresses short-range temporal correlations while preserving long-range motion structure. Without retraining or modifying the backbone, SAGA can be directly applied to existing chunk-wise autoregressive diffusion models, which is the prevalent setting for high-quality generation. Extensive experiments show that SAGA consistently improves temporal quality across multiple autoregressive diffusion models. On Self-Forcing, SAGA improves Temporal Quality from 97.30 to 97.91 and Image Quality from 69.60 to 70.51. Moreover, spectral analysis and human preference studies demonstrate that SAGA reduces temporal instability while maintaining visual fidelity.

Figures

Figures reproduced from arXiv: 2607.08020 by Minh-Triet Tran, Tam V. Nguyen, Thanh-Nhan Vo, Trong-Thuan Nguyen, Trung-Hoang Le.

Figure 1
Figure 1. Figure 1: End-to-end SAGA workflow for autoregressive text-to-video generation. A frozen causal DiT rolls out latent chunks using SG (Spectral Guidance 4.3) and SAN (Structured AR Noise 4.2), then decodes the generated trajectory z1:T into the video. Detailed modules are shown in [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of our proposed approach, termed SAGA, a novel training-free ap￾proach for stable acceleration guidance in autoregressive video generation. Our ap￾proach stabilizes autoregressive video diffusion rollout through two complementary components: (Top) a structured frequency-aware initialization that establishes a kine￾matically neutral motion prior, and (Bottom) a training-free spectral guidance mecha… view at source ↗
Figure 3
Figure 3. Figure 3: Global and local acceleration spectra over 100 matched videos, together with paired reductions in acceleration RMS, total acceleration power, global high-frequency power, and local high-frequency power. The high-frequency band is defined as f > 0.375, and error bars denote paired-bootstrap 95% confidence intervals. matched videos generated by Self-Forcing and Self-Forcing+SAGA. As shown in [PITH_FULL_IMAG… view at source ↗
Figure 4
Figure 4. Figure 4: Human preference study with tie judgments. Pairwise comparisons between Self-Forcing and Self-Forcing+SAGA are reported at both the individual-vote and video-level majority aggregation levels, with ties retained as a separate outcome [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative comparison between Self-Forcing and Self-Forcing+SAGA using uniformly sampled video frames. Red dashed boxes highlight temporal artifacts in the Self-Forcing baseline, including structural drift, background inconsistency, and local appearance flickering. Best viewed in color and zoomed in. 58 videos versus 34, with only 8 ties. Excluding ties, SAGA achieves 58.5% of vote-level and 63.0% of vide… view at source ↗
Figure 6
Figure 6. Figure 6: Long-horizon evalua￾tion of Self-Forcing and Self￾Forcing+SAGA at different video durations. Autoregressive video generation becomes increasingly challenging over long rollout horizons, as local temporal perturbations can propagate through recursively generated causal context. To assess SAGA under ex￾tended rollout, we compare Self-Forcing [10] and Self-Forcing+SAGA at durations of 5, 10, and 30 seconds us… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

29 extracted references · 29 canonical work pages · 4 internal anchors

  1. [1]

    In: International Conference on Learning Representations

    Bansal,A.,Chu,H.M.,Schwarzschild,A.,Sengupta,R.,Goldblum,M.,Geiping,J., Goldstein, T.: Universal guidance for diffusion models. In: International Conference on Learning Representations. vol. 2024, pp. 51304–51323 (2024) 4

  2. [2]

    In: Proceedings of the IEEE/CVF international conference on computer vision

    Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 9650–9660 (2021) 8

  3. [3]

    Advances in Neural Information Processing Systems37, 24081–24125 (2025) 1, 3

    Chen, B., Martí Monsó, D., Du, Y., Simchowitz, M., Tedrake, R., Sitzmann, V.: Diffusion forcing: Next-token prediction meets full-sequence diffusion. Advances in Neural Information Processing Systems37, 24081–24125 (2025) 1, 3

  4. [4]

    Chen, H., Xia, M., He, Y., Zhang, Y., Cun, X., Yang, S., Xing, J., Liu, Y., Chen, Q., Wang, X., Weng, C., Shan, Y.: Videocrafter1: Open diffusion models for high- quality video generation (2023) 1, 9

  5. [5]

    In: The Eleventh International Conference on Learning Representations (2023) 4

    Chung, H., Kim, J., Mccann, M.T., Klasky, M.L., Ye, J.C.: Diffusion posterior sam- pling for general noisy inverse problems. In: The Eleventh International Conference on Learning Representations (2023) 4

  6. [6]

    In: International Conference on Learning Representations

    Deng, H., Pan, T., Diao, H., Luo, Z., Cui, Y., Lu, H., Shan, S., Qi, Y., Wang, X.: Autoregressive video generation without vector quantization. In: International Conference on Learning Representations. vol. 2025, pp. 44730–44745 (2025) 3

  7. [7]

    Advances in neural information processing systems34, 8780–8794 (2021) 4

    Dhariwal, P., Nichol, A.: Diffusion models beat gans on image synthesis. Advances in neural information processing systems34, 8780–8794 (2021) 4

  8. [8]

    In: ICML 2026 Workshop on Foundations of Deep Generative Models: Understanding Memoriza- tion, Generalization, and Reasoning (2026) 2, 3

    Han, W., Kang, S., Jun, Y., Chen, M.H., Yang, F.E., Hwang, S.J.: Physics in 2- steps: Locking motion priors before visual refinement erases them. In: ICML 2026 Workshop on Foundations of Deep Generative Models: Understanding Memoriza- tion, Generalization, and Reasoning (2026) 2, 3

  9. [9]

    Ho, J., Salimans, T.: Classifier-free diffusion guidance (2022) 4

  10. [10]

    Advances in Neural Information Processing Systems38, 167283–167308 (2026) 1, 2, 3, 4, 8, 9, 10, 14

    Huang, X., Li, Z., He, G., Zhou, M., Shechtman, E.: Self forcing: Bridging the train-test gap in autoregressive video diffusion. Advances in Neural Information Processing Systems38, 167283–167308 (2026) 1, 2, 3, 4, 8, 9, 10, 14

  11. [11]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024) 2, 8, 9, 14

    Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., Wang, Y., Chen, X., Wang, L., Lin, D., Qiao, Y., Liu, Z.: VBench: Comprehensive benchmark suite for video generative models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024) 2, 8, 9, 14

  12. [12]

    In: Proceedings of the IEEE/CVF international conference on computer vision

    Ke, J., Wang, Q., Wang, Y., Milanfar, P., Yang, F.: Musiq: Multi-scale image quality transformer. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 5148–5157 (2021) 8

  13. [13]

    arXiv preprint arXiv:2602.14027 (2026) 4

    Li, J., Fu, X., Peng, X., Chen, W., Zheng, Y., Zhao, T., Wang, J., Chen, F., Wang, X., So, H.K.H.: Train short, inference long: Training-free horizon extension for autoregressive video generation. arXiv preprint arXiv:2602.14027 (2026) 4

  14. [14]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Li, Z., Zhu, Z.L., Han, L.H., Hou, Q., Guo, C.L., Cheng, M.M.: Amt: All-pairs multi-field transforms for efficient frame interpolation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9801– 9810 (2023) 8

  15. [15]

    In: European conference on computer vision

    Liang, J., Fan, Y., Zhang, K., Timofte, R., Van Gool, L., Ranjan, R.: Movideo: Motion-aware video generation with diffusion model. In: European conference on computer vision. pp. 56–74. Springer (2024) 2, 3 16 Thanh-Nhan Vo et al

  16. [16]

    Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

    Liu, Y., Zhang, K., Li, Y., Yan, Z., Gao, C., Chen, R., Yuan, Z., Huang, Y., Sun, H., Gao, J., et al.: Sora: A review on background, technology, limitations, and opportunities of large vision models. arXiv preprint arXiv:2402.17177 (2024) 9

  17. [17]

    arXiv preprint arXiv:2507.16869 (2025) 3

    Ma, Y., Feng, K., Hu, Z., Wang, X., Wang, Y., Zheng, M., Wang, B., Wang, Q., He, X., Wang, H., et al.: Controllable video generation: A survey. arXiv preprint arXiv:2507.16869 (2025) 3

  18. [18]

    Movie Gen: A Cast of Media Foundation Models

    Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A., Vyas, A., Shi, B., Ma, C.Y., Chuang, C.Y., et al.: Movie gen: A cast of media foundation models. arXiv preprint arXiv:2410.13720 (2024) 8, 10

  19. [19]

    In: International conference on machine learning

    Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning. pp. 8748–8763. PmLR (2021) 8

  20. [20]

    Bell System Technical Journal57(5), 1371–1430 (1978) 2, 4

    Slepian, D.: Prolate spheroidal wave functions, fourier analysis, and uncer- tainty—v: The discrete case. Bell System Technical Journal57(5), 1371–1430 (1978) 2, 4

  21. [21]

    Proceedings of the IEEE70(9), 1055–1096 (1982) 2, 4, 9

    Thomson, D.J.: Spectrum estimation and harmonic analysis. Proceedings of the IEEE70(9), 1055–1096 (1982) 2, 4, 9

  22. [22]

    Wan: Open and Advanced Large-Scale Video Generative Models

    Wan, T., Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.W., Chen, D., Yu, F., Zhao, H., Yang, J., Zeng, J., Wang, J., Zhang, J., Zhou, J., Wang, J., Chen, J., Zhu, K., Zhao, K., Yan, K., Huang, L., Feng, M., Zhang, N., Li, P., Wu, P., Chu, R., Feng, R., Zhang, S., Sun, S., Fang, T., Wang, T., Gui, T., Weng, T., Shen, T., Lin, W., Wang, W., Wang, W., Zhou, W.,...

  23. [23]

    arXiv preprint arXiv:2503.10704 (2025) 2, 4

    Wang, J., Zhang, F., Li, X., Tan, V.Y., Pang, T., Du, C., Sun, A., Yang, Z.: Er- ror analyses of auto-regressive video diffusion models: A unified framework. arXiv preprint arXiv:2503.10704 (2025) 2, 4

  24. [24]

    ACM Computing Surveys57(2), 1–42 (2024) 1, 3

    Xing, Z., Feng, Q., Chen, H., Dai, Q., Hu, H., Xu, H., Wu, Z., Jiang, Y.G.: A survey on video diffusion models. ACM Computing Surveys57(2), 1–42 (2024) 1, 3

  25. [25]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision

    Yang, X., Li, B., Zhang, Y., Yin, Z., Bai, L., Ma, L., Wang, Z., Cai, J., Wong, T.T., Lu, H., et al.: Vlipp: Towards physically plausible video generation with vision and language informed physical prior. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 12360–12370 (2025) 2, 3

  26. [26]

    In: International Conference on Learning Representations

    Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., et al.: Cogvideox: Text-to-video diffusion models with an expert transformer. In: International Conference on Learning Representations. vol. 2025, pp. 83048–83077 (2025) 1, 9

  27. [27]

    In: CVPR (2025) 1, 3, 4, 8, 9

    Yin, T., Zhang, Q., Zhang, R., Freeman, W.T., Durand, F., Shechtman, E., Huang, X.: From slow bidirectional to fast autoregressive video diffusion models. In: CVPR (2025) 1, 3, 4, 8, 9

  28. [28]

    ACM Computing Surveys58(12), 1–41 (2026) 3

    Yin, Z., Chen, K., Bai, X., Jiang, R., Li, J., Li, H., Liu, J., Xiang, Y., Yu, J., Zhang, M.: A survey: spatiotemporal consistency in video generation. ACM Computing Surveys58(12), 1–41 (2026) 3

  29. [29]

    Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

    Zhu, H., Zhao, M., He, G., Su, H., Li, C., Zhu, J.: Causal forcing: Autoregressive diffusion distillation done right for high-quality real-time interactive video genera- tion. arXiv preprint arXiv:2602.02214 (2026) 1, 3, 8, 9