REVIEW 4 major objections 5 minor 88 references
Generating time-consistent dynamics with discriminator-guided image diffusion models
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A time-consistency discriminator lets pretrained image diffusion models generate stable, realistic dynamics over centuries.
desk verdict A genuinely useful inference-time discriminator guidance for image diffusion models, with broad and serious evaluation; the centennial-stability claim overreaches its global-mean evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the time-consistency discriminator $D_\theta(x_t^{n+1}; x_0^{n-1:n}, t)$, a binary classifier conditioned on the two most recent denoised frames and on the diffusion noise time. Its guidance term is the gradient with respect to $x_t^{n+1}$ of $\log(D_\theta/(1-D_\theta))$, which equals $\nabla \log[p(x_t^{n+1}\mid \text{past}) / p(x_t^{n+1})]$; adding it to the unconditional score turns the reverse SDE into a conditional sampler. The discriminator is trained with cross-entropy on real next frames versus importance-sampled corrupted frames and random crops, then applied in both solver steps of the stochastic EDM sampler.
What would settle it
Disable the discriminator guidance partway through a guided precipitation rollout and continue with unconditional sampling; if the autocorrelation, Hovmöller statistics, and global mean stay stable for decades, then the guidance is not what enforces long-run stability, whereas a rapid drift would support the paper's causal claim.
Extended reading notes
Core claim
On its own terms, the paper establishes that temporal consistency can be imposed on an unconditionally trained image diffusion model by adding the score-like guidance term $d_\theta(x_t^{n+1}; x_0^{n-1:n}, t) = \nabla_{x_t^{n+1}} \log(D_\theta/(1-D_\theta))$ to the reverse SDE, where $D_\theta$ is a discriminator trained to separate the conditional density $p(x^{n+1}\mid x^n, x^{n-1})$ from the marginal $p(x^{n+1})$ at noise level $t$. This term is the gradient of the log-density ratio, so it steers the denoising trajectory toward frames that belong after the already-generated frames. The paper argues, and demonstrates on 2D Navier-Stokes turbulence and ERA5 daily precipitation, that this suffices to turn a pretrained image diffusion model into a dynamical emulator with realistic autocorrelation, Hovmöller structure, extreme-event waiting times, and forecast calibration, and that the resulting autoregressive rollout remains stable over more than a century, whereas a video diffusion baseline exhibits drifting global means.
Load-bearing premise
The load-bearing premise is that a discriminator trained on only the local one-step transition (the last two denoised frames) produces gradients that keep an autoregressive rollout accurate for hundreds of steps; the century-scale stability is demonstrated empirically in ten 100-year runs and one 170-year run, but it is not theoretically guaranteed, and the video diffusion baseline fails the same test.
Editorial extensions
If this is right
- Any pretrained image diffusion model with access to clean conditioning frames can be converted into a dynamical emulator without retraining; the discriminator trains separately on target data.
- Long autoregressive rollouts (more than 100 years at daily steps for precipitation) remain stable under guidance, while the video diffusion baseline develops mean drift, suggesting that guidance prevents error accumulation.
- Ensemble forecasts from the guided model are better calibrated (spread-skill ratio) and have lower spatial bias than the video diffusion model, though slightly worse CRPS at the shortest lead times.
- The guidance adds roughly 3-8% to generation time, so it can be attached to existing pretrained models and cheap discriminators.
- The recovered Hovmöller wave structures and extreme-event waiting time distributions show that the method reproduces dynamical statistics, not just pixel-level sharpness.
Reading between the lines
- If the discriminator gradient remains informative near the end of denoising, the same recipe should transfer to latent image diffusion models and to video processing tasks such as downscaling or inpainting, because it only needs a clean past frame.
- The $m=1$ conditioning makes the method naturally suited to first-order Markov dynamics; systems with longer memory or slower modes may need an $m>1$ discriminator plus a long-range statistic term, which would require retesting the stability claim.
- A direct extension the paper only mentions in passing: apply the same locally-trained discriminator to sampling from a video diffusion model, to see whether its centennial-scale drift is corrected by the same guidance.
- The balance of results suggests that for climate emulation, long-run stability and calibration may matter more than short-lead forecast skill, so the guided image model may be the more useful configuration for century-scale studies.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an inference-time guidance method that uses a separately trained time-consistency discriminator to steer pretrained image diffusion models toward temporally consistent autoregressive generation. The discriminator classifies whether a noised image is the next frame given the current and previous clean frames, and its logit gradient is added to the unconditional score in the reverse SDE. The method is evaluated on 2D Navier-Stokes turbulence and global ERA5 daily precipitation, with comparisons to an unconditional image diffusion model and a video diffusion model trained from scratch. Metrics include Wasserstein distances between Hovmöller-diagram rows, autocorrelation functions, CRPS, spread-skill ratios, EOFs, waiting-time distributions, bias maps, and long-run stability assessed by annual rolling global means. The paper claims comparable temporal consistency to the video diffusion model, improved calibration and lower bias, and stable centennial-scale climate simulations.
Significance. If the claims hold, the contribution is practically significant: it offers a way to reuse pretrained image diffusion models for spatiotemporal generation without architecture changes or finetuning, with modest inference overhead (reported as 3%-8%). The empirical evaluation is unusually broad, drawing on standard metrics from fluid dynamics and climate science, and the manuscript reports detailed hyperparameters, training configurations, and sampling pseudocode. The main uncertainty is whether the headline stability claim is supported by the evidence, given that it rests on a single aggregate statistic, and whether the reported quantitative advantages are robust to hyperparameter choices and statistical noise. These issues are addressable and do not undermine the core idea, but they are material to the paper's central claims.
major comments (4)
- [Section 5, Figure 7] The claim of 'stable centennial-scale climate simulations' is supported only by annual rolling global-mean plots. Because the discriminator is trained on local one-step transitions with m=1 (Eqs. 4-6), nothing in the method explicitly controls slow spatial patterns, regional biases, or multi-decadal variability; a flat global mean is compatible with compensating regional drifts or a gradual loss of spectral variance. The manuscript should add diagnostics for the 100-year and 170-year rollouts, such as spatial power spectra, EOF stability over time, local ACF in early versus late decades, and regional bias maps, ideally with trends of these quantities over the run. Without such validation, the stability advantage over the video DM is not established beyond a single aggregate statistic.
- [Sections 4-5, Table 1] The guidance strength λ and conditioning length m are tuned per dataset (λ=14 for vorticity, λ=68 for precipitation, m=1) without any sensitivity analysis. Since λ controls the relative weight of the discriminator gradient in Eq. (5), the reported improvements in calibration, bias, and long-run stability may depend critically on this choice. The authors should provide a sensitivity study over λ, and ideally m, for at least one headline metric per dataset, such as global-mean drift, ACF error, or CRPS, to demonstrate robustness and rule out hyperparameter selection effects.
- [Section 5, Figures 4c, 6c, 11, 12] The quantitative comparisons of CRPS and spread-skill ratio are reported without confidence intervals or significance tests. Differences between the guided DM and video DM at individual lead times are small, and statements such as 'improved calibration' and 'slightly outperforming' are not supported by uncertainty quantification. The authors should add bootstrap confidence intervals or significance tests over the 100 forecasts for CRPS, SSR, and the error curves in Figures 14-15, so that readers can assess whether the reported differences are meaningful.
- [Section 3 and Algorithm 1] The discriminator is trained using ground-truth clean conditioning frames (Eq. 6), but during autoregressive inference it is applied to previously generated frames. This train/inference distribution shift is not discussed and is a potential source of error accumulation in long rollouts, especially for the centennial-scale claim. The authors should discuss this issue explicitly and, ideally, measure how the discriminator's classification accuracy degrades when conditioned on generated frames rather than ground-truth frames.
minor comments (5)
- [Appendix B.1, Eq. (13)] Equation (13) is a tautology and does not justify Equation (14); the latter follows directly from Bayes' theorem. Consider replacing Eq. (13) with the correct factorization p(x^{n+1}|x^{(n-m):n}) = p(x^{n+1}) p(x^{(n-m):n}|x^{n+1}) / p(x^{(n-m):n}), or removing it entirely.
- [Figure 2 caption and main text] The phrase 'reserve diffusion process' should be 'reverse diffusion process' in the caption of Figure 2.
- [Figure 7] The time axis of Figure 7 extends beyond the ERA5 test period (2011-2020), but the ground-truth line appears constant; the authors should clarify how the ground-truth reference is represented for later years and state how the 170-year guided run is initialized.
- [Table 1] The sampling parameters Stmin and Stmax are not defined in the table or the main text; they should be defined consistently with the stochastic sampler notation in Appendix B.3.
- [Section 5] The statement that 'all generative DMs remain sharp' is qualitative; sharpness is not defined or measured. Either define a sharpness metric or soften the claim.
Circularity Check
No significant circularity: the central result is an empirical demonstration on held-out data; the guidance identity is a standard score-ratio decomposition, not a restatement of the target claim.
full rationale
The paper's derivation chain is self-contained. The time-consistency guidance is obtained from the standard optimal-discriminator identity (Eqs. 3-4 and 9-12), which expresses the conditional score as the unconditional score plus the gradient of the log-density ratio; this is a mathematical identity, not a restatement of the target claim. The discriminator is trained with cross-entropy on real consecutive frames versus corrupted or non-consecutive frames (Eq. 6), but the evaluation metrics — Wasserstein distances of Hovmöller rows, autocorrelation functions, CRPS, spread-skill ratio, waiting-time distributions, and EOFs — are computed on held-out test periods and are not equal to the discriminator's training objective. The centennial-stability claim is an empirical rollout result (Section 5, Figure 7), not a consequence of the training loss; whether the global-mean-only diagnostic is sufficient is a robustness or validation concern, not circularity. The authors' self-citations [33, 34, 56] appear only as background for related downscaling and weather/climate-generation work and are not load-bearing for the proposed method. No parameter is fitted to a target metric and then reported as a prediction of that metric.
Assumptions & free parameters
free parameters (3)
- Guidance strength lambda =
14 (vorticity), 68 (precipitation)
- Conditioning history length m =
1
- Negative-sample importance sampling parameters (mu, sigma_step) =
mu=1, sigma_step=2
assumptions (5)
- standard math Reverse-time SDE and learned score function (Eqs. 1-2) from Song et al. 2021 and Karras et al. 2022 are valid for the data distributions used.
- domain assumption An optimal discriminator has the density-ratio form in Eq. 3, and the cross-entropy loss trains toward it.
- domain assumption The gradient of the trained discriminator provides a useful approximation of the conditional score ratio at every diffusion noise level.
- domain assumption Local temporal conditioning with m=1 past frame is sufficient to keep long autoregressive rollouts stable and unbiased.
- domain assumption ERA5 reanalysis and the Navier-Stokes simulation are treated as ground-truth target distributions for training and evaluation.
Cite this review
Pith. "Pith review of Generating time-consistent dynamics with discriminator-guided image diffusion models." pith.science (2026). https://pith.science/paper/5CTWVYW7
@misc{pith2026250509089,
author = {Pith},
title = {Pith review of: Generating time-consistent dynamics with discriminator-guided image diffusion models},
year = {2026},
howpublished = {\url{https://pith.science/paper/5CTWVYW7}},
note = {Machine review of arXiv:2505.09089}
}
read the original abstract
Realistic temporal dynamics are crucial for many video generation, processing and modelling applications, e.g. in computational fluid dynamics, weather prediction, or long-term climate simulations. Video diffusion models (VDMs) are the current state-of-the-art method for generating highly realistic dynamics. However, training VDMs from scratch can be challenging and requires large computational resources, limiting their wider application. Here, we propose a time-consistency discriminator that enables pretrained image diffusion models to generate realistic spatiotemporal dynamics. The discriminator guides the sampling inference process and does not require extensions or finetuning of the image diffusion model. We compare our approach against a VDM trained from scratch on an idealized turbulence simulation and a real-world global precipitation dataset. Our approach performs equally well in terms of temporal consistency, shows improved uncertainty calibration and lower biases compared to the VDM, and achieves stable centennial-scale climate simulations at daily time steps.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video Diffusion Models. Advances in Neural Information Processing Systems, 35:8633–8646, December 2022
2022
-
[2]
Kingma, Ben Poole, Mohammad Norouzi, David J
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. Imagen Video: High Definition Video Generation with Diffusion Models, October 2022. arXiv:2210.02303 [cs]
arXiv 2022
-
[3]
Open-Sora: Democratizing Efficient Video Production for All, December 2024
Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-Sora: Democratizing Efficient Video Production for All, December 2024. arXiv:2412.20404 [cs]
arXiv 2024
-
[4]
Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models, January
Wen Wang, Yan Jiang, Kangyang Xie, Zide Liu, Hao Chen, Yue Cao, Xinlong Wang, and Chunhua Shen. Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models, January
-
[5]
Dimakis, Morteza Mardani, Nikola B
Giannis Daras, Weili Nie, Karsten Kreis, Alexandros G. Dimakis, Morteza Mardani, Nikola B. Kovachki, and Arash Vahdat. Warped Diffusion: Solving Video Inverse Problems with Image Diffusion Models. Advances in Neural Information Processing Systems, 37:101116–101143, December 2024
2024
-
[6]
Conditional neural field latent diffusion model for generating spatiotemporal turbulence
Pan Du, Meet Hemant Parikh, Xiantao Fan, Xin-Yang Liu, and Jian-Xun Wang. Conditional neural field latent diffusion model for generating spatiotemporal turbulence. Nature Communi- cations, 15(1):10416, November 2024
2024
-
[7]
T. Li, L. Biferale, F. Bonaccorso, M. A. Scarpolini, and M. Buzzicotti. Synthetic Lagrangian turbulence by generative diffusion models. Nature Machine Intelligence, 6(4):393–403, April 2024
2024
-
[8]
From Zero to Turbulence: Generative Modeling for 3D Flow Simulation, March 2024
Marten Lienen, David Lüdke, Jan Hansen-Palmus, and Stephan Günnemann. From Zero to Turbulence: Generative Modeling for 3D Flow Simulation, March 2024. arXiv:2306.01776 [physics]
arXiv 2024
Show all 88 references
-
[9]
Andersson, Andrew El-Kadi, Do- minic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson
Ilan Price, Alvaro Sanchez-Gonzalez, Ferran Alet, Tom R. Andersson, Andrew El-Kadi, Do- minic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson. Probabilistic weather forecasting with machine learning. Nature, 637(8044):84–90,...
2025
-
[10]
Generative emula- tion of weather forecast ensembles with diffusion models
Lizao Li, Robert Carver, Ignacio Lopez-Gomez, Fei Sha, and John Anderson. Generative emula- tion of weather forecast ensembles with diffusion models. Science Advances, 10(13):eadk4489, March 2024. Publisher: American Association for the Advancement of Science
2024
-
[11]
Bretherton, and Stephan Mandt
Prakhar Srivastava, Ruihan Yang, Gavin Kerrigan, Gideon Dresdner, Jeremy McGibbon, Christo- pher S. Bretherton, and Stephan Mandt. Precipitation Downscaling with Spatiotemporal Video Diffusion. Advances in Neural Information Processing Systems, 37:56374–56400, December 2024
2024
-
[12]
DiffESM: Conditional Emulation of Temperature and Precipitation in Earth System Models With 3D Diffusion Models
Seth Bassetti, Brian Hutchinson, Claudia Tebaldi, and Ben Kravitz. DiffESM: Conditional Emulation of Temperature and Precipitation in Earth System Models With 3D Diffusion Models. Journal of Advances in Modeling Earth Systems, 16(10):e2023MS004194, October 2024
2024
-
[13]
Bretherton, and Rose Yu
Salva Rühling Cachay, Brian Henn, Oliver Watt-Meyer, Christopher S. Bretherton, and Rose Yu. Probablistic Emulation of a Global Climate Model with Spherical DYffusion. Advances in Neural Information Processing Systems, 37:127610–127644, December 2024
2024
-
[14]
ArchesWeather & ArchesWeatherGen: a deterministic and generative model for efficient ML weather forecasting, December 2024
Guillaume Couairon, Renu Singh, Anastase Charantonis, Christian Lessig, and Claire Mon- teleoni. ArchesWeather & ArchesWeatherGen: a deterministic and generative model for efficient ML weather forecasting, December 2024. arXiv:2412.12971 [cs]. 11
2024 arXiv
-
[15]
Generative Modeling by Estimating Gradients of the Data Distribution
Yang Song and Stefano Ermon. Generative Modeling by Estimating Gradients of the Data Distribution. In Advances in Neural Information Processing Systems , volume 32. Curran Associates, Inc., 2019
2019
-
[16]
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising Diffusion Probabilistic Models. In Advances in Neural Information Processing Systems , volume 33, pages 6840–6851. Curran Associates, Inc., 2020
2020
-
[17]
Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-Based Generative Modeling through Stochastic Differential Equations, February 2021. arXiv:2011.13456 [cs, stat]
2021 arXiv
-
[18]
Photorealistic Video Generation with Diffusion Models, December 2023
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, and José Lezama. Photorealistic Video Generation with Diffusion Models, December 2023. arXiv:2312.06662 [cs]
2023 arXiv
-
[19]
VideoComposer: Compositional Video Synthesis with Motion Controllability, June 2023
Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. VideoComposer: Compositional Video Synthesis with Motion Controllability, June 2023. arXiv:2306.02018 [cs]
2023 arXiv
-
[20]
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer, October 2024
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, Da Yin, Xiaotao Gu, Yuxuan Zhang, Weihan Wang, Yean Cheng, Ting Liu, Bin Xu, Yuxiao Dong, and Jie Tang. CogVideoX: Text-to-Video Diffusion Models ...
2024 arXiv
-
[21]
Autoregressive Video Generation without Vector Quantization, March 2025
Haoge Deng, Ting Pan, Haiwen Diao, Zhengxiong Luo, Yufeng Cui, Huchuan Lu, Shiguang Shan, Yonggang Qi, and Xinlong Wang. Autoregressive Video Generation without Vector Quantization, March 2025. arXiv:2412.14169 [cs]
2025 arXiv
-
[22]
HunyuanVideo: A Systematic Framework For Large Video Generative Models, January 2025
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, Kathrina Wu, Qin Lin, Junkun Yuan, Yanxin Long, Aladdin Wang, Andong Wang, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Hongmei Wang, Jacob Song, Jiawang Bai, J...
2025 arXiv
-
[23]
Efficiency-optimized Video Diffusion Models
Zijun Deng, Xiangteng He, and Yuxin Peng. Efficiency-optimized Video Diffusion Models. In Proceedings of the 31st ACM International Conference on Multimedia , MM ’23, pages 7295–7303, New York, NY , USA, October 2023. Association for Computing Machinery
2023
-
[24]
Diffusion Model-Based Video Editing: A Survey, June 2024
Wenhao Sun, Rong-Cheng Tu, Jingyi Liao, and Dacheng Tao. Diffusion Model-Based Video Editing: A Survey, June 2024. arXiv:2407.07111 [cs]
2024 arXiv
-
[25]
Pascal Chang, Jingwei Tang, Markus Gross, and Vinicius C. Azevedo. How I Warped Your Noise: a Temporally-Correlated Noise Prior for Diffusion Models. In The Twelfth International Conference on Learning Representations, October 2023
2023
-
[26]
Huang, and Niloy J
Duygu Ceylan, Chun-Hao P. Huang, and Niloy J. Mitra. Pix2Video: Video Editing using Image Diffusion. pages 23206–23217, 2023
2023
-
[27]
ControlVideo: Training-free Controllable Text-to-video Generation
Yabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang, Wangmeng Zuo, and Qi Tian. ControlVideo: Training-free Controllable Text-to-video Generation. October 2023
2023
-
[28]
FateZero: Fusing Attentions for Zero-shot Text-based Video Editing
Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. FateZero: Fusing Attentions for Zero-shot Text-based Video Editing. pages 15932–15942, 2023
2023
-
[29]
Debias Coarsely, Sample Conditionally: Statistical Downscaling through Optimal Transport and Probabilistic Diffusion Models, May 2023
Zhong Yi Wan, Ricardo Baptista, Yi-fan Chen, John Anderson, Anudhyan Boral, Fei Sha, and Leonardo Zepeda-Núñez. Debias Coarsely, Sample Conditionally: Statistical Downscaling through Optimal Transport and Probabilistic Diffusion Models, May 2023. arXiv:2305.15618 [physics]. 12
2023 arXiv
-
[30]
Unpaired Downscaling of Fluid Flows with Diffusion Bridges
Tobias Bischoff and Katherine Deck. Unpaired Downscaling of Fluid Flows with Diffusion Bridges. Artificial Intelligence for the Earth Systems, 3(2), May 2024. Publisher: American Meteorological Society Section: Artificial Intelligence for the Earth Systems
2024
-
[31]
Residual Diffusion Modeling for Km-scale Atmospheric Downscaling, January 2024
Morteza Mardani, Noah Brenowitz, Yair Cohen, Jaideep Pathak, Chieh-Yu Chen, Cheng-Chin Liu, Arash Vahdat, Karthik Kashinath, Jan Kautz, and Mike Pritchard. Residual Diffusion Modeling for Km-scale Atmospheric Downscaling, January 2024. ISSN: 2693-5015
2024
-
[32]
Étienne Plésiat, Robert J. H. Dunn, Markus G. Donat, and Christopher Kadow. Artificial intelligence reveals past climate extremes by reconstructing historical records. Nature Commu- nications, 15(1):9191, October 2024. Publisher: Nature Publishing Group
2024
-
[33]
Fast, scale-adaptive and uncertainty-aware downscaling of Earth system model fields with generative machine learning
Philipp Hess, Michael Aich, Baoxiang Pan, and Niklas Boers. Fast, scale-adaptive and uncertainty-aware downscaling of Earth system model fields with generative machine learning. Nature Machine Intelligence, 7(3):363–373, March 2025. Publisher: Nature Publishing Group
2025
-
[34]
Conditional diffusion models for downscaling & bias correction of Earth system model precipi- tation, April 2024
Michael Aich, Philipp Hess, Baoxiang Pan, Sebastian Bathiany, Yu Huang, and Niklas Boers. Conditional diffusion models for downscaling & bias correction of Earth system model precipi- tation, April 2024. arXiv:2404.14416 [physics]
2024
-
[35]
Behera, Dachao Jin, Baoxiang Pan, Huidong Jiang, and Toshio Yamagata
Fenghua Ling, Zeyu Lu, Jing-Jia Luo, Lei Bai, Swadhin K. Behera, Dachao Jin, Baoxiang Pan, Huidong Jiang, and Toshio Yamagata. Diffusion model-based probabilistic downscaling for 180-year East Asian climate reconstruction. npj Climate and Atmospheric Science, 7(1):1–11, June 2...
2024
-
[36]
Machine learning emulation of precipitation from km-scale regional climate simulations using a diffusion model, July 2024
Henry Addison, Elizabeth Kendon, Suman Ravuri, Laurence Aitchison, and Peter AG Watson. Machine learning emulation of precipitation from km-scale regional climate simulations using a diffusion model, July 2024. arXiv:2407.14158 [physics]
2024
-
[37]
Make-A-Video: Text-to-Video Generation without Text-Video Data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-A-Video: Text-to-Video Generation without Text-Video Data. September 2022
2022
-
[38]
Structure and Content-Guided Video Synthesis with Diffusion Models
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and Content-Guided Video Synthesis with Diffusion Models. pages 7346–7356, 2023
2023
-
[39]
Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation. pages 7623–7633, 2023
2023
-
[40]
Lumiere: A Space-Time Diffusion Model for Video Generation
Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Herrmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Amit Raj, Yuanzhen Li, Michael Rubinstein, Tomer Michaeli, Oliver Wang, Deqing Sun, Tali Dekel, and Inbar Mosseri. Lumiere: A Space-Time Diffusion Model for ...
2024
-
[41]
Adversarial Video Generation on Complex Datasets, September 2019
Aidan Clark, Jeff Donahue, and Karen Simonyan. Adversarial Video Generation on Complex Datasets, September 2019. arXiv:1907.06571 [cs]
2019 arXiv
-
[42]
Skilful precipitation nowcasting using deep generative models of radar
Suman Ravuri, Karel Lenc, Matthew Willson, Dmitry Kangin, Remi Lam, Piotr Mirowski, Megan Fitzsimons, Maria Athanassiadou, Sheleem Kashem, Sam Madge, Rachel Prudden, Amol Mandhane, Aidan Clark, Andrew Brock, Karen Simonyan, Raia Hadsell, Niall Robinson, Ellen Clancy, Alberto A...
2021 arXiv
-
[43]
Puja Das, August Posch, Nathan Barber, Michael Hicks, Kate Duffy, Thomas Vandal, Debjani Singh, Katie van Werkhoven, and Auroop R. Ganguly. Hybrid physics-AI outperforms numerical weather prediction for extreme precipitation nowcasting. npj Climate and Atmospheric Science, 7(1...
2024
-
[44]
Generating Videos with Scene Dy- namics
Carl V ondrick, Hamed Pirsiavash, and Antonio Torralba. Generating Videos with Scene Dy- namics. In Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc., 2016
2016
-
[45]
Deep multi-scale video prediction beyond mean square error, February 2016
Michael Mathieu, Camille Couprie, and Yann LeCun. Deep multi-scale video prediction beyond mean square error, February 2016. arXiv:1511.05440 [cs]
2016 arXiv
-
[46]
Temporal Generative Adversarial Nets With Singular Value Clipping
Masaki Saito, Eiichi Matsumoto, and Shunta Saito. Temporal Generative Adversarial Nets With Singular Value Clipping. In Proceedings of the IEEE International Conference on Computer Vision, pages 2830–2839, 2017
2017
-
[47]
MoCoGAN: Decomposing Motion and Content for Video Generation
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz. MoCoGAN: Decomposing Motion and Content for Video Generation. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1526–1535, Salt Lake City, UT, June 2018. IEEE
2018
-
[48]
Train Sparsely, Gen- erate Densely: Memory-Efficient Unsupervised Training of High-Resolution Temporal GAN
Masaki Saito, Shunta Saito, Masanori Koyama, and Sosuke Kobayashi. Train Sparsely, Gen- erate Densely: Memory-Efficient Unsupervised Training of High-Resolution Temporal GAN. International Journal of Computer Vision, 128(10):2586–2606, November 2020
2020
-
[49]
tempoGAN: a temporally coherent, volumetric GAN for super-resolution fluid flow
You Xie, Erik Franz, Mengyu Chu, and Nils Thuerey. tempoGAN: a temporally coherent, volumetric GAN for super-resolution fluid flow. ACM Trans. Graph., 37(4):95:1–95:15, July 2018
2018
-
[50]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. In M. Ranzato, A. Beygelzimer, Y . Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in neural information processing systems , volume 34, pages 8780–8794. Curran Associates, Inc., 2021
2021
-
[51]
Latent Video Diffusion Models for High-Fidelity Long Video Generation, March 2023
Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent Video Diffusion Models for High-Fidelity Long Video Generation, March 2023. arXiv:2211.13221 [cs]
2023 arXiv
-
[52]
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, November 2023
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Do- minik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, November ...
2023 arXiv
-
[53]
Latte: Latent Diffusion Transformer for Video Generation, January 2024
Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent Diffusion Transformer for Video Generation, January 2024. arXiv:2401.03048 [cs]
2024 arXiv
-
[54]
Pyramidal Flow Matching for Efficient Video Generative Modeling, October 2024
Yang Jin, Zhicheng Sun, Ningyuan Li, Kun Xu, Kun Xu, Hao Jiang, Nan Zhuang, Quzhe Huang, Yang Song, Yadong Mu, and Zhouchen Lin. Pyramidal Flow Matching for Efficient Video Generative Modeling, October 2024. arXiv:2410.05954 [cs]
2024
-
[55]
History-Guided Video Diffusion, February 2025
Kiwhan Song, Boyuan Chen, Max Simchowitz, Yilun Du, Russ Tedrake, and Vincent Sitzmann. History-Guided Video Diffusion, February 2025. arXiv:2502.06764 [cs]
2025 arXiv
-
[56]
Generative assimilation and prediction for weather and climate, March 2025
Shangshang Yang, Congyi Nai, Xinyan Liu, Weidong Li, Jie Chao, Jingnan Wang, Leyi Wang, Xichen Li, Xi Chen, Bo Lu, Ziniu Xiao, Niklas Boers, Huiling Yuan, and Baoxiang Pan. Generative assimilation and prediction for weather and climate, March 2025. arXiv:2503.03038 [cs]
2025 arXiv
-
[57]
Learning spatiotemporal dynamics with a pretrained generative model
Zeyu Li, Wang Han, Yue Zhang, Qingfei Fu, Jingxuan Li, Lizi Qin, Ruoyu Dong, Hao Sun, Yue Deng, and Lijun Yang. Learning spatiotemporal dynamics with a pretrained generative model. Nature Machine Intelligence, 6(12):1566–1579, December 2024. Publisher: Nature Publishing Group
2024
-
[58]
DiffObs: Generative Diffusion for Global Forecasting of Satellite Observations, April 2024
Jason Stock, Jaideep Pathak, Yair Cohen, Mike Pritchard, Piyush Garg, Dale Durran, Morteza Mardani, and Noah Brenowitz. DiffObs: Generative Diffusion for Global Forecasting of Satellite Observations, April 2024. arXiv:2404.06517
2024 arXiv
-
[59]
Andersson, Andrew El-Kadi, Do- minic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson
Ilan Price, Alvaro Sanchez-Gonzalez, Ferran Alet, Tom R. Andersson, Andrew El-Kadi, Do- minic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson. GenCast: Diffusion-based ensemble forecasting for medium-range weather, May 2024....
2024 arXiv
-
[60]
TokenFlow: Consistent Diffusion Features for Consistent Video Editing
Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. TokenFlow: Consistent Diffusion Features for Consistent Video Editing. In The Twelfth International Conference on Learning Representations, October 2023
2023
-
[61]
Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators
Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators. In Proceedings of the IEEE/CVF International Conference on Computer ...
2023
-
[62]
Score-based Data Assimilation for a Two-Layer Quasi- Geostrophic Model, November 2023
François Rozet and Gilles Louppe. Score-based Data Assimilation for a Two-Layer Quasi- Geostrophic Model, November 2023. arXiv:2310.01853 [stat]
2023 arXiv
-
[63]
Spa- tiotemporally Coherent Probabilistic Generation of Weather from Climate, January 2025
Jonathan Schmidt, Luca Schmidt, Felix Strnad, Nicole Ludwig, and Philipp Hennig. Spa- tiotemporally Coherent Probabilistic Generation of Weather from Climate, January 2025. arXiv:2412.15361 [cs]
2025
-
[64]
Weather Prediction with Diffusion Guided by Realistic Forecast Processes, February 2024
Zhanxiang Hua, Yutong He, Chengqian Ma, and Alexandra Anderson-Frey. Weather Prediction with Diffusion Guided by Realistic Forecast Processes, February 2024. arXiv:2402.06666 [physics]
2024 arXiv
-
[65]
Dueben, and Torsten Hoefler
Langwen Huang, Lukas Gianinazzi, Yuejiang Yu, Peter D. Dueben, and Torsten Hoefler. DiffDA: a diffusion model for weather-scale data assimilation. In Proceedings of the 41st International Conference on Machine Learning , volume 235 of ICML’24, pages 19798–19815, Vienna, Austri...
2024
-
[66]
Align Your Latents: High-Resolution Video Synthesis With Latent Diffusion Models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align Your Latents: High-Resolution Video Synthesis With Latent Diffusion Models. pages 22563–22575, 2023
2023
-
[67]
Diffusion-GAN: Training GANs with Diffusion
Zhendong Wang, Huangjie Zheng, Pengcheng He, Weizhu Chen, and Mingyuan Zhou. Diffusion-GAN: Training GANs with Diffusion. September 2022
2022
-
[68]
DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination Capability
Runhui Huang, Jianhua Han, Guansong Lu, Xiaodan Liang, Yihan Zeng, Wei Zhang, and Hang Xu. DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination Capability. pages 15713–15723, 2023
2023
-
[69]
Adversarial Training of Denoising Diffusion Model Using Dual Discriminators for High-Fidelity Multi-Speaker TTS
Myeongjin Ko, Euiyeon Kim, and Yong-Hoon Choi. Adversarial Training of Denoising Diffusion Model Using Dual Discriminators for High-Fidelity Multi-Speaker TTS. IEEE Open Journal of Signal Processing, 5:577–587, 2024. Conference Name: IEEE Open Journal of Signal Processing
2024
-
[70]
Adversarial Diffusion Distillation
Axel Sauer, Dominik Lorenz, Andreas Blattmann, and Robin Rombach. Adversarial Diffusion Distillation. In Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol, editors, Computer Vision – ECCV 2024, pages 87–103, Cham, 2024. Springer Nature ...
2024
-
[71]
Improved Distribution Matching Distillation for Fast Image Synthesis
Tianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang, Eli Shechtman, Fredo Durand, and Bill Freeman. Improved Distribution Matching Distillation for Fast Image Synthesis. Advances in Neural Information Processing Systems, 37:47455–47487, December 2024
2024
-
[72]
Refining Generative Process with Discriminator Guidance in Score-based Diffusion Models, June 2023
Dongjun Kim, Yeongmin Kim, Se Jung Kwon, Wanmo Kang, and Il-Chul Moon. Refining Generative Process with Discriminator Guidance in Score-based Diffusion Models, June 2023. arXiv:2211.17091 [cs]
2023 arXiv
-
[73]
Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation, May 2024
Akio Hayakawa, Masato Ishii, Takashi Shibuya, and Yuki Mitsufuji. Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation, May 2024. arXiv:2405.17842
2024 arXiv
-
[74]
Discriminator Guidance for Autoregressive Diffusion Models
Filip Ekström Kelvinius and Fredrik Lindsten. Discriminator Guidance for Autoregressive Diffusion Models. In Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, pages 3403–3411. PMLR, April 2024. ISSN: 2640-3498
2024
-
[75]
Kerby and Kevin R
Thomas J. Kerby and Kevin R. Moon. Training-Free Guidance for Discrete Diffusion Models for Molecular Generation, September 2024. arXiv:2409.07359 [stat]. 15
2024 arXiv
-
[76]
Brian D. O. Anderson. Reverse-time diffusion equation models. Stochastic Processes and their Applications, 12(3):313–326, May 1982
1982
-
[77]
Elucidating the Design Space of Diffusion-Based Generative Models, October 2022
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the Design Space of Diffusion-Based Generative Models, October 2022. arXiv:2206.00364 [cs, stat]
2022 arXiv
-
[78]
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Nets. In Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K. Q. Weinberger, editors, Advances in Neural Information Pro...
2014
-
[79]
Hans Hersbach, Bill Bell, Paul Berrisford, Shoji Hirahara, András Horányi, Joaquín Muñoz- Sabater, Julien Nicolas, Carole Peubey, Raluca Radu, Dinand Schepers, Adrian Simmons, Cornel Soci, Saleh Abdalla, Xavier Abellan, Gianpaolo Balsamo, Peter Bechtold, Gionata Biavati, Jean ...
1999
-
[80]
The Trough-and-Ridge diagram
Ernest Hovmöller. The Trough-and-Ridge diagram. Tellus, 1(2):62–66, 1949. _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.2153-3490.1949.tb01260.x
1949
-
[81]
Tilmann Gneiting, Fadoua Balabdaoui, and Adrian E. Raftery. Probabilistic forecasts, calibra- tion and sharpness. Journal of the Royal Statistical Society: Series B (Statistical Methodol- ogy), 69(2):243–268, 2007. _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.146...
2007
-
[82]
Sperber, Pe- ter J
Min-Seop Ahn, Daehyun Kim, Daehyun Kang, Jiwoo Lee, Kenneth R. Sperber, Pe- ter J. Gleckler, Xianan Jiang, Yoo-Geun Ham, and Hyemi Kim. MJO Propa- gation Across the Maritime Continent: Are CMIP6 Models Better Than CMIP5 Models? Geophysical Research Letters , 47(11):e2020GL0872...
2020 doi
-
[83]
Qiang Sun, and Pedram Hassanzadeh
Ashesh Chattopadhyay, Y . Qiang Sun, and Pedram Hassanzadeh. Challenges of learning multi- scale dynamics with AI weather models: Implications for stability and one solution, December
-
[84]
Accurate calculation of spherical and vector spherical harmonic expansions via spectral element grids
Bo Wang, Li-Lian Wang, and Ziqing Xie. Accurate calculation of spherical and vector spherical harmonic expansions via spectral element grids. Advances in Computational Mathematics , 44(3):951–985, June 2018
2018
-
[85]
arXiv:2304.07029 [physics]
-
[86]
Fortin, M
V . Fortin, M. Abaza, F. Anctil, and R. Turcotte. Why Should Ensemble Spread Match the RMSE of the Ensemble Mean? Journal of Hydrometeorology, 15(4):1708–1713, August 2014. Publisher: American Meteorological Society Section: Journal of Hydrometeorology. 16 A Diffusion models F...
2014
-
[87]
Constantinou, Gregory LeClaire Wagner, Lia Siegelman, Brodie C
Navid C. Constantinou, Gregory LeClaire Wagner, Lia Siegelman, Brodie C. Pearson, and André Palóczy. GeophysicalFlows.jl: Solvers for geophysical fluid dynamics problems in periodic domains on CPUs & GPUs. Journal of Open Source Software, 6(60):3053, April 2021
2021
-
[2024]
arXiv:2303.17599 [cs]
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.