REVIEW 3 major objections 5 minor 2 cited by
FlexCache: Flexible Approximate Cache System for Video Diffusion
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read FlexCache claims approximate caching can be made practical for text-to-video diffusion, reporting 1.26x throughput and 25% lower cost than a video-adapted image-cache baseline while keeping quality within the paper's stated imperceptible…
desk verdict Genuine new techniques for video-diffusion caching, but the decoupled stitching rests on an unvalidated assumption about object positions across denoising steps—worth a serious referee, not a desk reject. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the compressed, decoupled latent-state cache entry. Intra-step compression keeps only key frames within each cached denoising step and records a map that repeats those key frames on decompression. Inter-step compression stores the differential frames from one base step and reproduces other steps by multiplying with a per-step least-squares coefficient $\alpha_s$, saving the coefficient as a single scalar plus any extra frames unique to a step. For lookup, the system keeps three embedding indexes for the whole prompt, the object description, and the background description, and it chooses whichever yields higher similarity: the whole prompt or the minimum of the object and background scores. To stitch two cached videos, it uses per-frame, 1-bit segmentation masks produced by segmenting the final output video, then applies those masks to the earlier latent states; this transfer relies on the claim that object positions do not change as denoising proceeds. Cache eviction uses a priority of $f_i \times \text{step}_i / (\text{capacity}_i \times \text{duration}_i)$, where $f_i$ is access frequency, $\text{step}_i$ is the steps saved by the entry, $\text{capacity}_i$ is its size, and $\text{duration}_i$ is time since last access.
What would settle it
Run FlexCache on prompts whose objects visibly move or change shape during generation, compare the stitched latent states against the no-cache latent states at the same step, and check whether misalignment appears or the FVD difference rises above 50; if it does, the position-invariance assumption fails.
Extended reading notes
Core claim
On its own terms, the central discovery is that there is enough exploitable structure in the latent states of a video diffusion model to make approximate caching economical: frames inside a step are near-duplicates, and the differential values between key frames are almost identical across the first 25 denoising steps, so a cache entry can be reduced roughly 6.7x with almost no quality loss. Because a prompt's usable content splits into object and background, looking those parts up separately yields a 13.8 percentage-point hit-rate gain over whole-prompt lookup, and the two matched latent states can be stitched by transferring object and background masks from the final video onto early latent states. With a 1000 GB cache, the resulting system reports 1.26x the throughput of a video-adapted image-cache baseline at 25% lower cost, with FVD staying within the 50-point band the paper treats as imperceptible.
Load-bearing premise
The whole scheme relies on the assumption that objects in a generated video hold their positions throughout denoising, so a mask cut from the final video can be stamped onto early noisy latent states; if objects drift, the stitching misaligns and quality degrades.
Editorial extensions
If this is right
- With compression, a 1000 GB cache holds roughly 6.7x more video latent states, which is what makes a viable hit rate possible under a fixed storage budget.
- A service using an image-style cache baseline could switch to FlexCache and serve 1.26x the requests per hour at 25% lower per-video cost, with FVD within 50 of no-cache generation.
- Decoupled object and background lookup raises the hit rate by 13.8 percentage points over whole-prompt lookup, and the most common reuse in the evaluation is skipping 10 of 50 denoising steps.
- LRBU replacement yields higher computation savings than FIFO, LRU, and the image-cache policy at every tested cache size from 1 GB to 1000 GB.
- FlexCache inserts new cache entries even on hits before step 25, filling later cached steps, and falls back to an earlier cached step when the desired step has been evicted.
Reading between the lines
- Beyond the paper: object and background decoupling should help most in long-tail workloads where exact whole-prompt reuse is rare, so the reported 13.8-point hit-rate gain may widen on more diverse prompt streams.
- Beyond the paper: the inter-step differential sharing suggests a general principle—if two denoising steps differ by a near-linear factor, one cached tensor plus a scalar can stand for both—which could be reused by other iterative generative models that share temporal structure.
- Beyond the paper: if object positions do drift in some prompts, re-segmenting early latent states or estimating per-step masks with a motion model would extend the stitching without changing the compression scheme.
- Beyond the paper: the cost claim depends on storage pricing; under cheaper storage the compression advantage shifts from cost to hit rate and could be traded for more cached prompts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. FlexCache is an approximate caching system for text-to-video diffusion models. It compresses cached latent states through intra-step key-frame removal and inter-step differential scaling (Eq. 2), decouples object/background lookups and stitches the retrieved latents, and introduces an LRBU replacement policy that accounts for per-entry capacity and recency. The system is evaluated on VideoCrafter2 with 700k prompts from the VidProM trace, reporting 1.26x higher throughput than a video-adapted NIRVANA baseline, 25% lower cost, 6.7x cache-size reduction, a 98.4% hit rate, and an FVD difference of 24 relative to no caching. The main risk is that the decoupled stitching path assumes object positions are fixed across denoising steps; this assumption is asserted from a qualitative example and is not demonstrated for the workload, and the reported quality metrics do not isolate the decoupled-stitched generations.
Significance. If the reported numbers hold, FlexCache is a useful systems contribution: it is the first approximate caching system for text-to-video diffusion, it is evaluated on a large real-world prompt trace, and its least-squares compression coefficient in Eq. (2) is internally correct as a lossy compression scheme. The paper also makes a concrete cost model (Figure 22) and separates the effects of caching, compression, and hit-rate optimization (Figure 17). The needed work is not a new theoretical derivation but a validation that decoupled object/background stitching produces latents that are actually usable as starting points; without that validation, the headline hit-rate and throughput gains are not fully supported.
major comments (3)
- [§4.2.2] The decoupled stitching path depends on the claim that 'the positions of these objects are never changed as the denoising steps increase,' which is used to justify applying segmentation masks from the final video to early latent states. The only support is the qualitative example in Figure 2b, while Figure 6 measures similarity of inter-frame differential values, which is a different quantity from spatial alignment of object masks in early noisy latents. This is load-bearing because decoupled hits must produce valid starting latents for the diffusion model; if masks misalign, the stitched latent injects mismatched noise that the remaining steps may not repair. Please provide a quantitative evaluation on VidProM, for example by measuring mask-conditional latent similarity at steps 5-15 or by reporting FVD/CLIP metrics separately for whole-prompt hits, decoupled-stitched hits, and misses instead of only the Table 2 averages. Without such evidence, a failure of this assumption would reduce the effective hit rate toward the whole-prompt-only rate and shrink the reported 1.26x throughput / 25% cost advantage.
- [§3.1.2, Figure 7] The claimed 13.8% higher hit rate from decoupled lookup is derived from marginal similarity distributions for object and background prompts separately, but Section 3.2 defines a decoupled hit as requiring both the object and the background to pass the similarity threshold. The combined decoupled hit rate is the intersection of these two marginal events, not their maximum or sum, so Figure 7 does not establish the claimed gain. Please report the combined object+background hit rate under the same cache conditions and the fraction of the 98.4% hit rate in Figure 18 that comes from the decoupled path. In addition, the 0.65 similarity threshold is taken from NIRVANA without recalibration for video prompts; a sensitivity analysis over this threshold is needed to confirm that the comparison is not an artifact of the chosen cutoff.
- [§6.2.5, Table 2] The quality evidence for the central 'almost no degradation' claim is thin. Table 2 reports only point estimates: FVD 192 for FlexCache versus 168 for No Cache, with no confidence intervals or error bars, and the text's assertion that an FVD difference below 50 is imperceptible is not justified for this experimental setup. With 2k evaluated videos, the authors can report standard errors or confidence intervals, and they should also provide per-mode quality numbers (whole-prompt hit, decoupled stitch, miss) rather than averaging over all requests. Without this decomposition, the quality of the decoupled-stitched generations, which is the riskiest path, can be masked by the more common whole-prompt hits.
minor comments (5)
- [Abstract / §1] The abstract says 'FlexCache reach' and should say 'reaches'; the throughput improvement is reported as '1.26x' in the abstract and conclusion and '26%' in the introduction, which should be harmonized.
- [References] References [38] and [39] are the same FreeNoise paper, and [40], [41], and [42] are all the same CLIP paper; these duplicates should be merged and cited once.
- [§4.1.2] The notation around Eq. (1) and Eq. (2) is under-specified: the indices i,j are introduced awkwardly ('We further denotei,j'), and the sums in Eq. (2) should be written with explicit bounds or a clear convention for all key frames and pixels.
- [§6.2.5] The FVD reference set and preprocessing are not described; because FVD depends on the reference video distribution and the splitting approach, these details should be reported for reproducibility.
- [Figure 18] The stacked bars for 'Hit whole prompt' and 'Hit decoupled' are not quantified in the text except for the overall 98.4% hit rate; the figure should be accompanied by the per-mode percentages.
Circularity Check
No significant circularity: headline results are measured against external baselines and workloads, not derived from the paper's own definitions or self-citations.
full rationale
FlexCache's central claims (6.7x compression, 13.8% higher hit rate, 1.26x throughput, 25% cost savings) are evaluation results obtained by running VideoCrafter2 on the VidProM workload against a NIRVANA-derived baseline, not consequences of definitions. The inter-step compression coefficient alpha_s (Eq. 2) is a per-entry least-squares parameter fitted to each latent state; although the reported reconstruction similarity is in-sample for that fit, the paper separately validates end-to-end quality with FVD/CLIP metrics on generated videos, so the compression benefit is not a fitted quantity renamed as a prediction. The decoupled lookup hit-rate gain is an empirical comparison of whole-prompt vs object/background similarity thresholds, and the 0.65 threshold is inherited from NIRVANA rather than chosen to force the result. The LRBU replacement policy was designed after observing VidProM recency trends and is evaluated on VidProM, which is a potential workload-overfitting concern but not a derivation-level circularity. There are no load-bearing self-citations or imported uniqueness theorems; the Section 4.2.2 assumption that object positions are fixed across denoising steps is an unverified premise affecting quality, not a circular step. Accordingly, no circular step can be quoted and exhibited, and the paper merits a 0 circularity score.
Assumptions & free parameters
free parameters (4)
- compression similarity threshold =
0.99
- inter-step coefficient alpha_s =
per-entry, per-step least-squares fit
- similarity threshold for cache hit =
0.65
- cached step positions =
steps 5, 10, 15, 20, 25
assumptions (4)
- domain assumption Object positions do not change across denoising steps.
- domain assumption Differential values between key frames are nearly identical across the first 25 denoising steps.
- domain assumption VidProM prompt trace is representative of real text-to-video serving workloads.
- standard math Least-squares minimization is solved by Equation 2.
Cite this review
Pith. "Pith review of FlexCache: Flexible Approximate Cache System for Video Diffusion." pith.science (2026). https://pith.science/paper/2YI752IC
@misc{pith2026250104012,
author = {Pith},
title = {Pith review of: FlexCache: Flexible Approximate Cache System for Video Diffusion},
year = {2026},
howpublished = {\url{https://pith.science/paper/2YI752IC}},
note = {Machine review of arXiv:2501.04012}
}
read the original abstract
Text-to-Video applications receive increasing attention from the public. Among these, diffusion models have emerged as the most prominent approach, offering impressive quality in visual content generation. However, it still suffers from substantial computational complexity, often requiring several minutes to generate a single video. While prior research has addressed the computational overhead in text-to-image diffusion models, the techniques developed are not directly suitable for video diffusion models due to the significantly larger cache requirements and enhanced computational demands associated with video generation. We present FlexCache, a flexible approximate cache system that addresses the challenges in two main designs. First, we compress the caches before saving them to storage. Our compression strategy can reduce 6.7 times consumption on average. Then we find that the approximate cache system can achieve higher hit rate and computation savings by decoupling the object and background. We further design a tailored cache replacement policy to support the two techniques mentioned above better. Through our evaluation, FlexCache reaches 1.26 times higher throughput and 25% lower cost compared to the state-of-the-art diffusion approximate cache system.
Figures
Figures from the paper (10 more)
Forward citations
Cited by 2 Pith papers
-
DSTAR: Accelerating Diffusion Transformers via Spatial and Temporal Redundancy Reduction
DSTAR reports 7.33x latency speedup and 41.89x energy savings over an A100 GPU on seven diffusion transformers by quantizing differential activations to as few as 2 bits and reusing block-wise sparse attention scores.
-
ReFrame: Layer Caching for Accelerated Inference in Real-Time Rendering
Caching deep encoder features across frames, with a SMAPE-threshold refresh policy, yields about 1.4x average inference speedup on three real-time rendering networks with small perceptual loss at the high-sensitivity setting.
Reference graph
Works this paper leans on
-
[1]
https://github.com/ luca-medeiros/lang-segment-anything , 2024
Language segment-anything. https://github.com/ luca-medeiros/lang-segment-anything , 2024
work page 2024
-
[2]
Create with adobe firefly generative AI
Adobe. Create with adobe firefly generative AI. https: //www.adobe.com/products/firefly.html, 2023
work page 2023
-
[3]
Approximate caching for efficiently serving Text-to-Image diffusion models
Shubham Agarwal, Subrata Mitra, Sarthak Chakraborty, Srikrishna Karanam, Koyel Mukherjee, and Shiv Ku- mar Saini. Approximate caching for efficiently serving Text-to-Image diffusion models. In21st USENIX Sympo- sium on Networked Systems Design and Implementation (NSDI 24), pages 1173–1189, Santa Clara, CA, April
-
[4]
Yongqi An, Xu Zhao, Tao Yu, Haiyun Gu, Chaoyang Zhao, Ming Tang, and Jinqiao Wang. ZBS: Zero-shot background subtraction via instance-level background modeling and foreground selection. In IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), pages 6355–6364, 2023
work page 2023
-
[5]
Stable video dif- fusion: Scaling latent video diffusion models to large datasets, 2023
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video dif- fusion: Scaling latent video diffusion models to large datasets, 2023
work page 2023
-
[6]
Token merging for fast stable diffusion
Daniel Bolya and Judy Hoffman. Token merging for fast stable diffusion. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 4599–4603, June 2023
work page 2023
-
[7]
VideoCrafter2: Overcoming data limitations for high- quality video diffusion models
Haoxin Chen, Yong Zhang, Xiaodong Cun, Meng- han Xia, Xintao Wang, Chao Weng, and Ying Shan. VideoCrafter2: Overcoming data limitations for high- quality video diffusion models. In IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), pages 7310–7320, 2024
work page 2024
-
[8]
Gentron: Diffu- sion transformers for image and video generation
Shoufa Chen, Mengmeng Xu, Jiawei Ren, Yuren Cong, Sen He, Yanping Xie, Animesh Sinha, Ping Luo, Tao Xiang, and Juan-Manuel Perez-Rua. Gentron: Diffu- sion transformers for image and video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6441– 6451, June 2024
work page 2024
Show all 61 references
-
[9]
Textual grounding for open-vocabulary visual information extraction in layout-diversified documents
Mengjun Cheng, Chengquan Zhang, Chang Liu, Yuke Li, Bohan Li, Kun Yao, Xiawu Zheng, Rongrong Ji, and Jie Chen. Textual grounding for open-vocabulary visual information extraction in layout-diversified documents. In Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Rus- sakovsky, ...
2025
-
[10]
Compute engine pricing
Google Cloud. Compute engine pricing. https:// cloud.google.com/compute/all-pricing, 2024
2024
-
[11]
FlashAttention-2: Faster attention with bet- ter parallelism and work partitioning
Tri Dao. FlashAttention-2: Faster attention with bet- ter parallelism and work partitioning. In International Conference on Learning Representations (ICLR), 2024
2024
-
[12]
Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Tri Dao, Daniel Y . Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FlashAttention: Fast and memory- efficient exact attention with IO-awareness. InAdvances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[13]
Veo: Our most capable generative video model
DeepMind. Veo: Our most capable generative video model. https://deepmind.google/technologies/ veo, 2024
2024
-
[14]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. In M. Ranzato, A. Beygelzimer, Y . Dauphin, P.S. Liang, and J. Wort- man Vaughan, editors, Advances in Neural Information Processing Systems (NeurIPS), volume 34, pages 8780–
-
[15]
Neural inter-frame compression for video coding
Abdelaziz Djelouah, Joaquim Campos, Simone Schaub- Meyer, and Christopher Schroers. Neural inter-frame compression for video coding. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), October 2019
2019
-
[16]
xdit: an inference engine for diffusion transform- ers (dits) with massive parallelism, 2024
Jiarui Fang, Jinzhe Pan, Xibo Sun, Aoyu Li, and Jiannan Wang. xdit: an inference engine for diffusion transform- ers (dits) with massive parallelism, 2024
2024
-
[17]
Dysen-VDM: Empowering dynamics- aware text-to-video diffusion with llms
Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, and Tat-Seng Chua. Dysen-VDM: Empowering dynamics- aware text-to-video diffusion with llms. In IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 7641–7653, 2024
2024
-
[18]
Prompt cache: Modular attention reuse for low-latency inference
In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong. Prompt cache: Modular attention reuse for low-latency inference. In P. Gibbons, G. Pekhimenko, and C. De Sa, editors, Pro- ceedings of Machine Learning and Systems (MLSys) , volume 6, pages 32...
2024
-
[19]
Denois- ing diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denois- ing diffusion probabilistic models. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, edi- tors, Advances in Neural Information Processing Sys- tems (NeurIPS), volume 33, pages 6840–6851. Curran Associates, In...
2020
-
[20]
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Rep- resentations (ICLR), 2022
2022
-
[21]
VMC: Video motion customization using temporal at- tention adaption for text-to-video diffusion models
Hyeonho Jeong, Geon Yeong Park, and Jong Chul Ye. VMC: Video motion customization using temporal at- tention adaption for text-to-video diffusion models. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9212–9221, 2024
2024
-
[22]
Ground-a-video: Zero-shot grounded video editing using text-to-image diffusion models
Hyeonho Jeong and Jong Chul Ye. Ground-a-video: Zero-shot grounded video editing using text-to-image diffusion models. In The Twelfth International Confer- ence on Learning Representations (ICLR), 2024
2024
-
[23]
VideoBooth: Diffusion-based video generation with im- age prompts
Yuming Jiang, Tianxing Wu, Shuai Yang, Chenyang Si, Dahua Lin, Yu Qiao, Chen Change Loy, and Ziwei Liu. VideoBooth: Diffusion-based video generation with im- age prompts. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6689– 6700, 2024
2024
-
[24]
Kling ai: Next-generation ai creative studio
KLing. Kling ai: Next-generation ai creative studio. https://klingai.com/, 2024
2024
-
[25]
Cambricon- d: Full-network differential acceleration for diffusion models
Weihao Kong, Yifan Hao, Qi Guo, Yongwei Zhao, Xinkai Song, Xiaqing Li, Mo Zou, Zidong Du, Rui Zhang, Chang Liu, Yuanbo Wen, Pengwei Jin, Xing Hu, Wei Li, Zhiwei Xu, and Tianshi Chen. Cambricon- d: Full-network differential acceleration for diffusion models. In 2024 ACM/IEEE 51...
2024
-
[26]
xformers: A modular and hackable transformer modelling library
Benjamin Lefaudeux, Francisco Massa, Diana Liskovich, Wenhan Xiong, Vittorio Caggiano, Sean Naren, Min Xu, Jieru Hu, Marta Tintore, Susan Zhang, Patrick Labatut, Daniel Haziza, Luca Wehrstedt, Jeremy Reizenstein, and Grigory Sizov. xformers: A modular and hackable transformer ...
2022
-
[27]
Manmatha, Ashwin Swaminathan, Zhuowen Tu, Stefano Ermon, and Stefano Soatto
Hao Li, Yang Zou, Ying Wang, Orchid Majumder, Yusheng Xie, R. Manmatha, Ashwin Swaminathan, Zhuowen Tu, Stefano Ermon, and Stefano Soatto. On the scalability of diffusion-based text-to-image generation. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CV...
2024
-
[28]
Distrifusion: Distributed parallel inference for high-resolution diffusion models
Muyang Li, Tianle Cai, Jiaxin Cao, Qinsheng Zhang, Han Cai, Junjie Bai, Yangqing Jia, Kai Li, and Song Han. Distrifusion: Distributed parallel inference for high-resolution diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (...
2024
-
[29]
Swiftdiffusion: Efficient diffusion model serving with add-on modules, 2024
Suyi Li, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu, Zhipeng Di, Weiyi Lu, Jiawei Chen, Kan Liu, Yinghao Yu, Tao Lan, Guodong Yang, Lin Qu, Liping Zhang, and Wei Wang. Swiftdiffusion: Efficient diffusion model serving with add-on modules, 2024
2024
-
[30]
DPM-Solver++: Fast solver for guided sampling of diffusion probabilistic models
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongx- uan Li, and Jun Zhu. DPM-Solver++: Fast solver for guided sampling of diffusion probabilistic models. In The Eleventh International Conference on Learning Rep- resentations (ICLR), 2023
2023
-
[31]
Learning-to-cache: Accelerating diffusion transformer via layer caching, 2024
Xinyin Ma, Gongfan Fang, Michael Bi Mi, and Xin- chao Wang. Learning-to-cache: Accelerating diffusion transformer via layer caching, 2024
2024
-
[32]
Deep- cache: Accelerating diffusion models for free
Xinyin Ma, Gongfan Fang, and Xinchao Wang. Deep- cache: Accelerating diffusion models for free. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15762–15772, 2024
2024
-
[33]
Video generation models as world simulators
OpenAI. Video generation models as world simulators. https://openai.com/index/ video-generation-models-as-world-simulators , 2024
2024
-
[34]
PyTorch: an imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fan...
2019
-
[35]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 4172– 4182, 2023
2023
-
[36]
SDXL: Improving latent diffu- sion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. SDXL: Improving latent diffu- sion models for high-resolution image synthesis. In The Twelfth International Conference on Learning Represen- tations (ICLR), 2024
2024
-
[37]
vector database
Qdrant. vector database. https://qdrant.tech/, 2023
2023
-
[38]
Freenoise: Tuning-free longer video diffusion via noise reschedul- ing
Haonan Qiu, Menghan Xia, Yong Zhang, Yingqing He, Xintao Wang, Ying Shan, and Ziwei Liu. Freenoise: Tuning-free longer video diffusion via noise reschedul- ing. In The Twelfth International Conference on Learn- ing Representations (ICLR), 2024. 14
2024
-
[39]
Freenoise: Tuning-free longer video diffusion via noise reschedul- ing
Haonan Qiu, Menghan Xia, Yong Zhang, Yingqing He, Xintao Wang, Ying Shan, and Ziwei Liu. Freenoise: Tuning-free longer video diffusion via noise reschedul- ing. In The Twelfth International Conference on Learn- ing Representations (ICLR), 2024
2024
-
[42]
Learning transferable vi- sual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable vi- sual models from natural language supervision. In Marina Meila an...
2021
-
[43]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 10674–10685, 2022
2022
-
[44]
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention (MICCAI), pages 234–
-
[45]
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-a-video: Text-to-video generation without text-video data. In The Eleventh International Conference on Lea...
2023
-
[46]
De- noising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. De- noising diffusion implicit models. In International Con- ference on Learning Representations (ICLR), 2021
2021
-
[47]
Umie: Unified multimodal information extraction with instruction tuning
Lin Sun, Kai Zhang, Qingyuan Li, and Renze Lou. Umie: Unified multimodal information extraction with instruction tuning. In Proceedings of the AAAI Confer- ence on Artificial Intelligence (AAAI), 2024
2024
-
[48]
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Al- bert, Amjad Almahairi, Yasmine Babaei, Nikolay Bash- lykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernan- des, Jeremy...
2023
-
[49]
Video description with spatial-temporal attention
Yunbin Tu, Xishan Zhang, Bingtao Liu, and Chenggang Yan. Video description with spatial-temporal attention. In Proceedings of the 25th ACM International Confer- ence on Multimedia (ACM Multimedia), MM ’17, page 1014–1022, New York, NY , USA, 2017. Association for Computing Machinery
2017
-
[50]
FVD: A new metric for video generation
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Ku- rach, Raphaël Marinier, Marcin Michalski, and Sylvain Gelly. FVD: A new metric for video generation. In International Conference on Learning Representations (ICLR), 2019
2019
-
[51]
ZoLA: Zero-shot creative long ani- mation generation with short video model
Fu-Yun Wang, Zhaoyang Huang, Qiang Ma, Guanglu Song, Xudong Lu, Weikang Bian, Yijin Li, Yu Liu, and Hongsheng Li. ZoLA: Zero-shot creative long ani- mation generation with short video model. In Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gü...
2025
-
[52]
Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models, 2024
Wenhao Wang and Yi Yang. Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models, 2024. 15
2024
-
[53]
Videofac- tory: Swap attention in spatiotemporal diffusions for text-to-video generation
Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu. Videofac- tory: Swap attention in spatiotemporal diffusions for text-to-video generation. In The Twelfth International Conference on Learning Representations (ICLR), 2024
2024
-
[54]
Wang, Evan Montoya, David Munechika, Haoyang Yang, Benjamin Hoover, and Duen Horng Chau
Zijie J. Wang, Evan Montoya, David Munechika, Haoyang Yang, Benjamin Hoover, and Duen Horng Chau. DiffusionDB: A large-scale prompt gallery dataset for text-to-image generative models. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, ed- itors, Proceedings of the 61st A...
2023
-
[55]
Cache me if you can: Accelerating diffusion mod- els through block caching
Felix Wimbauer, Bichen Wu, Edgar Schoenfeld, Xi- aoliang Dai, Ji Hou, Zijian He, Artsiom Sanakoyeu, Peizhao Zhang, Sam Tsai, Jonas Kohler, Christian Rup- precht, Daniel Cremers, Peter Vajda, and Jialiang Wang. Cache me if you can: Accelerating diffusion mod- els through block ...
2024
-
[56]
Video compression through image interpolation
Chao-Yuan Wu, Nayan Singhal, and Philipp Krahenbuhl. Video compression through image interpolation. In Proceedings of the European Conference on Computer Vision (ECCV), September 2018
2018
-
[57]
Towards a better metric for text-to-video generation, 2024
Jay Zhangjie Wu, Guian Fang, Haoning Wu, Xintao Wang, Yixiao Ge, Xiaodong Cun, David Junhao Zhang, Jia-Wei Liu, Yuchao Gu, Rui Zhao, Weisi Lin, Wynne Hsu, Ying Shan, and Mike Zheng Shou. Towards a better metric for text-to-video generation, 2024
2024
-
[58]
DynamiCrafter: An- imating open-domain images with video diffusion pri- ors
Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang, Ying Shan, and Tien-Tsin Wong. DynamiCrafter: An- imating open-domain images with video diffusion pri- ors. In Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten...
2025
-
[59]
Alt- diffusion: A multilingual text-to-image diffusion model
Fulong Ye, Guang Liu, Xinya Wu, and Ledell Wu. Alt- diffusion: A multilingual text-to-image diffusion model. Proceedings of the AAAI Conference on Artificial Intel- ligence (AAAI), 38(7):6648–6656, Mar. 2024
2024
-
[60]
Accelerating text-to-image editing via cache-enabled sparse diffusion inference
Zihao Yu, Haoyang Li, Fangcheng Fu, Xupeng Miao, and Bin Cui. Accelerating text-to-image editing via cache-enabled sparse diffusion inference. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 38(15):16605–16613, Mar. 2024
2024
-
[61]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 3813–3824, 2023
2023
-
[62]
Hidiffusion: Unlocking higher-resolution creativity and efficiency in pretrained diffusion models
Shen Zhang, Zhaowei Chen, Zhenyu Zhao, Yuhao Chen, Yao Tang, and Jiajun Liang. Hidiffusion: Unlocking higher-resolution creativity and efficiency in pretrained diffusion models. In Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol, edit...
2025
-
[8794]
Curran Associates, Inc., 2021
2021
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.