REVIEW 4 major objections 5 minor 73 references
LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Identity-preserving video generation is improved by routing local facial structure into a diffusion transformer and correcting denoised latent tokens with chunk-wise autoregressive biases.
desk verdict The local router is a plausible contribution, but the temporal autoregressive module has no training objective in the paper's own equations, so its claimed gains are unsupported as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's argument is carried by two mechanisms. The local router is a learned weighting layer that routes latent tokens through local facial component tokens, expressing global face identity as a weighted combination of eyebrows, eyes, mouth, nose, skin, and hair; it is supervised by a segmentation-mask cross-entropy loss. The temporal autoregressive module is a post-denoising bias predictor: it groups latent tokens into temporal chunks and applies $c_k^* = c_k + \beta b_k$, with $b_k$ produced by cross-attention over the previously enhanced chunk; causal self-attention with rotary position embeddings and teacher forcing keep the refinement temporally ordered. Their work is to add spatial
What would settle it
Sweep $\beta$ in Eq. (4) from 0 to 1 on the same 90-prompt benchmark while watching FaceSim-Curricular and FID together, and rerun the ablation with the route loss removed. If face-identity scores keep climbing as FID worsens, or if identity gains disappear when the segmentation supervision is dropped, the reported advantage depends on correction scale or router supervision rather than being a stable property of the architecture.
Extended reading notes
Core claim
The central claim is that two modules, used together on a CogVideoX-style diffusion transformer, overcome the identity loss caused by global unstructured attention. The local router injects spatial structure: it extracts six facial components with an off-the-shelf segmentation network, encodes them into token sequences, computes per-token weights $w_m$ that score how much each component should influence the joint latent tokens, and refines the latents as $z^* = z + \alpha \sum_m w_m^\top \odot \varphi(l_m, z)$, where $\varphi$ is a cross-attention reconstruction using local facial tokens as keys and values. A route loss supervises these weights with the ground-truth segmentation mask, so the
Load-bearing premise
The method assumes the numeric biases it adds to the denoised video tokens keep those tokens inside the distribution the video decoder was trained to turn into images, so visual quality is not silently traded for higher face-similarity scores.
Editorial extensions
If this is right
- With the local router alone, FaceSim-Curricular/ArcFace rise from 0.312/0.291 to 0.361/0.334 on the ablation subset; with the temporal module alone they reach 0.358/0.327; together they reach 0.428/0.398.
- In the user study, LaVieID wins the largest share of preference votes on all four criteria, with identity-similarity preference 0.504 versus 0.155 for the next-best identity baseline.
- The temporal module changes motion behavior: it produces a 'softly gazes downward' action that the baseline and the local-router-only variant miss, suggesting text can drive subject motion when temporal dependencies are explicit.
- Because the autoregressive module acts on denoised latents without modifying the base model's latent space, the whole enhancement can be trained on a single GPU (about 60 hours for 10K steps), making it a cheap plug-in rather than a full retrain.
- The best FID is still held by CogVideoX+IPA (167.117 vs 174.121), which the authors attribute to identity-preserving generation being an out-of-distribution task; LaVieID trades a little raw realism for much higher identity fidelity.
Reading between the lines
- The paper does not report how sensitive the result is to $\beta$ or how the corrected latents compare with the decoder's training distribution; a beta sweep with FID and FaceSim would show whether the identity gain is a genuine latent-space improvement or a metric-side effect.
- Because the router's supervision signal comes from a face-segmentation network, the method inherits that network's limits: on occluded, heavily angled, or stylized faces where segmentation fails, identity preservation should degrade, and removing the segmentation loss would isolate that dependency.
- The benchmark uses 30 evaluation subjects and 90 prompts; the claim of state-of-the-art identity preservation should be read as demonstrated on that distribution until tested across more identities, longer videos, and more varied motion.
- The chunk-wise bias formulation suggests a direct extension to longer videos by keeping the chunk architecture and increasing the number of chunks, but then bias corrections accumulate, so the stability of teacher forcing under longer rollout remains an open test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes LaVieID, an identity-preserving text-to-video framework built on a CogVideoX-style diffusion transformer. It adds two modules: a local router that uses fine-grained facial-component tokens to reweight latent tokens spatially, and a temporal autoregressive module (TAM) that splits denoised latent tokens into chunks and refines each chunk by predicting biases conditioned on the previous chunk. The reported training objective is a diffusion loss plus a router cross-entropy loss. Experiments compare LaVieID against four open-source baselines on a 49-frame 720x480 benchmark, reporting improved FaceSim, subject/background consistency, and user-study identity-similarity scores. Code and models are promised.
Significance. If the central claim holds, the paper offers a simple, low-cost recipe (single A100 GPU, 10K steps) for improving identity preservation in DiT video models. The local-router idea and the post-denoising bias-refinement scheme are interesting and are presented with clear equations and external baselines. However, the manuscript has a load-bearing internal inconsistency: the TAM appears to receive no gradient from the stated training objective. The empirical support also lacks error bars and uses different evaluation subsets across tables. These issues must be addressed before the proposed contributions can be accepted.
major comments (4)
- [Sec. 3.3–3.4, Eq. (6)] The total objective in Eq. (6) contains only L_diff (Eq. 1) and L_route (Eq. 5). L_diff is a noise-prediction loss on ε_θ at arbitrary timestep t, before TAM is applied, and L_route depends only on the router weights w_m. Neither term is a function of the TAM outputs b_k or c*_k defined in Eq. (4). Consequently, if Eq. (6) is the complete training objective, the TAM parameters (ψ and φ) receive zero gradient and cannot be learned. This contradicts the statement in Sec. 3.2 that the local encoder is fine-tuned 'alongside the local router and the temporal autoregressive module' and the Table 2 attribution of gains to TAM (e.g., FaceSim-Curricular 0.312→0.358). Please specify the actual loss used to train TAM (e.g., a reconstruction or perceptual loss on decoded frames), or state clearly that TAM is not trained; in the latter case the performance attribution is unsupported.
- [Sec. 4.3, Table 2] The ablation study uses 'a subset of the test dataset' and reports FID=168.714 for LaVieID, while Table 1 reports FID=174.121 for the same model on the full test set. Because the numbers in Tables 1 and 2 come from different evaluation sets, the ablation cannot be used to claim improvements that transfer to the main comparison. Please report results on the full test set, or at least list the subset size and composition and provide the baselines on the same subset, so readers can calibrate the gains.
- [Sec. 4.2, Table 1] All metrics in Table 1 are single-run point estimates with no error bars, confidence intervals, or multi-seed variance. The margins over the strongest baseline ConsisID are small: +0.013 FaceSim-Curricular, +0.012 FaceSim-ArcFace, +0.016 subject consistency, +0.008 background consistency, and the FID is 6.3 worse than CogVideoX+IPA. Without variance estimates or significance tests, the claimed 'state-of-the-art' improvement may be within run-to-run noise. Please add bootstrapped confidence intervals or multiple-seed runs.
- [Sec. 4.4] The user study uses 30 participants and 50 questions, but no confidence intervals, significance tests, or agreement measures are reported. Since the IDS preference for LaVieID (0.504) is a key subjective claim, please provide at least a binomial confidence interval and a description of how ties and random ordering were handled.
minor comments (5)
- [Eq. (1) vs Sec. 3.3] The notation z0 is used both for the clean latent in Eq. (1) and for the fully denoised latent tokens in Sec. 3.3 (e.g., Eq. (4)). Please disambiguate these two meanings.
- [Sec. 3.2, Eq. (5)] The mapping from the ground-truth segmentation mask y_m to latent-token positions is not described. How are token-level labels obtained for L'=17750 latent tokens? Please clarify.
- [Sec. 3.3, Discussion] The claim that TAM 'will not alter the latent space of the baseline DiT' is only true if the decoder is used as-is; adding β b_k changes the input distribution to the decoder. The stability argument would benefit from a sensitivity analysis over β or a distributional check on the corrected latents.
- [References and figures] There are duplicated VideoPoet references ([30] and [31]) and some garbled text in the Figure 3 caption; please clean these up before publication.
- [Sec. 3.2] The text first says the local encoder architecture is adopted from [66] and later says it is fine-tuned. Please state explicitly which parameters are initialized from [66] and which are frozen/unfrozen during training.
Circularity Check
No significant circularity: the central claims are evaluated against external baselines and no target metric is fitted by construction.
full rationale
The paper's derivation chain is self-contained against external benchmarks: FaceSim, VBench subject/background consistency, and FID are computed on the ConsisID dataset with external baselines, and the reported gains are attributable to the proposed modules via ablations. The local router is supervised by segmentation masks (Eq. 5), not by the identity metrics, and the temporal autoregressive module's bias correction (Eq. 4) is applied post-hoc with fixed hyperparameters alpha=1 and beta=0.2; no parameter is fitted to the reported identity scores. The only self-citation ([26], ConsistentID) is used to position the method as a contrast and is not load-bearing. A separate non-circularity concern: in Sec. 3.4, Eq. (6) includes only L_diff and L_route, so as written the temporal autoregressive module has no objective that depends on b_k or c*_k; this is an internal-consistency/omission issue that would affect correctness, not a circular reduction, and therefore does not raise the circularity score.
Assumptions & free parameters
free parameters (5)
- alpha (spatial enhancement scale) =
1
- beta (temporal bias scale) =
0.2
- K (number of temporal chunks) =
4
- lambda_route (mask loss weight) =
1
- N (layers in TAM transformer) =
6
assumptions (5)
- domain assumption CogVideoX DiT is an adequate base model and its latent space is compatible with inserted router and post-denoising corrections
- domain assumption BiSeNet segmentation masks correctly delineate identity-relevant facial components
- domain assumption Early DiT blocks dominate subject structure
- domain assumption FaceSim and VBench metrics are faithful proxies for identity preservation and temporal consistency
- standard math Standard latent diffusion loss in Eq. (1) is the correct training objective for the conditional generation task
Cite this review
Pith. "Pith review of LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation." pith.science (2026). https://pith.science/paper/L5MIH5LH
@misc{pith2026250807603,
author = {Pith},
title = {Pith review of: LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation},
year = {2026},
howpublished = {\url{https://pith.science/paper/L5MIH5LH}},
note = {Machine review of arXiv:2508.07603}
}
read the original abstract
In this paper, we present LaVieID, a novel \underline{l}ocal \underline{a}utoregressive \underline{vi}d\underline{e}o diffusion framework designed to tackle the challenging \underline{id}entity-preserving text-to-video task. The key idea of LaVieID is to mitigate the loss of identity information inherent in the stochastic global generation process of diffusion transformers (DiTs) from both spatial and temporal perspectives. Specifically, unlike the global and unstructured modeling of facial latent states in existing DiTs, LaVieID introduces a local router to explicitly represent latent states by weighted combinations of fine-grained local facial structures. This alleviates undesirable feature interference and encourages DiTs to capture distinctive facial characteristics. Furthermore, a temporal autoregressive module is integrated into LaVieID to refine denoised latent tokens before video decoding. This module divides latent tokens temporally into chunks, exploiting their long-range temporal dependencies to predict biases for rectifying tokens, thereby significantly enhancing inter-frame identity consistency. Consequently, LaVieID can generate high-fidelity personalized videos and achieve state-of-the-art performance. Our code and models are available at https://github.com/ssugarwh/LaVieID.
Reference graph
Works this paper leans on
-
[1]
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normaliza- tion. arXiv preprint arXiv:1607.06450 (2016). LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation MM ’25, October 27–31, 2025, Dublin, Ireland
arXiv 2016
-
[2]
Hila Chefer, Shiran Zada, Roni Paiss, Ariel Ephrat, Omer Tov, Michael Rubinstein, Lior Wolf, Tali Dekel, Tomer Michaeli, and Inbar Mosseri. 2024. Still-moving: Customized video generation without customized video data. ACM Transactions on Graphics 43, 6 (2024), 1–11
work page 2024
-
[3]
Boyuan Chen, Diego Martí Monsó, Yilun Du, Max Simchowitz, Russ Tedrake, and Vincent Sitzmann. 2024. Diffusion forcing: Next-token prediction meets full-sequence diffusion. Advances in Neural Information Processing Systems 37 (2024), 24081–24125
work page 2024
-
[4]
Pengtao Chen, Mingzhu Shen, Peng Ye, Jianjian Cao, Chongjun Tu, Christos- Savvas Bouganis, Yiren Zhao, and Tao Chen. 2024. 𝐷𝑒𝑙𝑡𝑎 -DiT: A Training- Free Acceleration Method Tailored for Diffusion Transformers. arXiv preprint arXiv:2406.01125 (2024)
arXiv 2024
-
[5]
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah
-
[6]
Jiahao Cui, Hui Li, Yao Yao, Hao Zhu, Hanlin Shang, Kaihui Cheng, Hang Zhou, Siyu Zhu, and Jingdong Wang. 2024. Hallo2: Long-duration and high-resolution audio-driven portrait image animation. arXiv preprint arXiv:2410.07718 (2024)
arXiv 2024
-
[7]
Haoge Deng, Ting Pan, Haiwen Diao, Zhengxiong Luo, Yufeng Cui, Huchuan Lu, Shiguang Shan, Yonggang Qi, and Xinlong Wang. 2024. Autoregressive Video Generation without Vector Quantization. arXiv preprint arXiv:2412.14169 (2024)
arXiv 2024
-
[8]
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. 2019. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 4690–4699
work page 2019
Show all 73 references
-
[9]
Michael Fuest, Vincent Tao Hu, and Björn Ommer. 2025. MaskFlow: Dis- crete Flows For Flexible and Efficient Long Video Generation. arXiv preprint arXiv:2502.11234 (2025)
2025 arXiv
-
[10]
Junyao Gao, SUN Yanan, Fei Shen, Xin Jiang, Zhening Xing, Kai Chen, and Cairong Zhao. [n. d.]. FaceShot: Bring Any Character into Life. In The Thirteenth International Conference on Learning Representations
-
[11]
Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin, Guan Pang, David Jacobs, Jia- Bin Huang, and Devi Parikh. 2022. Long video generation with time-agnostic vqgan and time-sensitive transformer. In Proceedings of the European Conference on Computer Vision. 102–118
2022
-
[12]
Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu, and Lingpeng Kong. 2022. DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models. In International Conference on Learning Representations
2022
-
[13]
Yuchao Gu, Weijia Mao, and Mike Zheng Shou. 2025. Long-Context Autoregres- sive Video Modeling with Next-Frame Prediction. arXiv preprint arXiv:2503.19325 (2025)
2025 arXiv
-
[14]
Jianzhu Guo, Dingyun Zhang, Xiaoqiang Liu, Zhizhou Zhong, Yuan Zhang, Pengfei Wan, and Di Zhang. 2024. Liveportrait: Efficient portrait animation with stitching and retargeting control. arXiv preprint arXiv:2407.03168 (2024)
2024 arXiv
-
[15]
Zinan Guo, Yanze Wu, Chen Zhuowei, Peng Zhang, Qian He, et al. 2024. Pulid: Pure and lightning id customization via contrastive alignment. Advances in Neural Information Processing Systems 37 (2024), 36777–36804
2024
-
[16]
Ligong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri, Kyle Olszewski, Shervin Minaee, Dimitris Metaxas, and Sergey Tulyakov. 2022. Show me what and tell me how: Video synthesis via multimodal conditioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern...
2022
-
[17]
Xuanhua He, Quande Liu, Shengju Qian, Xin Wang, Tao Hu, Ke Cao, Keyu Yan, and Jie Zhang. 2024. Id-animator: Zero-shot identity-preserving human video generation. arXiv preprint arXiv:2404.15275 (2024)
2024 arXiv
-
[18]
Yefei He, Yuanyu He, Shaoxuan He, Feng Chen, Hong Zhou, Kaipeng Zhang, and Bohan Zhuang. 2025. Neighboring autoregressive modeling for efficient visual generation. arXiv preprint arXiv:2503.10696 (2025)
2025 arXiv
-
[19]
Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. 2022. Latent video diffusion models for high-fidelity long video generation. arXiv preprint arXiv:2211.13221 (2022)
2022 arXiv
-
[20]
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30 (2017)
2017
-
[21]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems 33 (2020), 6840–6851
2020
-
[22]
Jonathan Ho and Tim Salimans. [n. d.]. Classifier-Free Diffusion Guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications
2021
-
[23]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022. Video diffusion models. Advances in Neu- ral Information Processing Systems 35 (2022), 8633–8646
2022
-
[24]
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. 2022. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868 (2022)
2022 arXiv
-
[25]
Jinyi Hu, Shengding Hu, Yuxuan Song, Yufei Huang, Mingxuan Wang, Hao Zhou, Zhiyuan Liu, Wei-Ying Ma, and Maosong Sun. 2024. ACDiT: Interpolating Autoregressive Conditional Modeling and Diffusion Transformer. arXiv preprint arXiv:2412.07720 (2024)
2024
-
[26]
Jiehui Huang, Xiao Dong, Wenhui Song, Zheng Chong, Zhenchao Tang, Jun Zhou, Yuhao Cheng, Long Chen, Hanhui Li, Yiqiang Yan, et al . 2024. Consistentid: Portrait generation with multimodal fine-grained identity preserving. arXiv preprint arXiv:2404.16771 (2024)
2024 arXiv
-
[27]
Jiancheng Huang, Mingfu Yan, Songyan Chen, Yi Huang, and Shifeng Chen. 2024. MagicFight: Personalized Martial Arts Combat Video Generation. In Proceedings of the ACM International Conference on Multimedia . 10833–10842
2024
-
[28]
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuan- han Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. 2024. Vbench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision...
2024
-
[29]
Liming Jiang, Qing Yan, Yumin Jia, Zichuan Liu, Hao Kang, and Xin Lu. 2025. InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity. arXiv preprint arXiv:2503.16418 (2025)
2025 arXiv
-
[30]
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, et al. 2023. Videopoet: A large language model for zero-shot video generation. arXiv preprint arXiv:2312.14125 (2023)
2023 arXiv
-
[31]
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming Chang Chiu, et al. 2024. VideoPoet: A Large Language Model for Zero-Shot Video Generation. International Conference on Machine Learning 235 (202...
2024
-
[32]
PKU-Yuan Lab, Tuzhan AI, et al. 2024. Open-sora-plan
2024
-
[33]
Duong H Le, Tuan Pham, Sangho Lee, Christopher Clark, Aniruddha Kembhavi, Stephan Mandt, Ranjay Krishna, and Jiasen Lu. 2025. One Diffusion to Generate Them All. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
2025
-
[34]
Hengjia Li, Lifan Jiang, Xi Xiao, Tianyang Wang, Hongwei Yi, Boxi Wu, and Deng Cai. 2025. MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization. arXiv preprint arXiv:2503.12689 (2025)
2025 arXiv
-
[35]
Hengjia Li, Haonan Qiu, Shiwei Zhang, Xiang Wang, Yujie Wei, Zekun Li, Yingya Zhang, Boxi Wu, and Deng Cai. 2024. PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation. arXiv preprint arXiv:2411.17048 (2024)
2024 arXiv
-
[36]
Xiang Li, Kai Qiu, Hao Chen, Jason Kuen, Zhe Lin, Rita Singh, and Bhiksha Raj
-
[37]
Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. 2024. Latte: Latent diffusion transformer for video generation. arXiv preprint arXiv:2401.03048 (2024)
2024 arXiv
-
[38]
Yiyang Ma, Xingchao Liu, Xiaokang Chen, Wen Liu, Chengyue Wu, Zhiyu Wu, Zizheng Pan, Zhenda Xie, Haowei Zhang, Liang Zhao, et al . 2025. Janusflow: Harmonizing autoregression and rectified flow for unified multimodal under- standing and generation. Proceedings of the IEEE Conf...
2025
-
[39]
Ze Ma, Daquan Zhou, Chun-Hsiao Yeh, Xue-She Wang, Xiuyu Li, Huanrui Yang, Zhen Dong, Kurt Keutzer, and Jiashi Feng. 2024. Magic-me: Identity-specific video customized diffusion. arXiv preprint arXiv:2402.09368 (2024)
2024 arXiv
-
[40]
OpenAI. 2023. Video generation models as world simulators
2023
-
[41]
William Peebles and Saining Xie. 2023. Scalable diffusion models with trans- formers. In Proceedings of the IEEE International Conference on Computer Vision . 4195–4205
2023
-
[42]
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In International Conference on Machine Learning . 8821–8831
2021
-
[43]
Shuhuai Ren, Shuming Ma, Xu Sun, and Furu Wei. 2025. Next Block Predic- tion: Video Generation via Semi-Auto-Regressive Modeling. arXiv preprint arXiv:2502.07737 (2025)
2025 arXiv
-
[44]
Zhongwei Ren, Yunchao Wei, Xun Guo, Yao Zhao, Bingyi Kang, Jiashi Feng, and Xiaojie Jin. 2025. VideoWorld: Exploring Knowledge Learning from Unlabeled Videos. arXiv preprint arXiv:2501.09781 (2025)
2025 arXiv
-
[45]
David Ruhe, Jonathan Heek, Tim Salimans, and Emiel Hoogeboom. 2024. Rolling diffusion models. In International Conference on Machine Learning . 42818–42835
2024
-
[46]
Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi. 2022. Palette: Image-to-image diffusion models. In Proceedings of the ACM SIGGRAPH Conference . 1–10
2022
-
[47]
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. 2023. Make-A-Video: Text-to-Video Generation without Text-Video Data. InInternational Conference on Learning Representations
2023
-
[48]
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2021. Denoising Diffusion Implicit Models. In International Conference on Learning Representations
2021
-
[49]
J Su, H Zhang, X Li, J Zhang, and Y RoFormer Li. 2021. Enhanced transformer with rotary position embedding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (AC...
2021
-
[50]
Mingzhen Sun, Weining Wang, Gen Li, Jiawei Liu, Jiahui Sun, Wanquan Feng, Shanshan Lao, SiYu Zhou, Qian He, and Jing Liu. 2025. Ar-diffusion: Asyn- chronous video generation with auto-regressive diffusion. In Proceedings of the Computer Vision and Pattern Recognition Conferenc...
2025
-
[51]
Yusuke Tashiro, Jiaming Song, Yang Song, and Stefano Ermon. 2021. Csdi: Con- ditional score-based diffusion models for probabilistic time series imputation. Advances in Neural Information Processing Systems 34 (2021), 24804–24816
2021
-
[52]
Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang. 2024. Visual autoregressive modeling: Scalable image generation via next-scale prediction. Advances in Neural Information Processing Systems 37 (2024), 84839–84865
2024
-
[53]
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. 2024. Eyes wide shut? exploring the visual shortcomings of multimodal llms. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 9568–9578
2024
-
[54]
Aaron Van Den Oord, Oriol Vinyals, et al. 2017. Neural discrete representation learning. Advances in Neural Information Processing Systems 30 (2017)
2017
-
[55]
Huawei Wei, Zejun Yang, and Zhisheng Wang. 2024. Aniportrait: Audio-driven synthesis of photorealistic portrait animation. arXiv preprint arXiv:2403.17694 (2024)
2024 arXiv
-
[56]
Yujie Wei, Shiwei Zhang, Zhiwu Qing, Hangjie Yuan, Zhiheng Liu, Yu Liu, Yingya Zhang, Jingren Zhou, and Hongming Shan. 2024. Dreamvideo: Composing your dream videos with customized subject and motion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni...
2024
-
[57]
Ronald J Williams and David Zipser. 1989. A learning algorithm for continually running fully recurrent neural networks.Neural computation 1, 2 (1989), 270–280
1989
-
[58]
Zhenyu Xie, Haoye Dong, Yufei Gao, Zehua Ma, and Xiaodan Liang. 2024. DreamVTON: Customizing 3D Virtual Try-on with Personalized Diffusion Mod- els. In Proceedings of the ACM International Conference on Multimedia . 10784– 10793
2024
-
[59]
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. 2021. Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:2104.10157 (2021)
2021 arXiv
-
[60]
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al . 2024. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072 (2024)
2024 arXiv
-
[61]
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. 2023. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv preprint arXiv:2308.06721 (2023)
2023 arXiv
-
[62]
Tianwei Yin, Qiang Zhang, Richard Zhang, William T Freeman, Fredo Durand, Eli Shechtman, and Xun Huang. 2024. From slow bidirectional to fast autoregressive video diffusion models. arXiv preprint arXiv:2412.07772 2 (2024)
2024
-
[63]
Changqian Yu, Jingbo Wang, Chao Peng, Changxin Gao, Gang Yu, and Nong Sang. 2018. Bisenet: Bilateral segmentation network for real-time semantic segmentation. In Proceedings of the European Conference on Computer Vision . 325–341
2018
-
[64]
Lijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama, Han Zhang, Huiwen Chang, Alexander G Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, et al. 2023. Magvit: Masked generative video transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition...
2023
-
[65]
Lijun Yu, Jose Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G Haupt- mann, et al. 2024. Language Model Beats Diffusion-Tokenizer is key to visual generation. In International Conference on Learning ...
2024
-
[66]
Shenghai Yuan, Jinfa Huang, Xianyi He, Yunyuan Ge, Yujun Shi, Liuhan Chen, Jiebo Luo, and Li Yuan. 2025. Identity-Preserving Text-to-Video Generation by Frequency Decomposition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
2025
-
[67]
Yuechen Zhang, Yaoyang Liu, Bin Xia, Bohao Peng, Zexin Yan, Eric Lo, and Jiaya Jia. 2025. Magic Mirror: ID-Preserved Video Generation in Video Diffusion Transformers. arXiv preprint arXiv:2501.03931 (2025)
2025
-
[68]
Yunpeng Zhang, Qiang Wang, Fan Jiang, Yaqi Fan, Mu Xu, and Yonggang Qi
-
[69]
Dingcheng Zhen, Shunshun Yin, Shiyang Qin, Hou Yi, Ziwei Zhang, Siyuan Liu, Gan Qi, and Ming Tao. 2025. Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation. In Proceedings of the Computer Vision and Pattern Recognition Conference ....
2025
-
[70]
Yong Zhong, Zhuoyi Yang, Jiayan Teng, Xiaotao Gu, and Chongxuan Li. 2025. Concat-ID: Towards Universal Identity-Preserving Video Synthesis. arXiv preprint arXiv:2503.14151 (2025)
2025 arXiv
-
[2023]
IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 10850–10869
Diffusion models in vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 10850–10869
2023
-
[2024]
ControlVAR: Exploring Controllable Visual Autoregressive Modeling.arXiv preprint arXiv:2406.09750 (2024)
2024 arXiv
-
[2025]
arXiv preprint arXiv:2502.13995 (2025)
Fantasyid: Face knowledge enhanced id-preserving video generation. arXiv preprint arXiv:2502.13995 (2025)
2025 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.