{"work":{"id":"1f9d1d3b-a6d6-45a9-9f13-51393c03be8a","openalex_id":"https://openalex.org/W4383994075","doi":"10.48550/arxiv.2307.04725","arxiv_id":"2307.04725","raw_key":null,"title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","authors":null,"authors_text":"Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao","year":2023,"venue":"cs.CV","abstract":"With the advance of text-to-image (T2I) diffusion models (e.g., Stable Diffusion) and corresponding personalization techniques such as DreamBooth and LoRA, everyone can manifest their imagination into high-quality images at an affordable cost. However, adding motion dynamics to existing high-quality personalized T2Is and enabling them to generate animations remains an open challenge. In this paper, we present AnimateDiff, a practical framework for animating personalized T2I models without requiring model-specific tuning. At the core of our framework is a plug-and-play motion module that can be trained once and seamlessly integrated into any personalized T2Is originating from the same base T2I. Through our proposed training strategy, the motion module effectively learns transferable motion priors from real-world videos. Once trained, the motion module can be inserted into a personalized T2I model to form a personalized animation generator. We further propose MotionLoRA, a lightweight fine-tuning technique for AnimateDiff that enables a pre-trained motion module to adapt to new motion patterns, such as different shot types, at a low training and data collection cost. We evaluate AnimateDiff and MotionLoRA on several public representative personalized T2I models collected from the community. The results demonstrate that our approaches help these models generate temporally smooth animation clips while preserving the visual quality and motion diversity. Codes and pre-trained weights are available at https://github.com/guoyww/AnimateDiff.","external_url":"https://arxiv.org/abs/2307.04725","cited_by_count":84,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2307.04725","created_at":"2026-05-09T06:15:38.154562+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","render_title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning"},"hub":{"state":{"work_id":"1f9d1d3b-a6d6-45a9-9f13-51393c03be8a","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":121,"external_cited_by_count":84,"distinct_field_count":6,"first_pith_cited_at":"2023-10-30T13:12:40+00:00","last_pith_cited_at":"2026-07-09T17:59:52+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T08:09:24.014206+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":28},{"context_role":"method","n":5},{"context_role":"baseline","n":4}],"polarity_counts":[{"context_polarity":"background","n":28},{"context_polarity":"use_method","n":5},{"context_polarity":"baseline","n":4}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","claims":[{"claim_text":"With the advance of text-to-image (T2I) diffusion models (e.g., Stable Diffusion) and corresponding personalization techniques such as DreamBooth and LoRA, everyone can manifest their imagination into high-quality images at an affordable cost. However, adding motion dynamics to existing high-quality personalized T2Is and enabling them to generate animations remains an open challenge. In this paper, we present AnimateDiff, a practical framework for animating personalized T2I models without requiring model-specific tuning. At the core of our framework is a plug-and-play motion module that can be","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Flicker. Motion Smooth. Dynamic Degree Aesthetic Quality Imaging Quality Object Class Generation-only Models ModelScope [113] 1.7B 78.05 66.54 89.87 95.29 98.28 95.79 66.39 52.06 58.57 82.25 LaVie [117] 3B 78.78 70.31 91.41 97.47 98.30 96.38 49.72 54.94 61.90 91.82 Show-1 [144] 6B 80.42 72.98 95.53 98.02 99.12 98.24 44.44 57.35 58.66 93.07 AnimateDiff-V2 [36] - 82.90 69.75 95.30 97.68 98.75 97.76 40.83 67.16 70.10 90.90 VideoCrafter-2.0 [10] - 82.20 73.42 96.85 98.22 98.41 97.73 42.50 63.13 67.2","claim_type":"baseline","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"Flicker. Motion Smooth. Dynamic Degree Aesthetic Quality Imaging Quality Object Class Generation-only Models ModelScope [112] 1.7B 78.05 66.54 89.87 95.29 98.28 95.79 66.39 52.06 58.57 82.25 LaVie [116] 3B 78.78 70.31 91.41 97.47 98.30 96.38 49.72 54.94 61.90 91.82 Show-1 [143] 6B 80.42 72.98 95.53 98.02 99.12 98.24 44.44 57.35 58.66 93.07 AnimateDiff-V2 [35] - 82.90 69.75 95.30 97.68 98.75 97.76 40.83 67.16 70.10 90.90 VideoCrafter-2.0 [10] - 82.20 73.42 96.85 98.22 98.41 97.73 42.50 63.13 67.2","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Ego2Exo generation within a single continuous sequence model. 2 Related Works Diffusion-based Video Generation. Recent advances in diffusion models and large-scale datasets have rapidly improved the quality and diversity of video generation [4,13,14,17,28,29,34,38, 40,51,53,54]. Representative systems such as Tune-A-Video [47], Video Diffusion Models [14], Stable Video Diffusion [3], AnimateDiff [9], LTX 2 [10], Hunyan- Video [21] and WAN2.2 [45] demonstrate the ability of diffusion models to ge","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023. [24] Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang, Pengfei Wan, Di Zhang, and Wanli Ouyang. Flow-grpo: Training flow matching models via online rl. In Adv. Neural Inform. Process. Syst., 2025. 13 [25] Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Ani","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"-We develop a Spatially-Structured Co-Generation paradigm using an asym- metric co-attention mask to embed physical interaction rules into the DiT. Thisapproachforcesthemodeltorespectgeometricconstraintsandsubstan- tially reduces hand-object interpenetration, while ensuring zero additional computational cost at inference. 4 X. Luo et al. 2 Related Works 2.1 Video Diffusion Models Video diffusion models [1,13,20,21,42] have evolved rapidly from image diffusion to temporally coherent video synthes","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"1 Video Generation Models In recent years, diffusion models have been widely applied to image synthesis [25], image editing [2, 7, 9], video generation [1, 27, 3, 6, 20, 19, 21, 22, 17], and procedural generation [28, 29, 31, 39]. Within video generation, the currently prevalent DiT architecture [8] has progressively surpassed earlier GAN-based [23] and UNet-based methods [6, 28] by significantly enhancing visual fidelity and temporal consistency. Today, rapidly evolving DiT frameworks form the ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (27 contexts).","role_counts":[{"n":27,"context_role":"background"},{"n":5,"context_role":"method"},{"n":4,"context_role":"baseline"}]},"error":null,"updated_at":"2026-06-30T21:30:30.895269+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"a467d605-7b49-4edc-9201-b84ad4dea927","orcid":null,"display_name":"Yuwei Guo"},{"id":"52682ca5-2b74-4185-be53-fa1a9d950236","orcid":null,"display_name":"Ceyuan Yang"},{"id":"8c616772-abfe-454d-bb1e-ec8df5768210","orcid":null,"display_name":"Anyi Rao"},{"id":"391cd685-7db9-4521-ab3c-ff9c2600c9b6","orcid":null,"display_name":"Zhengyang Liang"},{"id":"6c6cf51d-5458-4dcd-bb90-622abce3e395","orcid":null,"display_name":"Yaohui Wang"},{"id":"044a7cb9-9f77-43f0-b616-05b4e95843b7","orcid":null,"display_name":"Yu Qiao"}]},"error":null,"updated_at":"2026-06-30T21:30:30.891822+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T14:51:36.442323+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":28},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":28},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":22},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":19},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":14},{"title":"Make-A-Video: Text-to-Video Generation without Text-Video Data","work_id":"52a801fc-a707-45a1-a8cd-0d6702f124ab","shared_citers":13},{"title":"Imagen Video: High Definition Video Generation with Diffusion Models","work_id":"bb20d241-dc6f-4b0a-b071-fd43a2cbd57f","shared_citers":10},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":9},{"title":"CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers","work_id":"2dbd6bcd-fc98-4fbf-b586-f6d94fe1abd2","shared_citers":9},{"title":"ModelScope Text-to-Video Technical Report","work_id":"1b1baf78-58ec-44d0-b700-84dff57b2f1f","shared_citers":9},{"title":"CameraCtrl: Enabling Camera Control for Text-to-Video Generation","work_id":"1c05c278-c023-4ef0-a359-25a41f1065eb","shared_citers":8},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":8},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":8},{"title":"Magicvideo: Efficient video generation with latent diffusion models","work_id":"aad71b40-2721-438d-8e8c-97f84063ed39","shared_citers":8},{"title":"Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion","work_id":"53e58ef9-7932-4b83-b757-34ac14db3e0f","shared_citers":8},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":7},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":7},{"title":"Latte: Latent Diffusion Transformer for Video Generation","work_id":"5328e907-7278-4781-a2bb-c5ef40dc87fb","shared_citers":7},{"title":"Score-Based Generative Modeling through Stochastic Differential Equations","work_id":"d9110e53-a5d4-4794-a4c5-a575e91c31ad","shared_citers":7},{"title":"VideoGPT: Video Generation using VQ-VAE and Transformers","work_id":"703c74c3-fa5e-455c-8c00-697c83511fcf","shared_citers":7},{"title":"Videopoet: A large language model for zero-shot video generation","work_id":"5cc3572d-7e2f-4431-ae42-d9282a42a800","shared_citers":7},{"title":"Latent video diffusion models for high-fidelity video generation with arbitrary lengths","work_id":"23338b3d-620a-4954-904f-bab6a577b8a5","shared_citers":6},{"title":"LTX-Video: Realtime Video Latent Diffusion","work_id":"cee5c521-3ce9-466e-a035-1e42f89254f4","shared_citers":6},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":6}],"time_series":[{"n":1,"year":2023},{"n":7,"year":2024},{"n":39,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T14:51:47.517510+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T14:51:38.955824+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","claims":[{"claim_text":"With the advance of text-to-image (T2I) diffusion models (e.g., Stable Diffusion) and corresponding personalization techniques such as DreamBooth and LoRA, everyone can manifest their imagination into high-quality images at an affordable cost. However, adding motion dynamics to existing high-quality personalized T2Is and enabling them to generate animations remains an open challenge. In this paper, we present AnimateDiff, a practical framework for animating personalized T2I models without requiring model-specific tuning. At the core of our framework is a plug-and-play motion module that can be","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Flicker. Motion Smooth. Dynamic Degree Aesthetic Quality Imaging Quality Object Class Generation-only Models ModelScope [113] 1.7B 78.05 66.54 89.87 95.29 98.28 95.79 66.39 52.06 58.57 82.25 LaVie [117] 3B 78.78 70.31 91.41 97.47 98.30 96.38 49.72 54.94 61.90 91.82 Show-1 [144] 6B 80.42 72.98 95.53 98.02 99.12 98.24 44.44 57.35 58.66 93.07 AnimateDiff-V2 [36] - 82.90 69.75 95.30 97.68 98.75 97.76 40.83 67.16 70.10 90.90 VideoCrafter-2.0 [10] - 82.20 73.42 96.85 98.22 98.41 97.73 42.50 63.13 67.2","claim_type":"baseline","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"Flicker. Motion Smooth. Dynamic Degree Aesthetic Quality Imaging Quality Object Class Generation-only Models ModelScope [112] 1.7B 78.05 66.54 89.87 95.29 98.28 95.79 66.39 52.06 58.57 82.25 LaVie [116] 3B 78.78 70.31 91.41 97.47 98.30 96.38 49.72 54.94 61.90 91.82 Show-1 [143] 6B 80.42 72.98 95.53 98.02 99.12 98.24 44.44 57.35 58.66 93.07 AnimateDiff-V2 [35] - 82.90 69.75 95.30 97.68 98.75 97.76 40.83 67.16 70.10 90.90 VideoCrafter-2.0 [10] - 82.20 73.42 96.85 98.22 98.41 97.73 42.50 63.13 67.2","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Ego2Exo generation within a single continuous sequence model. 2 Related Works Diffusion-based Video Generation. Recent advances in diffusion models and large-scale datasets have rapidly improved the quality and diversity of video generation [4,13,14,17,28,29,34,38, 40,51,53,54]. Representative systems such as Tune-A-Video [47], Video Diffusion Models [14], Stable Video Diffusion [3], AnimateDiff [9], LTX 2 [10], Hunyan- Video [21] and WAN2.2 [45] demonstrate the ability of diffusion models to ge","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023. [24] Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang, Pengfei Wan, Di Zhang, and Wanli Ouyang. Flow-grpo: Training flow matching models via online rl. In Adv. Neural Inform. Process. Syst., 2025. 13 [25] Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Ani","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"-We develop a Spatially-Structured Co-Generation paradigm using an asym- metric co-attention mask to embed physical interaction rules into the DiT. Thisapproachforcesthemodeltorespectgeometricconstraintsandsubstan- tially reduces hand-object interpenetration, while ensuring zero additional computational cost at inference. 4 X. Luo et al. 2 Related Works 2.1 Video Diffusion Models Video diffusion models [1,13,20,21,42] have evolved rapidly from image diffusion to temporally coherent video synthes","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"1 Video Generation Models In recent years, diffusion models have been widely applied to image synthesis [25], image editing [2, 7, 9], video generation [1, 27, 3, 6, 20, 19, 21, 22, 17], and procedural generation [28, 29, 31, 39]. Within video generation, the currently prevalent DiT architecture [8] has progressively surpassed earlier GAN-based [23] and UNet-based methods [6, 28] by significantly enhancing visual fidelity and temporal consistency. Today, rapidly evolving DiT frameworks form the ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (27 contexts).","role_counts":[{"n":27,"context_role":"background"},{"n":5,"context_role":"method"},{"n":4,"context_role":"baseline"}]},"error":null,"updated_at":"2026-06-30T21:30:30.897739+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","claims":[{"claim_text":"With the advance of text-to-image (T2I) diffusion models (e.g., Stable Diffusion) and corresponding personalization techniques such as DreamBooth and LoRA, everyone can manifest their imagination into high-quality images at an affordable cost. However, adding motion dynamics to existing high-quality personalized T2Is and enabling them to generate animations remains an open challenge. In this paper, we present AnimateDiff, a practical framework for animating personalized T2I models without requiring model-specific tuning. At the core of our framework is a plug-and-play motion module that can be","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T14:51:47.520100+00:00"}},"summary":{"title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","claims":[{"claim_text":"With the advance of text-to-image (T2I) diffusion models (e.g., Stable Diffusion) and corresponding personalization techniques such as DreamBooth and LoRA, everyone can manifest their imagination into high-quality images at an affordable cost. However, adding motion dynamics to existing high-quality personalized T2Is and enabling them to generate animations remains an open challenge. In this paper, we present AnimateDiff, a practical framework for animating personalized T2I models without requiring model-specific tuning. At the core of our framework is a plug-and-play motion module that can be","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":28},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":28},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":22},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":19},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":14},{"title":"Make-A-Video: Text-to-Video Generation without Text-Video Data","work_id":"52a801fc-a707-45a1-a8cd-0d6702f124ab","shared_citers":13},{"title":"Imagen Video: High Definition Video Generation with Diffusion Models","work_id":"bb20d241-dc6f-4b0a-b071-fd43a2cbd57f","shared_citers":10},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":9},{"title":"CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers","work_id":"2dbd6bcd-fc98-4fbf-b586-f6d94fe1abd2","shared_citers":9},{"title":"ModelScope Text-to-Video Technical Report","work_id":"1b1baf78-58ec-44d0-b700-84dff57b2f1f","shared_citers":9},{"title":"CameraCtrl: Enabling Camera Control for Text-to-Video Generation","work_id":"1c05c278-c023-4ef0-a359-25a41f1065eb","shared_citers":8},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":8},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":8},{"title":"Magicvideo: Efficient video generation with latent diffusion models","work_id":"aad71b40-2721-438d-8e8c-97f84063ed39","shared_citers":8},{"title":"Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion","work_id":"53e58ef9-7932-4b83-b757-34ac14db3e0f","shared_citers":8},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":7},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":7},{"title":"Latte: Latent Diffusion Transformer for Video Generation","work_id":"5328e907-7278-4781-a2bb-c5ef40dc87fb","shared_citers":7},{"title":"Score-Based Generative Modeling through Stochastic Differential Equations","work_id":"d9110e53-a5d4-4794-a4c5-a575e91c31ad","shared_citers":7},{"title":"VideoGPT: Video Generation using VQ-VAE and Transformers","work_id":"703c74c3-fa5e-455c-8c00-697c83511fcf","shared_citers":7},{"title":"Videopoet: A large language model for zero-shot video generation","work_id":"5cc3572d-7e2f-4431-ae42-d9282a42a800","shared_citers":7},{"title":"Latent video diffusion models for high-fidelity video generation with arbitrary lengths","work_id":"23338b3d-620a-4954-904f-bab6a577b8a5","shared_citers":6},{"title":"LTX-Video: Realtime Video Latent Diffusion","work_id":"cee5c521-3ce9-466e-a035-1e42f89254f4","shared_citers":6},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":6}],"time_series":[{"n":1,"year":2023},{"n":7,"year":2024},{"n":39,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"8c616772-abfe-454d-bb1e-ec8df5768210","orcid":null,"display_name":"Anyi Rao","source":"manual","import_confidence":0.72},{"id":"52682ca5-2b74-4185-be53-fa1a9d950236","orcid":null,"display_name":"Ceyuan Yang","source":"manual","import_confidence":0.72},{"id":"6c6cf51d-5458-4dcd-bb90-622abce3e395","orcid":null,"display_name":"Yaohui Wang","source":"manual","import_confidence":0.72},{"id":"044a7cb9-9f77-43f0-b616-05b4e95843b7","orcid":null,"display_name":"Yu Qiao","source":"manual","import_confidence":0.72},{"id":"a467d605-7b49-4edc-9201-b84ad4dea927","orcid":null,"display_name":"Yuwei Guo","source":"manual","import_confidence":0.72},{"id":"391cd685-7db9-4521-ab3c-ff9c2600c9b6","orcid":null,"display_name":"Zhengyang Liang","source":"manual","import_confidence":0.72}]}}