{"work":{"id":"52a801fc-a707-45a1-a8cd-0d6702f124ab","openalex_id":"https://openalex.org/W4298185919","doi":"10.48550/arxiv.2209.14792","arxiv_id":"2209.14792","raw_key":null,"title":"Make-A-Video: Text-to-Video Generation without Text-Video Data","authors":null,"authors_text":"Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang","year":2022,"venue":"cs.CV","abstract":"We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from paired text-image data, and learn how the world moves from unsupervised video footage. Make-A-Video has three advantages: (1) it accelerates training of the T2V model (it does not need to learn visual and multimodal representations from scratch), (2) it does not require paired text-video data, and (3) the generated videos inherit the vastness (diversity in aesthetic, fantastical depictions, etc.) of today's image generation models. We design a simple yet effective way to build on T2I models with novel and effective spatial-temporal modules. First, we decompose the full temporal U-Net and attention tensors and approximate them in space and time. Second, we design a spatial temporal pipeline to generate high resolution and frame rate videos with a video decoder, interpolation model and two super resolution models that can enable various applications besides T2V. In all aspects, spatial and temporal resolution, faithfulness to text, and quality, Make-A-Video sets the new state-of-the-art in text-to-video generation, as determined by both qualitative and quantitative measures.","external_url":"https://arxiv.org/abs/2209.14792","cited_by_count":314,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2209.14792","created_at":"2026-05-09T05:55:29.327628+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Make-A-Video: Text-to-Video Generation without Text-Video Data","render_title":"Make-A-Video: Text-to-Video Generation without Text-Video Data"},"hub":{"state":{"work_id":"52a801fc-a707-45a1-a8cd-0d6702f124ab","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":122,"external_cited_by_count":314,"distinct_field_count":10,"first_pith_cited_at":"2022-10-05T14:41:38+00:00","last_pith_cited_at":"2026-07-09T17:59:52+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T21:49:24.211282+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":28},{"context_role":"baseline","n":2}],"polarity_counts":[{"context_polarity":"background","n":27},{"context_polarity":"baseline","n":2},{"context_polarity":"unclear","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Make-A-Video: Text-to-Video Generation without Text-Video Data","claims":[{"claim_text":"We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from paired text-image data, and learn how the world moves from unsupervised video footage. Make-A-Video has three advantages: (1) it accelerates training of the T2V model (it does not need to learn visual and multimodal representations from scratch), (2) it does not require paired text-video data, and (3) the generated videos inherit the vastness (diversity in ae","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"[89] Soshi Shimada, Vladislav Golyanik, Weipeng Xu, and Christian Theobalt. Physcap: Physically plausible monocular 3d motion capture in real time. ACM Transactions on Graphics (ToG), 39(6):1-16, 2020. [90] Eftychios Sifakis and Jernej Barbic. Fem simulation of 3d deformable solids: a practitioner's guide to theory, discretization and model reduction. In Acm siggraph 2012 courses, pages 1-50. 2012. [91] Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yan","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"The surfer's expression is one of exhilaration and focus. A mid-shot from a low-angle perspective capturing the surfer's motion and the wave's power. B Related Works Video Diffusion Models.Video generation is of great benefit in neural simu- lators [2,3,9] and world models [4,5,23,37,46]. Synthesizing photorealistic videos using video diffusion models [7,8,14,20,28,29,33,34,36,49,67,73,82,90,100,108, 110] has become the community standard, following the substantial success of image diffusion mod","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":", Kautz, J.: Mocogan: Decomposing motion and content for video generation. In: Proceedings of the IEEE Conference on Computer Vision and 22 Pattern Recognition, pp. 1526-1535 (2018) [41] Li, Y., Min, M., Shen, D., Carlson, D., Carin, L.: Video generation from text. In: Proceed- ings of the AAAI Conference on Artificial Intelligence, vol. 32 (2018) [42] Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al.: Make-a-video: Text- to-video gene","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Signe Nørly, Srivatsan Srinivasan, Tobias Pfaff, Tom Hume, Vikas Verma, Weizhe Hua, William Zhu, Xinchen Yan, Xinyu Wang, Yelin Kim, Yuqing Du, and Yutian Chen. Veo. 2024. 11 [33] Chenyang Si, Ziqi Huang, Yuming Jiang, and Ziwei Liu. Freeu: Free lunch in diffusion u-net. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4733-4743, 2024. [34] Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"pattern recognition, pages 10684-10695, 2022. [31] Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Sali- mans, et al. Photorealistic text-to-image diffusion mod- els with deep language understanding.Advances in neu- ral information processing systems, 35:36479-36494, 2022. [32] Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"its exceptional capability in modeling plausible spatial appearance and coherent temporal dynamics. References [1] Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models.arXiv preprint arXiv:2210.02303, 2022. [2] Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Make-A-Video: Text-to-Video Generation without Text-Video Data because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (28 contexts).","role_counts":[{"n":28,"context_role":"background"},{"n":2,"context_role":"baseline"}]},"error":null,"updated_at":"2026-06-29T17:29:10.363775+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"46983c5e-881b-4d61-bfda-f727b70125ab","orcid":null,"display_name":"Uriel Singer"},{"id":"9f5d8ace-8e4d-427a-ba25-63c1aef83482","orcid":null,"display_name":"Adam Polyak"},{"id":"cd1b678c-b258-4e70-b5c6-dca9b58195a7","orcid":null,"display_name":"Thomas Hayes"},{"id":"6e4bc21b-4bc4-4b60-b6f3-2f13064aac7f","orcid":null,"display_name":"Xi Yin"},{"id":"798e0a0e-767c-4ee2-9fd3-429c27455832","orcid":null,"display_name":"Jie An"},{"id":"99b4149d-99b7-4361-aa15-ba5557443e1f","orcid":null,"display_name":"Songyang Zhang"}]},"error":null,"updated_at":"2026-06-29T17:29:10.936527+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T16:52:45.517776+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Imagen Video: High Definition Video Generation with Diffusion Models","work_id":"bb20d241-dc6f-4b0a-b071-fd43a2cbd57f","shared_citers":22},{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":21},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":17},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":14},{"title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","work_id":"1f9d1d3b-a6d6-45a9-9f13-51393c03be8a","shared_citers":13},{"title":"Score-Based Generative Modeling through Stochastic Differential Equations","work_id":"d9110e53-a5d4-4794-a4c5-a575e91c31ad","shared_citers":13},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":12},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":11},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":11},{"title":"CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers","work_id":"2dbd6bcd-fc98-4fbf-b586-f6d94fe1abd2","shared_citers":10},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":10},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":10},{"title":"Magicvideo: Efficient video generation with latent diffusion models","work_id":"aad71b40-2721-438d-8e8c-97f84063ed39","shared_citers":9},{"title":"Latte: Latent Diffusion Transformer for Video Generation","work_id":"5328e907-7278-4781-a2bb-c5ef40dc87fb","shared_citers":7},{"title":"LoRA: Low-Rank Adaptation of Large Language Models","work_id":"0426219a-789e-4964-adc8-a04538510818","shared_citers":7},{"title":"ModelScope Text-to-Video Technical Report","work_id":"1b1baf78-58ec-44d0-b700-84dff57b2f1f","shared_citers":7},{"title":"An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion","work_id":"ca618c21-3ba6-448e-bd86-bcecff3cdeb5","shared_citers":6},{"title":"arXiv:2210.02399 , year=","work_id":"a325cd53-6549-4726-b3e9-94509f0df168","shared_citers":6},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":6},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":6},{"title":"GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models","work_id":"34430d19-7919-48ce-88a5-17b3bfe2192e","shared_citers":6},{"title":"High-Resolution Image Synthesis with Latent Diffusion Models","work_id":"f0270d36-2952-47fb-84c1-95e3ec341126","shared_citers":6},{"title":"Movie Gen: A Cast of Media Foundation Models","work_id":"a6a118b0-002f-4b19-881f-7f1183e0d7d8","shared_citers":6},{"title":"Open-Sora: Democratizing Efficient Video Production for All","work_id":"8b29ba7b-3d84-4281-85b7-9eaf905afd7f","shared_citers":6}],"time_series":[{"n":1,"year":2022},{"n":6,"year":2023},{"n":4,"year":2024},{"n":33,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T17:06:38.874721+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T16:52:53.026183+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Make-A-Video: Text-to-Video Generation without Text-Video Data","claims":[{"claim_text":"We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from paired text-image data, and learn how the world moves from unsupervised video footage. Make-A-Video has three advantages: (1) it accelerates training of the T2V model (it does not need to learn visual and multimodal representations from scratch), (2) it does not require paired text-video data, and (3) the generated videos inherit the vastness (diversity in ae","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"[89] Soshi Shimada, Vladislav Golyanik, Weipeng Xu, and Christian Theobalt. Physcap: Physically plausible monocular 3d motion capture in real time. ACM Transactions on Graphics (ToG), 39(6):1-16, 2020. [90] Eftychios Sifakis and Jernej Barbic. Fem simulation of 3d deformable solids: a practitioner's guide to theory, discretization and model reduction. In Acm siggraph 2012 courses, pages 1-50. 2012. [91] Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yan","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"The surfer's expression is one of exhilaration and focus. A mid-shot from a low-angle perspective capturing the surfer's motion and the wave's power. B Related Works Video Diffusion Models.Video generation is of great benefit in neural simu- lators [2,3,9] and world models [4,5,23,37,46]. Synthesizing photorealistic videos using video diffusion models [7,8,14,20,28,29,33,34,36,49,67,73,82,90,100,108, 110] has become the community standard, following the substantial success of image diffusion mod","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":", Kautz, J.: Mocogan: Decomposing motion and content for video generation. In: Proceedings of the IEEE Conference on Computer Vision and 22 Pattern Recognition, pp. 1526-1535 (2018) [41] Li, Y., Min, M., Shen, D., Carlson, D., Carin, L.: Video generation from text. In: Proceed- ings of the AAAI Conference on Artificial Intelligence, vol. 32 (2018) [42] Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al.: Make-a-video: Text- to-video gene","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Signe Nørly, Srivatsan Srinivasan, Tobias Pfaff, Tom Hume, Vikas Verma, Weizhe Hua, William Zhu, Xinchen Yan, Xinyu Wang, Yelin Kim, Yuqing Du, and Yutian Chen. Veo. 2024. 11 [33] Chenyang Si, Ziqi Huang, Yuming Jiang, and Ziwei Liu. Freeu: Free lunch in diffusion u-net. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4733-4743, 2024. [34] Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"pattern recognition, pages 10684-10695, 2022. [31] Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Sali- mans, et al. Photorealistic text-to-image diffusion mod- els with deep language understanding.Advances in neu- ral information processing systems, 35:36479-36494, 2022. [32] Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"its exceptional capability in modeling plausible spatial appearance and coherent temporal dynamics. References [1] Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models.arXiv preprint arXiv:2210.02303, 2022. [2] Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Make-A-Video: Text-to-Video Generation without Text-Video Data because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (28 contexts).","role_counts":[{"n":28,"context_role":"background"},{"n":2,"context_role":"baseline"}]},"error":null,"updated_at":"2026-06-29T17:29:10.361182+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Make-A-Video: Text-to-Video Generation without Text-Video Data","claims":[{"claim_text":"We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from paired text-image data, and learn how the world moves from unsupervised video footage. Make-A-Video has three advantages: (1) it accelerates training of the T2V model (it does not need to learn visual and multimodal representations from scratch), (2) it does not require paired text-video data, and (3) the generated videos inherit the vastness (diversity in ae","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Make-A-Video: Text-to-Video Generation without Text-Video Data because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T17:06:38.827093+00:00"}},"summary":{"title":"Make-A-Video: Text-to-Video Generation without Text-Video Data","claims":[{"claim_text":"We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from paired text-image data, and learn how the world moves from unsupervised video footage. Make-A-Video has three advantages: (1) it accelerates training of the T2V model (it does not need to learn visual and multimodal representations from scratch), (2) it does not require paired text-video data, and (3) the generated videos inherit the vastness (diversity in ae","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Make-A-Video: Text-to-Video Generation without Text-Video Data because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Imagen Video: High Definition Video Generation with Diffusion Models","work_id":"bb20d241-dc6f-4b0a-b071-fd43a2cbd57f","shared_citers":22},{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":21},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":17},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":14},{"title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","work_id":"1f9d1d3b-a6d6-45a9-9f13-51393c03be8a","shared_citers":13},{"title":"Score-Based Generative Modeling through Stochastic Differential Equations","work_id":"d9110e53-a5d4-4794-a4c5-a575e91c31ad","shared_citers":13},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":12},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":11},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":11},{"title":"CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers","work_id":"2dbd6bcd-fc98-4fbf-b586-f6d94fe1abd2","shared_citers":10},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":10},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":10},{"title":"Magicvideo: Efficient video generation with latent diffusion models","work_id":"aad71b40-2721-438d-8e8c-97f84063ed39","shared_citers":9},{"title":"Latte: Latent Diffusion Transformer for Video Generation","work_id":"5328e907-7278-4781-a2bb-c5ef40dc87fb","shared_citers":7},{"title":"LoRA: Low-Rank Adaptation of Large Language Models","work_id":"0426219a-789e-4964-adc8-a04538510818","shared_citers":7},{"title":"ModelScope Text-to-Video Technical Report","work_id":"1b1baf78-58ec-44d0-b700-84dff57b2f1f","shared_citers":7},{"title":"An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion","work_id":"ca618c21-3ba6-448e-bd86-bcecff3cdeb5","shared_citers":6},{"title":"arXiv:2210.02399 , year=","work_id":"a325cd53-6549-4726-b3e9-94509f0df168","shared_citers":6},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":6},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":6},{"title":"GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models","work_id":"34430d19-7919-48ce-88a5-17b3bfe2192e","shared_citers":6},{"title":"High-Resolution Image Synthesis with Latent Diffusion Models","work_id":"f0270d36-2952-47fb-84c1-95e3ec341126","shared_citers":6},{"title":"Movie Gen: A Cast of Media Foundation Models","work_id":"a6a118b0-002f-4b19-881f-7f1183e0d7d8","shared_citers":6},{"title":"Open-Sora: Democratizing Efficient Video Production for All","work_id":"8b29ba7b-3d84-4281-85b7-9eaf905afd7f","shared_citers":6}],"time_series":[{"n":1,"year":2022},{"n":6,"year":2023},{"n":4,"year":2024},{"n":33,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"9f5d8ace-8e4d-427a-ba25-63c1aef83482","orcid":null,"display_name":"Adam Polyak","source":"manual","import_confidence":0.72},{"id":"798e0a0e-767c-4ee2-9fd3-429c27455832","orcid":null,"display_name":"Jie An","source":"manual","import_confidence":0.72},{"id":"99b4149d-99b7-4361-aa15-ba5557443e1f","orcid":null,"display_name":"Songyang Zhang","source":"manual","import_confidence":0.72},{"id":"cd1b678c-b258-4e70-b5c6-dca9b58195a7","orcid":null,"display_name":"Thomas Hayes","source":"manual","import_confidence":0.72},{"id":"46983c5e-881b-4d61-bfda-f727b70125ab","orcid":null,"display_name":"Uriel Singer","source":"manual","import_confidence":0.72},{"id":"6e4bc21b-4bc4-4b60-b6f3-2f13064aac7f","orcid":null,"display_name":"Xi Yin","source":"manual","import_confidence":0.72}]}}