{"work":{"id":"8b29ba7b-3d84-4281-85b7-9eaf905afd7f","openalex_id":"https://openalex.org/W4405956131","doi":"10.48550/arxiv.2412.20404","arxiv_id":"2412.20404","raw_key":null,"title":"Open-Sora: Democratizing Efficient Video Production for All","authors":null,"authors_text":"Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu","year":2024,"venue":"cs.CV","abstract":"Vision and language are the two foundational senses for humans, and they build up our cognitive ability and intelligence. While significant breakthroughs have been made in AI language ability, artificial visual intelligence, especially the ability to generate and simulate the world we see, is far lagging behind. To facilitate the development and accessibility of artificial visual intelligence, we created Open-Sora, an open-source video generation model designed to produce high-fidelity video content. Open-Sora supports a wide spectrum of visual generation tasks, including text-to-image generation, text-to-video generation, and image-to-video generation. The model leverages advanced deep learning architectures and training/inference techniques to enable flexible video synthesis, which could generate video content of up to 15 seconds, up to 720p resolution, and arbitrary aspect ratios. Specifically, we introduce Spatial-Temporal Diffusion Transformer (STDiT), an efficient diffusion framework for videos that decouples spatial and temporal attention. We also introduce a highly compressive 3D autoencoder to make representations compact and further accelerate training with an ad hoc training strategy. Through this initiative, we aim to foster innovation, creativity, and inclusivity within the community of AI content creation. By embracing the open-source principle, Open-Sora democratizes full access to all the training/inference/data preparation codes as well as model weights. All resources are publicly available at: https://github.com/hpcaitech/Open-Sora.","external_url":"https://arxiv.org/abs/2412.20404","cited_by_count":10,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2412.20404","created_at":"2026-05-09T06:05:34.489654+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Open-Sora: Democratizing Efficient Video Production for All","render_title":"Open-Sora: Democratizing Efficient Video Production for All"},"hub":{"state":{"work_id":"8b29ba7b-3d84-4281-85b7-9eaf905afd7f","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":133,"external_cited_by_count":10,"distinct_field_count":11,"first_pith_cited_at":"2025-01-23T18:55:41+00:00","last_pith_cited_at":"2026-07-09T17:59:52+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-20T17:59:46.154672+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":17},{"context_role":"baseline","n":3},{"context_role":"dataset","n":2}],"polarity_counts":[{"context_polarity":"background","n":17},{"context_polarity":"baseline","n":3},{"context_polarity":"use_dataset","n":2}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Open-Sora: Democratizing Efficient Video Production for All","claims":[{"claim_text":"Vision and language are the two foundational senses for humans, and they build up our cognitive ability and intelligence. While significant breakthroughs have been made in AI language ability, artificial visual intelligence, especially the ability to generate and simulate the world we see, is far lagging behind. To facilitate the development and accessibility of artificial visual intelligence, we created Open-Sora, an open-source video generation model designed to produce high-fidelity video content. Open-Sora supports a wide spectrum of visual generation tasks, including text-to-image generat","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"would be semantically appropriate for a given scene. Fol- lowing this, ConsistI2V [61] enhances visual consistency through spatiotemporal attention mechanisms, while Emu Video [11] factorizes video generation into separate T2I and I2V stages. Moreover, VideoCrafter1 [62] and VideoCrafter2 [63] provide toolkits supporting both text and image conditions. For large models, Open-Sora [64] and Open-Sora Plan [65] demonstrate that DiT can in- tegrate text conditioning seamlessly across billions of par","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"namic video generation, physical constraints become increasingly complex and important. 9 Fig. 6: Video Generations by GPT4Motion (Figure courtesy of [163]). TABLE 3: Performance of Video Generation Mod- els on representative physical and world modeling benchmarks. Model PhysicsIQ [164](↑) PhyGen [165](↑) VideoPhy [166](↑) WorldModel Bench [167](↑) Sora [168] 0.10 0.44 0.28 6.11 Pika [169] 0.13 0.44 0.29 - CogVideoX [170] - 0.45 0.49 7.31 LaVie [171] - 0.36 0.41 - Kling [172] - 0.49 - 8.82 Moder","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[45] Tianchen Zhao, Tongcheng Fang, Haofeng Huang, Enshu Liu, Rui Wan, Widyadewi Soedarmadji, Shiyao Li, Zinan Lin, Guohao Dai, Shengen Yan, et al. Vidit-q: Efficient and accurate quantization of diffusion transformers for image and video generation.arXiv preprint arXiv:2406.02540, 2024. [46] Xuanlei Zhao, Xiaolong Jin, Kai Wang, and Yang You. Real-time video generation with pyramid attention broadcast. arXiv preprint arXiv:2408.12588, 2024. [47] Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"VE-Bench[17] 169 1,170 6✓ ✓ ✗8 SD-based open-source (2024) EditBoard[13] - - 4✗- - - FiVE[14]∼100 420 6✗- - - OpenVE-3M[16] 1M 3M 8✓ ✗ ✓Open-source + agentic (2025) IVE-Bench[15] 600 - 8/35✗- - - VEFX-Dataset(Ours) 1,988 5,049 9/32 ✓ ✓ ✓ 4 Commercial + Open + Agentic (2026) 3.1 Data Collection Source videos.We curate source videos from open-source video datasets including Open-Sora [ 39] and OpenVid-1M [40], supplemented with privately collected footage for additional diversity. We filter the in","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ViVa: A Video-Generative Value Model for Robot Reinforcement Learning [57] Hongxiang Zhao, Xingchen Liu, Mutian Xu, Yiming Hao, Weikai Chen, and Xiaoguang Han. Taste- rob: Advancing video generation of task-oriented hand-object interaction for generalizable robotic manipulation. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 27683- 27693, 2025. 4 [58] Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[106] Zhixing Zhang, Yanyu Li, Yushu Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, Junli Cao, Dimitris Metaxas, Sergey Tulyakov, et al. Sf-v: Single forward video generation model. In Advances in Neural Information Processing Systems (NeurIPS), 2024. [107] Guangcong Zheng, Teng Li, Rui Jiang, Yehao Lu, Tao Wu, and Xi Li. Cami2v: Camera-controlled image- to-video diffusion model.arXiv preprint arXiv:2410.15957, 2024. [108] Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui S","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Open-Sora: Democratizing Efficient Video Production for All because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (17 contexts).","role_counts":[{"n":17,"context_role":"background"},{"n":3,"context_role":"baseline"},{"n":2,"context_role":"dataset"}]},"error":null,"updated_at":"2026-07-02T01:02:08.753460+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"43f7e8f1-2328-4e2f-adff-f48b2517c568","orcid":null,"display_name":"Zangwei Zheng"},{"id":"8535a87b-097b-4044-9fec-f1371c829a1a","orcid":null,"display_name":"Xiangyu Peng"},{"id":"bec1d02e-5191-4c86-ae9e-4d4c921691bb","orcid":null,"display_name":"Tianji Yang"},{"id":"8a6126ef-663c-47dc-a1ab-8071d8ca70db","orcid":null,"display_name":"Chenhui Shen"},{"id":"218b4ef5-3d11-4830-9dd8-f9ff40922808","orcid":null,"display_name":"Shenggui Li"},{"id":"025649eb-d371-44a4-92ea-11500891c6a1","orcid":null,"display_name":"Hongxin Liu"}]},"error":null,"updated_at":"2026-07-02T01:02:09.730802+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T17:59:40.890635+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":28},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":23},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":17},{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":14},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":9},{"title":"Movie Gen: A Cast of Media Foundation Models","work_id":"a6a118b0-002f-4b19-881f-7f1183e0d7d8","shared_citers":8},{"title":"LTX-Video: Realtime Video Latent Diffusion","work_id":"cee5c521-3ce9-466e-a035-1e42f89254f4","shared_citers":7},{"title":"arXiv preprint arXiv:2504.13074 (2025) 4 16 Xingtong Ge et al","work_id":"2ce11350-273e-4f0d-ae78-292aa3151060","shared_citers":6},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":6},{"title":"Make-A-Video: Text-to-Video Generation without Text-Video Data","work_id":"52a801fc-a707-45a1-a8cd-0d6702f124ab","shared_citers":6},{"title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","work_id":"1f9d1d3b-a6d6-45a9-9f13-51393c03be8a","shared_citers":5},{"title":"arXiv preprint arXiv:2406.01125 (2024) 3, 4","work_id":"a9eb5bf3-fc76-45d5-9721-87061c21a0ba","shared_citers":5},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":5},{"title":"Imagen Video: High Definition Video Generation with Diffusion Models","work_id":"bb20d241-dc6f-4b0a-b071-fd43a2cbd57f","shared_citers":5},{"title":"MAGI-1: Autoregressive Video Generation at Scale","work_id":"25e8bd3d-e51c-43ae-8126-4ea6ecdb3321","shared_citers":5},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":5},{"title":"Seedance 1.0: Exploring the Boundaries of Video Generation Models","work_id":"b2e36b5d-99e4-45b4-9358-64f6d3501983","shared_citers":5},{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":4},{"title":"$\\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization","work_id":"d1ad7304-d09a-49bc-809e-846439f6aff9","shared_citers":4},{"title":"arXiv preprint arXiv:2503.09642 (2025)","work_id":"c22f9060-268c-4cc9-8018-dee486a23da1","shared_citers":4},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":4},{"title":"CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers","work_id":"2dbd6bcd-fc98-4fbf-b586-f6d94fe1abd2","shared_citers":4},{"title":"Cosmos World Foundation Model Platform for Physical AI","work_id":"a2dba24c-318d-476a-8b21-4289c265810c","shared_citers":4},{"title":"Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference","work_id":"53b1d836-7feb-402c-97c7-87b9bc51196f","shared_citers":4}],"time_series":[{"n":1,"year":2025},{"n":37,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T17:59:19.576291+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T18:00:02.728062+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Open-Sora: Democratizing Efficient Video Production for All","claims":[{"claim_text":"Vision and language are the two foundational senses for humans, and they build up our cognitive ability and intelligence. While significant breakthroughs have been made in AI language ability, artificial visual intelligence, especially the ability to generate and simulate the world we see, is far lagging behind. To facilitate the development and accessibility of artificial visual intelligence, we created Open-Sora, an open-source video generation model designed to produce high-fidelity video content. Open-Sora supports a wide spectrum of visual generation tasks, including text-to-image generat","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"would be semantically appropriate for a given scene. Fol- lowing this, ConsistI2V [61] enhances visual consistency through spatiotemporal attention mechanisms, while Emu Video [11] factorizes video generation into separate T2I and I2V stages. Moreover, VideoCrafter1 [62] and VideoCrafter2 [63] provide toolkits supporting both text and image conditions. For large models, Open-Sora [64] and Open-Sora Plan [65] demonstrate that DiT can in- tegrate text conditioning seamlessly across billions of par","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"namic video generation, physical constraints become increasingly complex and important. 9 Fig. 6: Video Generations by GPT4Motion (Figure courtesy of [163]). TABLE 3: Performance of Video Generation Mod- els on representative physical and world modeling benchmarks. Model PhysicsIQ [164](↑) PhyGen [165](↑) VideoPhy [166](↑) WorldModel Bench [167](↑) Sora [168] 0.10 0.44 0.28 6.11 Pika [169] 0.13 0.44 0.29 - CogVideoX [170] - 0.45 0.49 7.31 LaVie [171] - 0.36 0.41 - Kling [172] - 0.49 - 8.82 Moder","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[45] Tianchen Zhao, Tongcheng Fang, Haofeng Huang, Enshu Liu, Rui Wan, Widyadewi Soedarmadji, Shiyao Li, Zinan Lin, Guohao Dai, Shengen Yan, et al. Vidit-q: Efficient and accurate quantization of diffusion transformers for image and video generation.arXiv preprint arXiv:2406.02540, 2024. [46] Xuanlei Zhao, Xiaolong Jin, Kai Wang, and Yang You. Real-time video generation with pyramid attention broadcast. arXiv preprint arXiv:2408.12588, 2024. [47] Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"VE-Bench[17] 169 1,170 6✓ ✓ ✗8 SD-based open-source (2024) EditBoard[13] - - 4✗- - - FiVE[14]∼100 420 6✗- - - OpenVE-3M[16] 1M 3M 8✓ ✗ ✓Open-source + agentic (2025) IVE-Bench[15] 600 - 8/35✗- - - VEFX-Dataset(Ours) 1,988 5,049 9/32 ✓ ✓ ✓ 4 Commercial + Open + Agentic (2026) 3.1 Data Collection Source videos.We curate source videos from open-source video datasets including Open-Sora [ 39] and OpenVid-1M [40], supplemented with privately collected footage for additional diversity. We filter the in","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ViVa: A Video-Generative Value Model for Robot Reinforcement Learning [57] Hongxiang Zhao, Xingchen Liu, Mutian Xu, Yiming Hao, Weikai Chen, and Xiaoguang Han. Taste- rob: Advancing video generation of task-oriented hand-object interaction for generalizable robotic manipulation. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 27683- 27693, 2025. 4 [58] Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[106] Zhixing Zhang, Yanyu Li, Yushu Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, Junli Cao, Dimitris Metaxas, Sergey Tulyakov, et al. Sf-v: Single forward video generation model. In Advances in Neural Information Processing Systems (NeurIPS), 2024. [107] Guangcong Zheng, Teng Li, Rui Jiang, Yehao Lu, Tao Wu, and Xi Li. Cami2v: Camera-controlled image- to-video diffusion model.arXiv preprint arXiv:2410.15957, 2024. [108] Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui S","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Open-Sora: Democratizing Efficient Video Production for All because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (17 contexts).","role_counts":[{"n":17,"context_role":"background"},{"n":3,"context_role":"baseline"},{"n":2,"context_role":"dataset"}]},"error":null,"updated_at":"2026-07-02T01:02:08.747818+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Open-Sora: Democratizing Efficient Video Production for All","claims":[{"claim_text":"Vision and language are the two foundational senses for humans, and they build up our cognitive ability and intelligence. While significant breakthroughs have been made in AI language ability, artificial visual intelligence, especially the ability to generate and simulate the world we see, is far lagging behind. To facilitate the development and accessibility of artificial visual intelligence, we created Open-Sora, an open-source video generation model designed to produce high-fidelity video content. Open-Sora supports a wide spectrum of visual generation tasks, including text-to-image generat","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Open-Sora: Democratizing Efficient Video Production for All because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:00:11.368497+00:00"}},"summary":{"title":"Open-Sora: Democratizing Efficient Video Production for All","claims":[{"claim_text":"Vision and language are the two foundational senses for humans, and they build up our cognitive ability and intelligence. While significant breakthroughs have been made in AI language ability, artificial visual intelligence, especially the ability to generate and simulate the world we see, is far lagging behind. To facilitate the development and accessibility of artificial visual intelligence, we created Open-Sora, an open-source video generation model designed to produce high-fidelity video content. Open-Sora supports a wide spectrum of visual generation tasks, including text-to-image generat","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Open-Sora: Democratizing Efficient Video Production for All because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":28},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":23},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":17},{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":14},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":9},{"title":"Movie Gen: A Cast of Media Foundation Models","work_id":"a6a118b0-002f-4b19-881f-7f1183e0d7d8","shared_citers":8},{"title":"LTX-Video: Realtime Video Latent Diffusion","work_id":"cee5c521-3ce9-466e-a035-1e42f89254f4","shared_citers":7},{"title":"arXiv preprint arXiv:2504.13074 (2025) 4 16 Xingtong Ge et al","work_id":"2ce11350-273e-4f0d-ae78-292aa3151060","shared_citers":6},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":6},{"title":"Make-A-Video: Text-to-Video Generation without Text-Video Data","work_id":"52a801fc-a707-45a1-a8cd-0d6702f124ab","shared_citers":6},{"title":"AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning","work_id":"1f9d1d3b-a6d6-45a9-9f13-51393c03be8a","shared_citers":5},{"title":"arXiv preprint arXiv:2406.01125 (2024) 3, 4","work_id":"a9eb5bf3-fc76-45d5-9721-87061c21a0ba","shared_citers":5},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":5},{"title":"Imagen Video: High Definition Video Generation with Diffusion Models","work_id":"bb20d241-dc6f-4b0a-b071-fd43a2cbd57f","shared_citers":5},{"title":"MAGI-1: Autoregressive Video Generation at Scale","work_id":"25e8bd3d-e51c-43ae-8126-4ea6ecdb3321","shared_citers":5},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":5},{"title":"Seedance 1.0: Exploring the Boundaries of Video Generation Models","work_id":"b2e36b5d-99e4-45b4-9358-64f6d3501983","shared_citers":5},{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":4},{"title":"$\\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization","work_id":"d1ad7304-d09a-49bc-809e-846439f6aff9","shared_citers":4},{"title":"arXiv preprint arXiv:2503.09642 (2025)","work_id":"c22f9060-268c-4cc9-8018-dee486a23da1","shared_citers":4},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":4},{"title":"CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers","work_id":"2dbd6bcd-fc98-4fbf-b586-f6d94fe1abd2","shared_citers":4},{"title":"Cosmos World Foundation Model Platform for Physical AI","work_id":"a2dba24c-318d-476a-8b21-4289c265810c","shared_citers":4},{"title":"Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference","work_id":"53b1d836-7feb-402c-97c7-87b9bc51196f","shared_citers":4}],"time_series":[{"n":1,"year":2025},{"n":37,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"8a6126ef-663c-47dc-a1ab-8071d8ca70db","orcid":null,"display_name":"Chenhui Shen","source":"manual","import_confidence":0.72},{"id":"025649eb-d371-44a4-92ea-11500891c6a1","orcid":null,"display_name":"Hongxin Liu","source":"manual","import_confidence":0.72},{"id":"218b4ef5-3d11-4830-9dd8-f9ff40922808","orcid":null,"display_name":"Shenggui Li","source":"manual","import_confidence":0.72},{"id":"bec1d02e-5191-4c86-ae9e-4d4c921691bb","orcid":null,"display_name":"Tianji Yang","source":"manual","import_confidence":0.72},{"id":"8535a87b-097b-4044-9fec-f1371c829a1a","orcid":null,"display_name":"Xiangyu Peng","source":"manual","import_confidence":0.72},{"id":"43f7e8f1-2328-4e2f-adff-f48b2517c568","orcid":null,"display_name":"Zangwei Zheng","source":"manual","import_confidence":0.72}]}}