{"work":{"id":"2ce11350-273e-4f0d-ae78-292aa3151060","openalex_id":null,"doi":null,"arxiv_id":"2504.13074","raw_key":null,"title":"SkyReels-V2: Infinite-length Film Generative Model","authors":null,"authors_text":"Guibin Chen, Dixuan Lin, Jiangping Yang, Chunze Lin, Junchen Zhu, Mingyuan Fan","year":2025,"venue":"cs.CV","abstract":"Recent advances in video generation have been driven by diffusion models and autoregressive frameworks, yet critical challenges persist in harmonizing prompt adherence, visual quality, motion dynamics, and duration: compromises in motion dynamics to enhance temporal visual quality, constrained video duration (5-10 seconds) to prioritize resolution, and inadequate shot-aware generation stemming from general-purpose MLLMs' inability to interpret cinematic grammar, such as shot composition, actor expressions, and camera motions. These intertwined limitations hinder realistic long-form synthesis and professional film-style generation. To address these limitations, we propose SkyReels-V2, an Infinite-length Film Generative Model, that synergizes Multi-modal Large Language Model (MLLM), Multi-stage Pretraining, Reinforcement Learning, and Diffusion Forcing Framework. Firstly, we design a comprehensive structural representation of video that combines the general descriptions by the Multi-modal LLM and the detailed shot language by sub-expert models. Aided with human annotation, we then train a unified Video Captioner, named SkyCaptioner-V1, to efficiently label the video data. Secondly, we establish progressive-resolution pretraining for the fundamental video generation, followed by a four-stage post-training enhancement: Initial concept-balanced Supervised Fine-Tuning (SFT) improves baseline quality; Motion-specific Reinforcement Learning (RL) training with human-annotated and synthetic distortion data addresses dynamic artifacts; Our diffusion forcing framework with non-decreasing noise schedules enables long-video synthesis in an efficient search space; Final high-quality SFT refines visual fidelity. All the code and models are available at https://github.com/SkyworkAI/SkyReels-V2.","external_url":"https://arxiv.org/abs/2504.13074","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-10T01:46:41.058772+00:00","pith_arxiv_id":"2504.13074","created_at":"2026-05-09T06:05:34.424231+00:00","updated_at":"2026-07-10T01:46:41.058772+00:00","title_quality_ok":true,"display_title":"SkyReels-V2: Infinite-length Film Generative Model","render_title":"SkyReels-V2: Infinite-length Film Generative Model"},"hub":{"state":{"work_id":"2ce11350-273e-4f0d-ae78-292aa3151060","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":77,"external_cited_by_count":null,"distinct_field_count":4,"first_pith_cited_at":"2025-05-08T17:58:45+00:00","last_pith_cited_at":"2026-07-09T17:59:11+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T08:09:25.838729+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":13},{"context_role":"baseline","n":4},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":13},{"context_polarity":"baseline","n":4},{"context_polarity":"use_method","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}