Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:52:24.986017Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2505.20827.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:52:24.986017Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:33.443014Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T16:39:34.724653Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 326a447c-60ba-46d3-a187-c57f939e92e8 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lumiere: A space-time diffusion model for video generation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d58735dd-48a3-4813-89ea-9807fa66f72b · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a57090df-d86f-46bc-bfce-6382860c7b8d · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Align your latents: High-resolution video synthesis with latent diffusion models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d51cc5a-4ff0-4606-a8f0-7645f5a55768 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Video generation models as world simulators
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44bba26c-5ef4-4aac-a2ab-af2857f75a54 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Activitynet: A large-scale video benchmark for human activity understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3721de0-7be4-461d-a563-a61dbd409232 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Diffusion forcing: Next-token prediction meets full-sequence diffusion
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3038fb96-a49c-4bb0-a81f-681711bf0412 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes SkyReels-V2: Infinite-length Film Generative Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2cd86fe-fcb3-4001-a2cc-2ac94e8811e4 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2f7a23e-1389-4327-b6fa-4601eda84a77 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Learning temporal coherence via self-supervision for gan-based video generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15b9771c-e22f-4231-adeb-d36464455b75 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Factorizing text-to-video generation by explicit image conditioning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa7f08fb-4ff9-4ff0-b7f1-7d20d9a4f1f1 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b625c430-1f3c-43d7-ab62-2dd8f00864fd · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Long Context Tuning for Video Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80296d13-3267-4b87-99b6-ede1ccec4070 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 438697a2-2e0d-43c2-b8ab-abc715c43277 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Autoregressive diffusion models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b9fdd70-f981-4764-a3ae-3e19a022e48e · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b13f9e56-6b9f-4ea2-a976-c9b6e453d8ee · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes FIFO-Diffusion: Generating Infinite Videos from Text without Training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12f45af-4f3b-4d6f-9679-3cce42ef9b65 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a307adf3-e2ac-4d93-8f34-e9cdbdb6e240 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes A Survey on Long Video Generation: Challenges, Methods, and Prospects
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4852276-2f62-44ff-88f6-1bece02499fe · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Unified Video Action Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ecca3c2-367f-4ce3-a333-0d009eac451e · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Open-Sora Plan: Open-Source Large Video Generation Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a076bb32-7a26-45d0-a597-80329400b0b5 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Videostudio: Generating consistent-content and multi-scene videos
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8eed2f6-a220-4570-a159-41bfa82eb6c3 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e77a915f-003e-4247-86e2-ad20030aba1a · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mevg: Multi-event video generation with text-to-video models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9b86993-63dc-4c26-ba95-210aa7a8b470 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Scalable diffusion models with transformers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f227f7fc-c4cc-4f16-ad55-e6e0e7674011 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cb4abcb-3de5-4d37-a93b-5e1a93019b13 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Freenoise: Tuning-free longer video diffusion via noise rescheduling
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c69b16a1-53d6-411e-920f-de6d91f91e68 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70b1d677-ea3a-4e73-802a-d1c568ef8139 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 212a9cf1-50d9-48f8-bbad-7a0f8af8487d · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lightweight, Pre-trained Transformers for Remote Sensing Timeseries
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1d42ff8-58c3-4633-ba3c-39b2f7a4c21f · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mocogan: Decomposing motion and content for video generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55bd74b1-3bf0-40f5-9594-e1a9d0f0218d · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Wan: Open and Advanced Large-Scale Video Generative Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f999d2fb-6d40-4027-9760-581b287740b2 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5bea886-b2d0-455b-af90-1c8903bcb576 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5e26b97-9690-40db-85af-400913c64a0e · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ffd8aa8-5ad0-40cc-a4ae-1faf228dae00 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lvbench: An extreme long video understanding benchmark, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7246b089-de6e-49b3-9bd5-6dd77cde997e · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b4f6d54-734f-4648-98e3-9798e16b796c · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83bf8499-8c7c-4f69-83f8-60277cbbbae4 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Imaginator: Condi- tional spatio-temporal gan for video generation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e9afcf7-8ae4-48b7-97b1-71253170f820 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Advancing high-resolution video-language representation with large-scale video transcriptions
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40835e76-76e2-4297-bbc0-8aeaf401f988 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1861851a-92b6-47e7-94ad-99d54fc6287e · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68c7dbf-cef9-4122-9acc-69273d4938b4 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Merlot: Multimodal neural script knowledge models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4cd00ff3-e07a-4ac9-b454-5e03a94706d1 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Moviedreamer: Hierarchical generation for coherent long visual sequence
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfaca529-075e-4c00-81e6-caaaf624d3f9 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a69e4fa-55ef-4a37-bc3e-75fd0c59ff2a · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Videogen-of-thought: A collaborative framework for multi-shot video generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63da9f05-e6a4-45aa-a677-27469349e31c · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Towards automatic learning of procedures from web instructional videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 614f1169-0a76-4a97-82d3-70953de36695 · outbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes Storydiffusion: Consistent self-attention for long-range image and video generation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 672da83a-bf59-448e-b9c2-ffe3a00488e4 · inbound
LoViC: Efficient Long Video Generation with Context Compression Frame-Level Captions for Long Video Generation with Complex Multi Scenes
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.