Pith. sign in

Paper Citation Record · LEDGER

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality

As of 22 August 2026, this Paper Citation Record lists 100 of 110 outbound references and 1 inbound Pith citation observation for arXiv:2507.07202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07202 v1

Coverage vector

measured 100 of 110 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:51:31.757497Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:41:42.111186Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 110 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b94a691a-eb77-415d-84b9-0688d3d65428 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.002704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.002704Z digest=sha256:05ed725433dd65cf2da8c6b299a43f251c3ceea381308feed0f5318983a084a0

Observation e31a678e-b35e-4cab-af6a-b1a2482c01da · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.068306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.068306Z digest=sha256:c1aeca9c45c7913c4475ed5e12b894266e3a694371327cfede7ce1ec58671562

Observation a3e5038a-6c1b-4730-8a7d-dee5bdec4704 · outbound

This paper cites Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Ma- chine Learning (ICML), pages 813–824.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Ma- chine Learning (ICML), pages 813–824

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.169776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.169776Z digest=sha256:6ef54a245847f2c51e9c1d26306c160619339ca27369decc1816fb235fead07e

Observation e8e5648f-1b7b-4433-85bb-64a28f7f72c9 · outbound

This paper cites Generating Long Videos of Dynamic Scenes.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Generating Long Videos of Dynamic Scenes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.241885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.241885Z digest=sha256:4b0293db7ff6db117f57f9641bc25efe9503522a5e9431e4e8c19f236ec8b2f8

Observation 157891b2-3d5e-4a19-b32e-a82f76fdbc3b · outbound

This paper cites an unresolved cited work.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.367980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.367980Z digest=sha256:a66300e69f84ceaf01082d71e651474dfd9e508c188c5f2d3e3642d08839616d

Observation a939f0a1-a0a5-4ace-ab90-69434ac45f4b · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality SkyReels-V2: Infinite-length Film Generative Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.412898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.412898Z digest=sha256:0751998325172aca42497166881e9ce4485649ddc3205d3001a7dd0fca13683a

Observation 5b10905a-0c5e-4e86-8436-51d7b60dede4 · outbound

This paper cites Pixart- α: Fast train- ing of diffusion transformer for photorealistic text-to-image synthesis, 2023.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Pixart- α: Fast train- ing of diffusion transformer for photorealistic text-to-image synthesis, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.497048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.497048Z digest=sha256:0eedc67a5110535cd039502edcf3f76486ddb448cbf0a022ca1f4a6477bd461c

Observation 9b81b3cc-7a17-45a0-be57-63b4dca53bfd · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.568977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.568977Z digest=sha256:fe9165ed567d15c74ae528aa82f64764b076caf6ec9099ecd57d422ad4e2d5ce

Observation 4ae97d87-06df-42d8-83df-1ca90836f6c6 · outbound

This paper cites Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.668995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.668995Z digest=sha256:e9b7b3e1aeee6f88962de2551398af9bc3cebe9dda2722b962ed01d1ccb8dc36

Observation da6cfe17-97e9-49b9-9301-a12ff7e27cb2 · outbound

This paper cites Multi-subject open-set personalization in video generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Multi-subject open-set personalization in video generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.735298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.735298Z digest=sha256:df4b7aeea20ed73eba46a971d1ab0c7111260ed56e093e09de3642438d7f3c29

Observation 1b6bbcf0-5e04-4977-adf3-2be0877510d1 · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.807647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.807647Z digest=sha256:ffb496bdc8501b2473ca70d561b6571b5b383e5b12170f5d448d646d6e048e29

Observation 79609944-7dee-4bd5-aeee-e6dcfefb663d · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.893203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.893203Z digest=sha256:d7b74833e679a837ec071a9c22be8bd13c2d8e350f5e5c7a3fbd408039abcaa1

Observation 97c88819-f091-44d0-80e6-7f03cec56924 · outbound

This paper cites Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.996882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.996882Z digest=sha256:93fcff3c719b25977eb146f5eccc9c8031683c040d6cf32335d636ea99226018

Observation dec23150-c339-4940-89cf-4e4bcef08eec · outbound

This paper cites Zhao, Yanping Huang, An- drew M.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Zhao, Yanping Huang, An- drew M

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.096771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.096771Z digest=sha256:749640bc026fb1521cc073ee973b33005d97d4607ff702f166e3166c049b21d3

Observation 8c3ace60-0f84-4185-97fc-77933b4cf28c · outbound

This paper cites Efficient video prediction via sparsely conditioned flow matching.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Efficient video prediction via sparsely conditioned flow matching

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.165135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.165135Z digest=sha256:7b1ff52be942ee7c45776c55f75dcddb4171006999fc3371e8d97ba3d0417bbd

Observation 09d01f88-6286-4db2-a306-09979a1da62b · outbound

This paper cites an unresolved cited work.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.262633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.262633Z digest=sha256:5e6a6b53b958be97b56a9c1d23410eb8e53551694859b7829f8990cc3511b819

Observation 467800da-d809-462b-929f-532a496a77f5 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Scaling rectified flow transformers for high-resolution image synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.343816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.343816Z digest=sha256:648e9e90f175a005166d409aed3a582e8d82666b203bf1ecc3d37fe18c0f67c8

Observation 01543e06-f156-483f-9450-0f16dc4bce26 · outbound

This paper cites Ca2-vdm: Efficient autoregres- sive video diffusion model with causal generation and cache sharing.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Ca2-vdm: Efficient autoregres- sive video diffusion model with causal generation and cache sharing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.427361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.427361Z digest=sha256:d301bed43d33a8f1d02dcd4dea2ac9b6ec9d29697593fbf6fbe23a6a98e9ae8a

Observation 8c26b3bf-c206-4a97-b9b7-d19f0d5f3a80 · outbound

This paper cites Berg, Arash Vah- dat, Alexei A.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Berg, Arash Vah- dat, Alexei A

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.487912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.487912Z digest=sha256:c13860c0aff277357b40ffa13e9cad803629c4772b13ecc456038fb902ea4ee6

Observation 12a2a0a7-d6a8-42dc-9b9e-ce9c5b091005 · outbound

This paper cites Seedance 1.0: Exploring the Boundaries of Video Generation Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Seedance 1.0: Exploring the Boundaries of Video Generation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.538623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.538623Z digest=sha256:82a0d64efa470ae127781f61bbfb72010c253650f0381ee777d1d7bea6fda3c8

Observation 0f4120ff-7f4d-4994-96ff-2b597a561a2c · outbound

This paper cites Zico Kolter, and Kaiming He.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Zico Kolter, and Kaiming He

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.587961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.587961Z digest=sha256:a80e204db1d72dffa27b77bf78bbc936718f2697d59328455942f3fe15f846a8

Observation bd2b6afa-ceff-42da-9e82-3472f370d1d2 · outbound

This paper cites Veo 3: Neural video generation with native audio.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Veo 3: Neural video generation with native audio

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.668222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.668222Z digest=sha256:696109231e18c184f495c5e4cf20b72f916474596435f8c31bda99b2189327e3

Observation d5f63aa6-20b5-4b09-9f41-02afe4000d2e · outbound

This paper cites Animatediff: Animate your personalized text-to-image models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Animatediff: Animate your personalized text-to-image models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.762753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.762753Z digest=sha256:4c4eee37db582018d667838e234508396b94231d78cf972ac8a4d818ac0ddd4b

Observation 17d3516f-6d2e-4467-877e-ccfcaba77b2a · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality LTX-Video: Realtime Video Latent Diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.825780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.825780Z digest=sha256:03d1c6567e8e6c2cbb7e694d37d8bdb491e74ebb0b34a69e09118e5e89825fd1

Observation ab2eda46-0dcb-4226-a0ed-2acf84203285 · outbound

This paper cites Clipscore: A reference-free eval- uation metric for image captioning.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Clipscore: A reference-free eval- uation metric for image captioning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.897788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.897788Z digest=sha256:0fb051d4ba5a4b9695a7e09c25054dca50f777378901fd50f9433c47985b0bb7

Observation a8ecd40b-fb18-4ba6-89da-10c64e9d3d41 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.982396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.982396Z digest=sha256:36ded723e3253e265ac4755feba0ee91c6ae245700135cb4d66a56fd626819ba

Observation 0d792ce2-6701-4d6a-885e-5cace2211c75 · outbound

This paper cites Denoising dif- fusion probabilistic models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Denoising dif- fusion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.076771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.076771Z digest=sha256:9b24ca694e2b029f5e4ec263337d0907bf1ea8e13a76608e9aac8914c68eefad

Observation 724d2876-d37d-46f0-b8d8-783c86ffe333 · outbound

This paper cites an unresolved cited work.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.143548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.143548Z digest=sha256:a80cfc628c9a202917414a5f57930b36a6b0ea5cca4863b6eb81aeb4bd744ea0

Observation a2cc2f49-fcca-40a1-acd1-754b6f86eb82 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.211695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.211695Z digest=sha256:c650e378d4dc4246e5c5c072944e82fc82e0dfd3730064c58880f59ea5363fb4

Observation b0fe2f5a-e804-4281-9262-52d1888ca8e0 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality LoRA: Low-Rank Adaptation of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.294142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.294142Z digest=sha256:f222e7f42f2330671d08cf7f90b1c29bb18267842e5c5aa3239ede6cd8bae29a

Observation b641c452-348b-4b38-88ec-e82f018c0546 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.356204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.356204Z digest=sha256:464da4ae82a292caa60568c76146a4af9c62650dfbbe876a78423e807fc7d004

Observation ba01330a-351c-4cf4-8100-3014152fd200 · outbound

This paper cites Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.438943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.438943Z digest=sha256:782b0a0a18d67fe2f373ab36a8f876a211061a3bf4c71d2d9d822839dd9d2374

Observation 5b2e4943-7699-4f24-afe0-f875b9a9d6f2 · outbound

This paper cites ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.486990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.486990Z digest=sha256:299c477a6e4876991c136fc5fba450ce7ee05850a29ae385309eea535fffef00

Observation c90a055f-9636-4d81-99f3-beee7dc447ae · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative mod- els.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Vbench: Comprehensive benchmark suite for video generative mod- els

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.585805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.585805Z digest=sha256:4e9b36f4264a85a8dd882c7640231e2ae6865e3bffe6f39737c2d753316b31aa

Observation 003ec7ca-ed6a-4fdf-8a0c-61d3aa62ac68 · outbound

This paper cites Pika 1.5: Realistic scene and motion synthesis,.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Pika 1.5: Realistic scene and motion synthesis,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.659409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.659409Z digest=sha256:28e9184e725c78b2d5945b0259528d761f7a81d4b19bf05728d453da0b590bf3

Observation fc8dc39c-f6fd-4297-bc0a-4401d128302d · outbound

This paper cites Sim2real: Synthetic video data for vision-based robotic manipulation learning.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Sim2real: Synthetic video data for vision-based robotic manipulation learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.858784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.858784Z digest=sha256:c59b85dc948b5b5a6dea580c0ec94c7d0eeb63f49205610ec30f63754654c618

Observation d0724895-c789-4cf7-9ec5-6f9c48184a28 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality VACE: All-in-One Video Creation and Editing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.935841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.935841Z digest=sha256:8b8ab372543fa5508a48545035b51c550d0f87c4fb5e7a1e9464d0a161fa6df6

Observation a8063218-9d9c-4d40-adb9-5d7ac640a87a · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Pyramidal flow matching for efficient video generative modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.973195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.973195Z digest=sha256:c59c0a8890f04a08a6e1adf794e3f2464c1cb028a3fd772f3af01df81ca086ce

Observation 332aff63-e3e3-4168-9dfb-fdc6d5a7d6c2 · outbound

This paper cites Miradata: A large-scale video dataset with long du- rations and structured captions, 2024.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Miradata: A large-scale video dataset with long du- rations and structured captions, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.028318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.028318Z digest=sha256:36fd148399b6a361e1ee5070f76a9530c7e6d29545184c38a8f6cdc3953069ae

Observation f890aa14-7922-4460-ab83-fc509f308c55 · outbound

This paper cites Miradata: A large-scale video dataset with long 10 durations and structured captions.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Miradata: A large-scale video dataset with long 10 durations and structured captions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.111928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.111928Z digest=sha256:d7650622f0047688f4b851ec8db79dd4b2fc7088503510d9df9339f77eebedba

Observation c7f7f318-975b-44cb-8d4a-372deeb3bf44 · outbound

This paper cites Cinediff: Diffusion models for cinematic video synthesis.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Cinediff: Diffusion models for cinematic video synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.149466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.149466Z digest=sha256:731a3a52a327d81cd2298793b16cd1723aa1ba5facd0504d67bf52b90d03e6bd

Observation c7410d0e-6651-492e-90a3-7e67b9c510fe · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.228352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.228352Z digest=sha256:72bd8dac78c68ee90f3f2ddcdda80e5b704c2932a43481accf47fd725d55233f

Observation 99aad170-c98e-46ec-af89-91034147206d · outbound

This paper cites Minimax hailuo: Scalable multi- subject video generation, 2025.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Minimax hailuo: Scalable multi- subject video generation, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.336902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.336902Z digest=sha256:a3f92f2919f5b9f390a04eb391adcd31d8db72cc4c39ed5e44135b43f629d1d3

Observation 017b6017-8503-4488-b8fb-9c2f7364b022 · outbound

This paper cites Openhumanvid: A large-scale high- quality dataset for enhancing human-centric video genera- tion.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Openhumanvid: A large-scale high- quality dataset for enhancing human-centric video genera- tion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.606674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:27.417606Z digest=sha256:d6fdd03a622b6e95e38eb53587ecf1a2fb7627aaa5592de6351169e949d30bb1

Observation 57e461e5-e01b-4fa5-b908-3544ad5da534 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Open-Sora Plan: Open-Source Large Video Generation Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.464912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.464912Z digest=sha256:ba66d90eab9fd040aae9411ca9af1e3ccba5d4b4739498d864d4414fac0e01aa

Observation 75a20dc8-5800-4d8b-9cbc-dcbd3b4c8144 · outbound

This paper cites VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.528264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.528264Z digest=sha256:2be19db620176526ae3397ff390bf0904b1d7b605de77623c32ece973a3aa39a

Observation 1be0cb62-7ebe-440e-91cb-cf0afe8f4a05 · outbound

This paper cites an unresolved cited work.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:51:33.597009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:27.591944Z digest=sha256:f70e42a6ab9626b3c65a4fa2bc97e2e54ad617d1f150ba753fb0a73ac455b1e7

Observation 78f0d332-7e44-4678-bd3e-d7f2777dd5a5 · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Phantom: Subject-consistent video generation via cross-modal alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.649430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.649430Z digest=sha256:2b339f6af3f33b6ff474f5115f541e04f4d93af1a83904d28d39c8da30c8a56a

Observation a395f1b4-e549-42f3-a11f-044a5ac1f887 · outbound

This paper cites AutoStory: Generating Diverse Storytelling Images with Minimal Human Effort.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality AutoStory: Generating Diverse Storytelling Images with Minimal Human Effort

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:51:33.048674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:27.714173Z digest=sha256:82c6c03490ded67d09781068da49b929441a90cdd58b2d23d213c830d1541657

Observation 14c0c2e1-baad-4300-a957-4c7151215b45 · outbound

This paper cites Latte: Latent diffusion transformer for video generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Latte: Latent diffusion transformer for video generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.588141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:27.753527Z digest=sha256:4e422bbce2858739ec9c650fbb673f4b38a96d5086cc08b5829f59dac56fcc9f

Observation 424ebe18-0559-4982-bacb-f234c5cf7497 · outbound

This paper cites Gamegan: Video generation for atari games.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Gamegan: Video generation for atari games

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.578596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:27.811437Z digest=sha256:eea048181cd26e614ce61c269db2af2844d5b758abe8c2c633e8c77a2b47d1fb

Observation 780a0571-1608-4adf-a561-3341a31ab6e5 · outbound

This paper cites Sora: Openai’s text-to-video generator, 2024.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Sora: Openai’s text-to-video generator, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.568719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:27.857440Z digest=sha256:09f5e59708abbdd00753aebc64a4f1d0ff332b92b83fe81abf4b0d445066a10e

Observation a03c9693-0c04-4c64-af21-97ddf8c8aa87 · outbound

This paper cites Scalable diffusion mod- els with transformers.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Scalable diffusion mod- els with transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.558545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:27.901975Z digest=sha256:a6406579c2aeedfbb6d1584cd3e3b43b423055b1730e9c7ebb2d31b548791900

Observation d0a17abb-1372-4d11-91c2-955192748180 · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.953286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.953286Z digest=sha256:fe9bcba092215f241e67774fbbc50e4bc9cec1b1620f802e596057b71a6744bb

Observation 4c1f6efc-ac82-4ad9-b8ea-a7da7c53f5e8 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Movie Gen: A Cast of Media Foundation Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.992227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.992227Z digest=sha256:e13a073bd792e0f7636f2577c9ae90b62cc2535de5e00ef6788a5d19b5f67d4d

Observation c73196b0-e227-4152-aff1-2f0008a20497 · outbound

This paper cites Learning transferable visual models from natural language supervision.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Learning transferable visual models from natural language supervision

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.548601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.043194Z digest=sha256:7921bff35f8a958c72cfe6c3c604bc973b65cee3d5fc1dad70298c1a2dd312a8

Observation 43dd8388-8eef-4038-bb7c-43218e56275e · outbound

This paper cites an unresolved cited work.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:51:33.538990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.093926Z digest=sha256:992f9e676facd46981a832d6fbb240fc1bf824e1d55eb5525b8dc87aa6c5d24f

Observation 193a288d-a51e-48db-9925-16bcbf086d0b · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Gener- ating diverse high-fidelity images with vq-vae-2

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.529569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.127742Z digest=sha256:48124d597c42e3b6f98734137396792b57d6f7bdb3b3b5b6084dda7f888c8a17

Observation 1923a6b5-f8d4-4bd4-b227-1f59a54c8e57 · outbound

This paper cites Runway gen-3: Advanced video syn- thesis platform, 2024.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Runway gen-3: Advanced video syn- thesis platform, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.520108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.154624Z digest=sha256:6ebb4f0b6559b438654e7848ab97ccd1da392fd45e2d7d985989c431621caa34

Observation 629336ec-88d4-4118-bcce-b4a194e5ae67 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality High-resolution image synthesis with latent diffusion models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.210651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.210651Z digest=sha256:37129ba6ae498670dbaf3e0e3efa966b0e862c663c2c83cc6059c38a609697f2

Observation 2bdb081f-9b27-4e7c-a486-11961415d555 · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality U- net: Convolutional networks for biomedical image segmen- tation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.505284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.272230Z digest=sha256:9baac4f50b72d9285da9f042551dc4e4d32b5007e465713a83fed06b4f74b6e9

Observation 2986ff15-fdb4-4378-a3d2-44752c23a84b · outbound

This paper cites Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.496424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.359682Z digest=sha256:f8cdfbb40155ada4458396b436b28a9159bef1c3f6ab978bf01125fa9b8a6af0

Observation 25d87521-01e9-43a9-b2c0-2a304a3eb5db · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision,.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Flashattention-3: Fast and accurate attention with asynchrony and low-precision,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.487809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.399654Z digest=sha256:dbfa295dc703e0959709fb1a7085a38dfee6759ead58700a4a61e7f9ed9123eb

Observation 837e33c7-e193-4212-9f6d-db048e6219c0 · outbound

This paper cites Videopoet: Large language models are zero-shot video generators.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Videopoet: Large language models are zero-shot video generators

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.478980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.461098Z digest=sha256:3b7620c5ef4057abec8f0beea797c9717dc8fe5be61cb64acbd517b858d93e80

Observation fc06091d-be55-4325-92c2-5669ad12e32c · outbound

This paper cites Denois- ing diffusion implicit models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Denois- ing diffusion implicit models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.469672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.496431Z digest=sha256:d110132eee7f10f082f14eb7c108b95226b20b0e5e3cfb40c18e4f57a5405990

Observation 594bba73-2e54-457b-8ad6-55eca9fd08dd · outbound

This paper cites Kling2.0: Proprietary high-fidelity video generation, 2024.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Kling2.0: Proprietary high-fidelity video generation, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.460582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.532366Z digest=sha256:cf6eefed60726e86a3c9a2730c3027b6d357a619424e481f22da087cb4099510

Observation 27325cd4-c1a1-4879-9699-5c1953e11ec8 · outbound

This paper cites MAGI-1: Autoregressive Video Generation at Scale.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality MAGI-1: Autoregressive Video Generation at Scale

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.578351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.578351Z digest=sha256:05d7e052a3fbbd99728607036ead3be22e5150b531d84c2b1de88aff3b576a42

Observation afbce4be-0622-49fc-88fa-eb122facfc40 · outbound

This paper cites Mocogan: Decomposing motion and content for video generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Mocogan: Decomposing motion and content for video generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.451254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.649535Z digest=sha256:b1ca9f1238120421063c5c751e208757bf400c13d7a6aa65673200b3589ea5ed

Observation df8baf8e-f220-41ab-ab0f-bb296487f7a1 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.755370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.755370Z digest=sha256:95c58d93bb4e02905808ff9c13f9bc5808ad8d1d60e9426aba8f1f525634579c

Observation 345d4139-8d48-45b2-8d12-b8ff27894019 · outbound

This paper cites Neural discrete representation learn- ing.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Neural discrete representation learn- ing

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.441921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:28.845014Z digest=sha256:c5c218f47ecbb6f2108e3a8bc6e1cf1c86bf3614cd2df281828fadc8561364c8

Observation 37e02c48-704a-483f-ab1a-ed5581b25862 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Wan: Open and Advanced Large-Scale Video Generative Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.893268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.893268Z digest=sha256:887c54a736193c83931a28c27c9c1148e1302a5a18df71f7e06f9ba0db375054

Observation c9402324-cd38-4aab-8a99-fb2ca61e9150 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality ModelScope Text-to-Video Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.967399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.967399Z digest=sha256:11c42760f51067af1b57982f6b9d6246264d764c94c11b901cc042e561c45448

Observation 55664aa2-5f58-4555-9c35-0de6339c5977 · outbound

This paper cites Text2video: Gener- ating educational videos from text.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Text2video: Gener- ating educational videos from text

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.433200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:29.028121Z digest=sha256:65b031b9f6b16d2f6681eb82c92a802961faebae7da90d873e76950e844c4d20

Observation 079cc167-5b3a-4d9f-89ed-0c0b57dbb5ff · outbound

This paper cites Koala- 36m: A large-scale video dataset improving consistency between fine-grained conditions and video content, 2024.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Koala- 36m: A large-scale video dataset improving consistency between fine-grained conditions and video content, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.424383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:29.134036Z digest=sha256:af7a1fb5c894b9476acf6f2854b7e5992b493e547f0644fe22085b06f0b5ce42

Observation 3db225ab-3e21-479c-999b-2be81a2f815a · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.180801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.180801Z digest=sha256:baaef8b708ac0d695b292dffcf57022a96daa6d933b236520bb288d7553e5e93

Observation 2c21f850-8a82-478d-9044-85fec28bfd67 · outbound

This paper cites Diffuse and Disperse: Image Generation with Representation Regularization.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Diffuse and Disperse: Image Generation with Representation Regularization

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.240432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.240432Z digest=sha256:7cdeb88a2337bd6b076c6f5b7de961e24c198bb9c147bccb666475399c10695f

Observation cfc808d4-8e26-44b4-83a7-5c2aba9e5c35 · outbound

This paper cites Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.322880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.322880Z digest=sha256:4e545c2a0abdf7083d17db50240f0d066f3a1e96d30122254b26d3a677d038b8

Observation 8196cb90-41fc-4e2d-af9f-03ceb81ba722 · outbound

This paper cites Videofactory: Swap attention in spatiotemporal diffusions for text-to-video gen- eration.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Videofactory: Swap attention in spatiotemporal diffusions for text-to-video gen- eration

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.415822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:29.387266Z digest=sha256:20782141322dee3393ba6ca21340c0e7364875838e8bbd58b2f42c9a421b226a

Observation f6648244-5260-4af2-b69d-1baa0b764ec2 · outbound

This paper cites Non-local neural networks.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Non-local neural networks

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.406460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:29.473010Z digest=sha256:4ab0e8fc9f95b0b4a883728c3452c7754b9a6b14e23d387d61dc3988e36c2b47

Observation b1bca7bc-db88-4221-8c66-7be59d4e6f92 · outbound

This paper cites Yuille, Zicheng Liu, and Emad Barsoum.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Yuille, Zicheng Liu, and Emad Barsoum

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.564360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.564360Z digest=sha256:42a29bc90f5cb19ea23b6e701cd2304ce0ae6a942ed093e24b913c4aac1e57a0

Observation 08562ac5-5180-4707-9435-0fbf8578180c · outbound

This paper cites Loong: Generating Minute-level Long Videos with Autoregressive Language Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.715057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.715057Z digest=sha256:8cc19a7f839731f4784e59b3e658d7cf38c196e3bb3fe7a1365a218c903eb309

Observation ec0cb88a-42e8-4d6a-bc3e-c70189248a6a · outbound

This paper cites Bovik, Hamid R.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Bovik, Hamid R

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.396397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:29.761047Z digest=sha256:79c52f5ecd8339269a685886fc675e112a518358ee4af4b61faeabf2845caa19

Observation 39361e02-e3e2-40d7-a1ca-bb2daa1e6d9a · outbound

This paper cites Humanvid: Demystifying train- ing data for camera-controllable human image animation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Humanvid: Demystifying train- ing data for camera-controllable human image animation

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.388182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:29.921360Z digest=sha256:8fecb99926f5b20dff15a75c0c3217a64b7937c01bcecaf8f55d0e31d974becc

Observation 9d980628-5c98-412d-a19f-c778cdc8fdca · outbound

This paper cites FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.977844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.977844Z digest=sha256:9f3a6cc0c57865c0f8d7dc61aa24c92cd82440e3787f69d21afa7de0b0fe398f

Observation d7efc319-0dfa-4f4a-b26b-a9d530746040 · outbound

This paper cites Panacea: Panoramic and controllable video generation for au- tonomous driving.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Panacea: Panoramic and controllable video generation for au- tonomous driving

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.379861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:30.077913Z digest=sha256:c4cda61e408b42cbc8e9cbc53dfaf5d58136ff47446adbf7208d0d254bb19e4a

Observation e1a52704-b453-4aa2-b5c1-63e31b1b411d · outbound

This paper cites MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:30.200027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:30.200027Z digest=sha256:50f336edd1a92cf31cfeb48e5a45d52a06cf1b5e75a889b3d780af5887aae20a

Observation 9cf082b0-8e2d-4d20-b1e9-8b58e686189b · outbound

This paper cites Automated Movie Generation via Multi-Agent CoT Planning.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Automated Movie Generation via Multi-Agent CoT Planning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:30.310486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:30.310486Z digest=sha256:16ed381f1601b6916c13ad23d254ddd3b8023b5009f17a68e56d041fc800c1a3

Observation 282a7c20-516f-4f2c-9ae3-0d1d9d22b5a3 · outbound

This paper cites Mind the time: Temporally- controlled multi-event video generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Mind the time: Temporally- controlled multi-event video generation

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.371071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:30.361446Z digest=sha256:5e7dc68e90b4d61ebd454f3ba2eb41085737664706883c1e95ae371cd8dce7ad

Observation bce22bbb-dd8e-43bd-b66a-05e822aad854 · outbound

This paper cites A survey on video diffusion models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality A survey on video diffusion models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.362057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:30.462880Z digest=sha256:c246d4019148df1228e7bf8ff78a9273f4bafd149be2e483572fa84ddda1796a

Observation 285c0893-c957-41e1-822a-9b8cbd99868b · outbound

This paper cites Surgi- cal video synthesis using generative models for procedure training.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Surgi- cal video synthesis using generative models for procedure training

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.353216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:30.537696Z digest=sha256:c688b3403e0bfa7d72974fc38e2f9f3a6574367989482b488b7acfc655df5038

Observation fd631bcc-e4b8-4e1e-98e4-b71d250ac301 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.344569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:30.659490Z digest=sha256:26a4ef423bc45b554cd0583435e0d4f6bad5b98be7a56c5afe317e41f48bad74

Observation d67194a9-c23e-4ad0-b33f-b54bc01b11b7 · outbound

This paper cites Temporally Consistent Transformers for Video Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Temporally Consistent Transformers for Video Generation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:30.809287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:30.809287Z digest=sha256:501cb10af62ff8d185fa682444ea4680436795c5c949cfca85c6f116b75fcddd

Observation 899138b4-ed67-4150-a186-aab007597afb · outbound

This paper cites Vript: A video is worth thousands of words.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Vript: A video is worth thousands of words

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.336035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:30.935301Z digest=sha256:9e4b67a9e9c7287e38d2cd27196d07bc62fa9f29a85028bf2c8a2b42110df048

Observation c90b442e-095d-4bab-a9e6-47c48015729f · outbound

This paper cites 360° vr video genera- tion with generative adversarial networks.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality 360° vr video genera- tion with generative adversarial networks

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.326333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:31.112264Z digest=sha256:e0e00143b5136b981b7b62b1765cf007a2a4cd41196afbe0e753ba0b58d3ded5

Observation 2a742e0c-8837-4f4f-8bb2-41edbf20e9c1 · outbound

This paper cites VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:31.233075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:31.233075Z digest=sha256:46c16b3d809fd8ced59c8bbea5d54309ea3dca72d804f8ea5df3e1d0d97f4340

Observation fcde2073-940e-4af3-ac8b-5c23df6659fc · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Cogvideox: Text-to-video diffusion models with an expert transformer

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.315924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:31.377341Z digest=sha256:909bbadf123e6c8e89386d507bd34127385296d7f2e66ba1a69a6a2eacfa0ed5

Observation 83421e8d-4098-4575-b067-bba14b442e4a · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:31.488724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:31.488724Z digest=sha256:1c2132121918c171f5e43fd597c4f68d643628ba9eabd24e42f167777e1bc0c7

Observation 0f6cd5e5-47de-4748-bb35-16492f705483 · outbound

This paper cites Framepack: Pack- ing input frame context for next-frame prediction models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Framepack: Pack- ing input frame context for next-frame prediction models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:31.599998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:31.599998Z digest=sha256:a869e44b32ec418a199dde3b6c66bbb40b057334cf0fa341946642732bdf9721

Observation 36715bad-77f5-4ede-bd11-a840d56fa124 · outbound

This paper cites Efros, Eli Shecht- man, and Oliver Wang.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Efros, Eli Shecht- man, and Oliver Wang

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.305757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T18:51:31.661328Z digest=sha256:8296dd31f1b532d4798292be55e019a8b87564ecfe4e167396d2e641c960f01a

Observation e1cda6db-b13e-47c4-b547-d3ec492efd5e · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:31.757497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:31.757497Z digest=sha256:9535ca9abe06b294380a25b192aab7ad4a7f37d406ec1b6dc4d77146f6bd2b5e

Pith citing papers

Observation 01be779b-48f3-4e5c-be87-d677397da751 · inbound

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling cites this paper.

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T13:41:42.111186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:41:42.111186Z digest=sha256:64d0c55a1fdc0b9d53e890ef72a159cef96b581b5b01330f5b1608c805ae1a17