Pith. sign in

Paper Citation Record · LEDGER

Frame-Level Captions for Long Video Generation with Complex Multi Scenes

As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2505.20827.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20827 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:52:24.986017Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:33.443014Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T16:39:34.724653Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 326a447c-60ba-46d3-a187-c57f939e92e8 · outbound

This paper cites Lumiere: A space-time diffusion model for video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lumiere: A space-time diffusion model for video generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:28.475980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:20.619862Z digest=sha256:ea7699ee2a915ccdc8c6af3a0685702a893be53be47005d675d9cb70fd78bfe0

Observation d58735dd-48a3-4813-89ea-9807fa66f72b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:20.736392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:20.736392Z digest=sha256:31ff487e737b8273562f8a8c4cceec86c3907aa5e615eaaab0318c6d29dac4e0

Observation a57090df-d86f-46bc-bfce-6382860c7b8d · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Align your latents: High-resolution video synthesis with latent diffusion models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:20.841396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:20.841396Z digest=sha256:d7bfa6b68990f54e498826138592a797aee781347d94d6c1282dc9b48f927129

Observation 1d51cc5a-4ff0-4606-a8f0-7645f5a55768 · outbound

This paper cites Video generation models as world simulators.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Video generation models as world simulators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:20.907024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:20.907024Z digest=sha256:ddd368d59caaf2f2b675a277578bfd7c6409f376bbfa2d9734d22e98397ac7e5

Observation 44bba26c-5ef4-4aac-a2ab-af2857f75a54 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Activitynet: A large-scale video benchmark for human activity understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:28.241900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:20.983081Z digest=sha256:93db27a0d3e78864841bf2b365b681c7da98a74e07dcbc527f0ae4fc08da96ab

Observation c3721de0-7be4-461d-a563-a61dbd409232 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffusion.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Diffusion forcing: Next-token prediction meets full-sequence diffusion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.064096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.064096Z digest=sha256:ae127495461b2f2d18517d7012d3ad59b11db3b9836f05bb9899c4e6a5d9b246

Observation 3038fb96-a49c-4bb0-a81f-681711bf0412 · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes SkyReels-V2: Infinite-length Film Generative Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.230585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.230585Z digest=sha256:b9a8b7ea59fa6de14e03b2c3dfbe23f7b926694196b175cf0d9abd0f4bbac77a

Observation f2cd86fe-fcb3-4001-a2cc-2ac94e8811e4 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:28.033266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:21.313961Z digest=sha256:95ab7c882eaf1de3732aec26399a691b19548431cc9f3420597738caa7eaf4cc

Observation b2f7a23e-1389-4327-b6fa-4601eda84a77 · outbound

This paper cites Learning temporal coherence via self-supervision for gan-based video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Learning temporal coherence via self-supervision for gan-based video generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.875131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:21.429053Z digest=sha256:0d58337825cbb65d25805f9e8b9368fca75560c48c788f7222550a88b838c624

Observation 15b9771c-e22f-4231-adeb-d36464455b75 · outbound

This paper cites Factorizing text-to-video generation by explicit image conditioning.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Factorizing text-to-video generation by explicit image conditioning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.757449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:21.543209Z digest=sha256:69eb70839b5795390c7f1d104cd4c2ef5753945f46de4ba6b2f8b5edccbece87

Observation fa7f08fb-4ff9-4ff0-b7f1-7d20d9a4f1f1 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.675607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.675607Z digest=sha256:a647a32e851b62ba3e8a0e5f10ae621d6b0fb8b31e1e0204b137077657ee6a05

Observation b625c430-1f3c-43d7-ab62-2dd8f00864fd · outbound

This paper cites Long Context Tuning for Video Generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Long Context Tuning for Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.800462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.800462Z digest=sha256:9fc2cc10b4b0c1cacfab3a33ba769b46187cc4f924199d2c0d323a647a0b2f8a

Observation 80296d13-3267-4b87-99b6-ede1ccec4070 · outbound

This paper cites StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.866461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.866461Z digest=sha256:4679890c3e828f5eb1ec30b1eb6cf6de70bd6b9ab7910662973639253e0fe3da

Observation 438697a2-2e0d-43c2-b8ab-abc715c43277 · outbound

This paper cites Autoregressive diffusion models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Autoregressive diffusion models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.602519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:21.937109Z digest=sha256:88cc8208230aab273aec4a7e0bf59419257f22415d195d59ea3e0c02b6523089

Observation 8b9fdd70-f981-4764-a3ae-3e19a022e48e · outbound

This paper cites Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.021206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.021206Z digest=sha256:8fa020de69e9d40ba6ecfbf87a2392d4854af67b4eefdeaf583ec269cfd0eb25

Observation b13f9e56-6b9f-4ea2-a976-c9b6e453d8ee · outbound

This paper cites FIFO-Diffusion: Generating Infinite Videos from Text without Training.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes FIFO-Diffusion: Generating Infinite Videos from Text without Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.117866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.117866Z digest=sha256:ba215b0afa811102d66802977041021067de2adab05bdaa6bd17658cbc157c80

Observation d12f45af-4f3b-4d6f-9679-3cce42ef9b65 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.237972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.237972Z digest=sha256:7bdaf16c7f20170a4e4e405477db18b5f6e0869225ae499d8b6e80ece21275ee

Observation a307adf3-e2ac-4d93-8f34-e9cdbdb6e240 · outbound

This paper cites A Survey on Long Video Generation: Challenges, Methods, and Prospects.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes A Survey on Long Video Generation: Challenges, Methods, and Prospects

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.315530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.315530Z digest=sha256:170c9a7b5b0ca052f965a104fbbc834cd5c97952cdea08cacc2cfa9681b08f5b

Observation c4852276-2f62-44ff-88f6-1bece02499fe · outbound

This paper cites Unified Video Action Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Unified Video Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.397770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.397770Z digest=sha256:5ac3c5f80627ee4242ddbe55d2d0ed25ae707bd31b70db4b8ddd50c8735fde73

Observation 4ecca3c2-367f-4ce3-a333-0d009eac451e · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Open-Sora Plan: Open-Source Large Video Generation Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.472441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.472441Z digest=sha256:61ae8c74400f967175485d9d8926c2426234e03551cf71781961604b7126f7cf

Observation a076bb32-7a26-45d0-a597-80329400b0b5 · outbound

This paper cites Videostudio: Generating consistent-content and multi-scene videos.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Videostudio: Generating consistent-content and multi-scene videos

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.452934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:22.546815Z digest=sha256:5bddfaa9cbbbb2884776c5c17687718d540ad00ed98b9ac6ce8ca9b3c25692b2

Observation e8eed2f6-a220-4570-a159-41bfa82eb6c3 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.696814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.696814Z digest=sha256:50a64ba72633d9c5a51836d13de63c7f1cceb3ca4405d8ab6ab628720b4256a3

Observation e77a915f-003e-4247-86e2-ad20030aba1a · outbound

This paper cites Mevg: Multi-event video generation with text-to-video models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mevg: Multi-event video generation with text-to-video models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.206936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:22.790070Z digest=sha256:383bb4c32568879bf539e4c95507f5f876178d1d662b3880568fa6e3c8d50645

Observation c9b86993-63dc-4c26-ba95-210aa7a8b470 · outbound

This paper cites Scalable diffusion models with transformers.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Scalable diffusion models with transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.881341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.881341Z digest=sha256:52a6bb267ff661c0674045e18e543855a4e9d7e772cd99eb7d15d0a95da24c2a

Observation f227f7fc-c4cc-4f16-ad55-e6e0e7674011 · outbound

This paper cites Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.951099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.951099Z digest=sha256:2bbc644603ac8e3f8a0914fc2c8b5fc3ca40342a5e59ead39ed61f74e3d11864

Observation 2cb4abcb-3de5-4d37-a93b-5e1a93019b13 · outbound

This paper cites Freenoise: Tuning-free longer video diffusion via noise rescheduling.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Freenoise: Tuning-free longer video diffusion via noise rescheduling

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.911330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:22.997892Z digest=sha256:37b11dbc9b497380b2406849aa7e307c5283244d5df787d0b45792d108a14f03

Observation c69b16a1-53d6-411e-920f-de6d91f91e68 · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.076920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.076920Z digest=sha256:10771ba63078e4a46d702eedc1befd2b3e8b020c832a3385dcbbd7e6c6b5a0d0

Observation 70b1d677-ea3a-4e73-802a-d1c568ef8139 · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.761609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:23.143892Z digest=sha256:efdad63311ced4972cd987234f8b10a50c962f508d8b66c313c0aa0363044917

Observation 212a9cf1-50d9-48f8-bbad-7a0f8af8487d · outbound

This paper cites Lightweight, Pre-trained Transformers for Remote Sensing Timeseries.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lightweight, Pre-trained Transformers for Remote Sensing Timeseries

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.239488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.239488Z digest=sha256:b756f23af94d86ecc0aafc6f7e9681606352478bd4d1222879ee22bc8bc67d58

Observation a1d42ff8-58c3-4633-ba3c-39b2f7a4c21f · outbound

This paper cites Mocogan: Decomposing motion and content for video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mocogan: Decomposing motion and content for video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.532001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:23.335228Z digest=sha256:ba2e8219643602aaf18f66f0027b9b05a82b09b922a179efb3fae1f37314a8ec

Observation 55bd74b1-3bf0-40f5-9594-e1a9d0f0218d · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Wan: Open and Advanced Large-Scale Video Generative Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.434779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.434779Z digest=sha256:6fb922e3bbc1606289d0c78b673b4f5834f3e324c62cd75d249197a9d21610be

Observation f999d2fb-6d40-4027-9760-581b287740b2 · outbound

This paper cites STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.548609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.548609Z digest=sha256:9f68acf6ba6b512b2b8ae70bbbbe89ad7bc6dd85029806862e874694ff69479a

Observation d5bea886-b2d0-455b-af90-1c8903bcb576 · outbound

This paper cites Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.659573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.659573Z digest=sha256:90cbd6ca482a5a5289ee39e8aa808326f1706b3fa7ade4ea456e81af7daa5d7b

Observation a5e26b97-9690-40db-85af-400913c64a0e · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.736263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.736263Z digest=sha256:5a0228944f5aefa3dfb54bee5973d898363e6a05bccef10d614c254e6edcd263

Observation 8ffd8aa8-5ad0-40cc-a4ae-1faf228dae00 · outbound

This paper cites Lvbench: An extreme long video understanding benchmark, 2024.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lvbench: An extreme long video understanding benchmark, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.833995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.833995Z digest=sha256:8ad82726e5afac9094f1db8cccf47f5276c278e9ffc33ba414e22a41c45ea58c

Observation 7246b089-de6e-49b3-9bd5-6dd77cde997e · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.917757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.917757Z digest=sha256:92abdb917e932ff6074fde45e25df35f91d96baeebec78b19310baebb5c531aa

Observation 0b4f6d54-734f-4648-98e3-9798e16b796c · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Vatex: A large-scale, high-quality multilingual dataset for video-and-language research

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.296919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:24.045232Z digest=sha256:047f1b1653b345025f4ae35b70d1c3f1f18d7e02a9f9137dd88a3f24f5826691

Observation 83bf8499-8c7c-4f69-83f8-60277cbbbae4 · outbound

This paper cites Imaginator: Condi- tional spatio-temporal gan for video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Imaginator: Condi- tional spatio-temporal gan for video generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.128349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:24.153790Z digest=sha256:237bd508a11a2262e8535d76cacfe80d1fa9122990701b818fd85c54aa9b4981

Observation 8e9afcf7-8ae4-48b7-97b1-71253170f820 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.252204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.252204Z digest=sha256:bd039cdea23442d3a280608724b9c5e34cf03d8d360ac350c18e5c167ca62fd9

Observation 40835e76-76e2-4297-bbc0-8aeaf401f988 · outbound

This paper cites Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.370109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.370109Z digest=sha256:dbde353212e0d0142ca610c7c3110f426dc7a113d951c50a8f7f7cc6b0fba6c8

Observation 1861851a-92b6-47e7-94ad-99d54fc6287e · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.467700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.467700Z digest=sha256:ffc076538f5a5c9656dc980aff9052e1303bf27d42a7266cd74700b898e2d335

Observation f68c7dbf-cef9-4122-9acc-69273d4938b4 · outbound

This paper cites Merlot: Multimodal neural script knowledge models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Merlot: Multimodal neural script knowledge models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:25.994105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:24.570316Z digest=sha256:d3b1bf709da6e237de37471e077bea6ccdc2bc398fa501dd2dc5ade477f002bd

Observation 4cd00ff3-e07a-4ac9-b454-5e03a94706d1 · outbound

This paper cites Moviedreamer: Hierarchical generation for coherent long visual sequence.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Moviedreamer: Hierarchical generation for coherent long visual sequence

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.643654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.643654Z digest=sha256:5b425543201b085ec559c9e2bf9e21c9f33aee98fed935144f5118b20c97a624

Observation bfaca529-075e-4c00-81e6-caaaf624d3f9 · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.723287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.723287Z digest=sha256:c08706ee01092b3a2f7164288bee34cda8025e9594ee2ff39ba69a05dc04c1d3

Observation 1a69e4fa-55ef-4a37-bc3e-75fd0c59ff2a · outbound

This paper cites Videogen-of-thought: A collaborative framework for multi-shot video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Videogen-of-thought: A collaborative framework for multi-shot video generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.804117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.804117Z digest=sha256:abacb9b07e0f769cfe571ad054fe9a262771184bc11153f4aeeaa9fa4b29c7d9

Observation 63da9f05-e6a4-45aa-a677-27469349e31c · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Towards automatic learning of procedures from web instructional videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:25.793585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:24.885503Z digest=sha256:4fb17416931194fc7a55f2fe01156af8ddcd4ed299e4a21f42939407b96cf2c5

Observation 614f1169-0a76-4a97-82d3-70953de36695 · outbound

This paper cites Storydiffusion: Consistent self-attention for long-range image and video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Storydiffusion: Consistent self-attention for long-range image and video generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:25.576965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:52:24.986017Z digest=sha256:275e1324b48e2750767674033487f111cc989feb60c1fbb22e712f570f69f3eb

Pith citing papers

Observation 672da83a-bf59-448e-b9c2-ffe3a00488e4 · inbound

LoViC: Efficient Long Video Generation with Context Compression cites this paper.

LoViC: Efficient Long Video Generation with Context Compression Frame-Level Captions for Long Video Generation with Complex Multi Scenes

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:39:34.798842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:39:33.443014Z digest=sha256:2d8ffedacd9c712c8f01a84efd9f85672ec2d2930a5a9059d3cbbd9ad668e54d