Pith. sign in

Paper Citation Record · LEDGER

Frame-Level Captions for Long Video Generation with Complex Multi Scenes

As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2505.20827.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20827 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:52:24.986017Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:33.443014Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T16:39:34.724653Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 326a447c-60ba-46d3-a187-c57f939e92e8 · outbound

This paper cites Lumiere: A space-time diffusion model for video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lumiere: A space-time diffusion model for video generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:28.475980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:20.619862Z digest=sha256:4d89d2c27d048e3df075f03bd968cfe2811e6a11da919451a972a23052437d62

Observation d58735dd-48a3-4813-89ea-9807fa66f72b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:20.736392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:20.736392Z digest=sha256:af91f560b5867da97903d51d35012d0a2bdb77cc1b7633f1ff39f2fa78ffac6a

Observation a57090df-d86f-46bc-bfce-6382860c7b8d · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Align your latents: High-resolution video synthesis with latent diffusion models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:20.841396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:20.841396Z digest=sha256:79bae0b0a7a5147a03c34fdb342b0286cb730b969731ab5afeba5d22e05bde04

Observation 1d51cc5a-4ff0-4606-a8f0-7645f5a55768 · outbound

This paper cites Video generation models as world simulators.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Video generation models as world simulators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:20.907024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:20.907024Z digest=sha256:f0e42f63b39a3f7b7d5e3dd0d20a56dde902a0ba078d3c98eb9ccf1ea2f9b5e5

Observation 44bba26c-5ef4-4aac-a2ab-af2857f75a54 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Activitynet: A large-scale video benchmark for human activity understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:28.241900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:20.983081Z digest=sha256:a002a316e4137844a75cc28f968a8d7e3af2ddfbec2054aa453c7e3c6fdf870a

Observation c3721de0-7be4-461d-a563-a61dbd409232 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffusion.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Diffusion forcing: Next-token prediction meets full-sequence diffusion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.064096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.064096Z digest=sha256:386cfc59a76a61627b2c1b3477d866cb2da5026a550f8d1dbddedc8e90976bfb

Observation 3038fb96-a49c-4bb0-a81f-681711bf0412 · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes SkyReels-V2: Infinite-length Film Generative Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.230585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.230585Z digest=sha256:d1eddbeeabd79026ab92232e02fd08a35dada9305b0d7431d703354d5575fbf4

Observation f2cd86fe-fcb3-4001-a2cc-2ac94e8811e4 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:28.033266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:21.313961Z digest=sha256:5aea0ef6b3ee4e6fd478897c0de515d3e7d6249088ad577d5d679fd6512db75b

Observation b2f7a23e-1389-4327-b6fa-4601eda84a77 · outbound

This paper cites Learning temporal coherence via self-supervision for gan-based video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Learning temporal coherence via self-supervision for gan-based video generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.875131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:21.429053Z digest=sha256:f1eed0705e74a538b8b9c8a0cbb808dfa11520db65fbb16b0f115c560991145c

Observation 15b9771c-e22f-4231-adeb-d36464455b75 · outbound

This paper cites Factorizing text-to-video generation by explicit image conditioning.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Factorizing text-to-video generation by explicit image conditioning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.757449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:21.543209Z digest=sha256:29917067b4234b0cd36a5f432d16350b46c8511c449e6ebf976ba4005df977f4

Observation fa7f08fb-4ff9-4ff0-b7f1-7d20d9a4f1f1 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.675607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.675607Z digest=sha256:4414ceb32131ef069da497b4b6e0b2ef67faaf4d1bae93819a30e737d56111b8

Observation b625c430-1f3c-43d7-ab62-2dd8f00864fd · outbound

This paper cites Long Context Tuning for Video Generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Long Context Tuning for Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.800462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.800462Z digest=sha256:fb174b8878cc476068286d5287785eea8ebb1b0999fc16d66e0c49aa6e9e2a0c

Observation 80296d13-3267-4b87-99b6-ede1ccec4070 · outbound

This paper cites StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.866461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.866461Z digest=sha256:51db32055432fc0736823eaebe42234bb80dba35496c49f9cdf664c07f53247c

Observation 438697a2-2e0d-43c2-b8ab-abc715c43277 · outbound

This paper cites Autoregressive diffusion models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Autoregressive diffusion models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.602519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:21.937109Z digest=sha256:759f8b89130a025cf5da333c437960a3b06aac8ff0f0e54e292ea4ca728d412c

Observation 8b9fdd70-f981-4764-a3ae-3e19a022e48e · outbound

This paper cites Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.021206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.021206Z digest=sha256:3171ade50f15aca85b37db909336e5243351702358576bfce93283a66e6817b0

Observation b13f9e56-6b9f-4ea2-a976-c9b6e453d8ee · outbound

This paper cites FIFO-Diffusion: Generating Infinite Videos from Text without Training.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes FIFO-Diffusion: Generating Infinite Videos from Text without Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.117866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.117866Z digest=sha256:0e067a60452d5515ebc4719b692ded799a9aedd1d176725d083fa0e4caa95a1d

Observation d12f45af-4f3b-4d6f-9679-3cce42ef9b65 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.237972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.237972Z digest=sha256:fe1c40c8e63015c62c554a1ee5e9ef8c1e841fb0d995063ca9ac3da80fedbab3

Observation a307adf3-e2ac-4d93-8f34-e9cdbdb6e240 · outbound

This paper cites A Survey on Long Video Generation: Challenges, Methods, and Prospects.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes A Survey on Long Video Generation: Challenges, Methods, and Prospects

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.315530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.315530Z digest=sha256:67c0add84098ee7354ad81c327a6ccc6b348935e526718a2497c017ff8c92ffc

Observation c4852276-2f62-44ff-88f6-1bece02499fe · outbound

This paper cites Unified Video Action Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Unified Video Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.397770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.397770Z digest=sha256:377430b415522048ef64832cc70c682a1e1c170394e2548e759f2e69177d5fe5

Observation 4ecca3c2-367f-4ce3-a333-0d009eac451e · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Open-Sora Plan: Open-Source Large Video Generation Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.472441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.472441Z digest=sha256:d965a8a768260eb89d1f519d50753c00215ef0067a6e912f9e7d38cb139feac9

Observation a076bb32-7a26-45d0-a597-80329400b0b5 · outbound

This paper cites Videostudio: Generating consistent-content and multi-scene videos.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Videostudio: Generating consistent-content and multi-scene videos

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.452934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:22.546815Z digest=sha256:3c6139ce147206f118238a4ab5b9c18f5a49bee38f477d3f540b3716c5122770

Observation e8eed2f6-a220-4570-a159-41bfa82eb6c3 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.696814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.696814Z digest=sha256:e270d6a0226372b5326077cff747fe07b8c156545ab6f86e71080774d3d789bc

Observation e77a915f-003e-4247-86e2-ad20030aba1a · outbound

This paper cites Mevg: Multi-event video generation with text-to-video models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mevg: Multi-event video generation with text-to-video models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.206936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:22.790070Z digest=sha256:c7ead35cd2dbc13368aa56d1c563211f07f3679032deae0b5dde0a25be40f637

Observation c9b86993-63dc-4c26-ba95-210aa7a8b470 · outbound

This paper cites Scalable diffusion models with transformers.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Scalable diffusion models with transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.881341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.881341Z digest=sha256:4eae55371fb5f8a0668b04bb20302b1906d9261cf78607b4e7e751f156028f59

Observation f227f7fc-c4cc-4f16-ad55-e6e0e7674011 · outbound

This paper cites Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.951099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.951099Z digest=sha256:73c3a22614f18a404a785fa7932c921a23ad8a49f1422c23aab6b8b48129c5a4

Observation 2cb4abcb-3de5-4d37-a93b-5e1a93019b13 · outbound

This paper cites Freenoise: Tuning-free longer video diffusion via noise rescheduling.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Freenoise: Tuning-free longer video diffusion via noise rescheduling

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.911330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:22.997892Z digest=sha256:c2e44489716a22a333cf717bddc4aac279c3f3abf834d6f88d8678ab6a1d5874

Observation c69b16a1-53d6-411e-920f-de6d91f91e68 · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.076920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.076920Z digest=sha256:a4c5504453ae5288a38d0e102a0b3c9848f003ae9253d7f2b450987df743f170

Observation 70b1d677-ea3a-4e73-802a-d1c568ef8139 · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.761609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:23.143892Z digest=sha256:ee1b9fad5eeab025a6cea529c87d7827f564a30617422cacab9ee8d866724ec3

Observation 212a9cf1-50d9-48f8-bbad-7a0f8af8487d · outbound

This paper cites Lightweight, Pre-trained Transformers for Remote Sensing Timeseries.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lightweight, Pre-trained Transformers for Remote Sensing Timeseries

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.239488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.239488Z digest=sha256:da091e4857eb8eb4b24669e54304989826e077135ee0024a5816a4f23518910d

Observation a1d42ff8-58c3-4633-ba3c-39b2f7a4c21f · outbound

This paper cites Mocogan: Decomposing motion and content for video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mocogan: Decomposing motion and content for video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.532001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:23.335228Z digest=sha256:8313666496ec6a4cd17ea7a0eb3a243efd3aa92102a65db429bfb7952addef57

Observation 55bd74b1-3bf0-40f5-9594-e1a9d0f0218d · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Wan: Open and Advanced Large-Scale Video Generative Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.434779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.434779Z digest=sha256:2b3c52f8ab4283e4e72974f912f34ae57b94d491d8445e4b4361f3875ecf482a

Observation f999d2fb-6d40-4027-9760-581b287740b2 · outbound

This paper cites STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.548609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.548609Z digest=sha256:c229346e94ce1f94120c0d5527b23f3ed48eeb30f9d7721addcf766695592de2

Observation d5bea886-b2d0-455b-af90-1c8903bcb576 · outbound

This paper cites Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.659573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.659573Z digest=sha256:1119865162b452731040c92bec700704ae9be4378551bfc4234fd394c69781fc

Observation a5e26b97-9690-40db-85af-400913c64a0e · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.736263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.736263Z digest=sha256:58448ee4da80b615f619f268756e71c2cf61206d1f61285c61f36cec03cc3faa

Observation 8ffd8aa8-5ad0-40cc-a4ae-1faf228dae00 · outbound

This paper cites Lvbench: An extreme long video understanding benchmark, 2024.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lvbench: An extreme long video understanding benchmark, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.833995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.833995Z digest=sha256:31c1f8bf75635349b4bbd0197f788625d63d5196f38481a88449b1542ea03083

Observation 7246b089-de6e-49b3-9bd5-6dd77cde997e · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.917757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.917757Z digest=sha256:3c01c047d3c1a80c05b8af9620813c72e768768bdb88cdea60b8a05fe531b469

Observation 0b4f6d54-734f-4648-98e3-9798e16b796c · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Vatex: A large-scale, high-quality multilingual dataset for video-and-language research

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.296919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:24.045232Z digest=sha256:656ed6b584786985118648de8829de57db7ef6c5e2ccfcc03ff382bbc768143e

Observation 83bf8499-8c7c-4f69-83f8-60277cbbbae4 · outbound

This paper cites Imaginator: Condi- tional spatio-temporal gan for video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Imaginator: Condi- tional spatio-temporal gan for video generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.128349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:24.153790Z digest=sha256:a4fce8faa22810f9becdac04382f650349da88e778dc6c30360be6c1bbfc305d

Observation 8e9afcf7-8ae4-48b7-97b1-71253170f820 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.252204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.252204Z digest=sha256:15be762d884eb25a3cb8ec801d729595859c914f4c2c4b296db5d7c01057e477

Observation 40835e76-76e2-4297-bbc0-8aeaf401f988 · outbound

This paper cites Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.370109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.370109Z digest=sha256:cb76f6eb56e375bb83655a63756fde1dc6cb8b93adfebb8105e98765099403d6

Observation 1861851a-92b6-47e7-94ad-99d54fc6287e · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.467700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.467700Z digest=sha256:85791b3a043d3dbe414c8cd41f75931a71fb34355c8b04dd169ab58e4cd8559a

Observation f68c7dbf-cef9-4122-9acc-69273d4938b4 · outbound

This paper cites Merlot: Multimodal neural script knowledge models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Merlot: Multimodal neural script knowledge models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:25.994105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:24.570316Z digest=sha256:40e405bc9d625897afd6d09eae32175f009fcfbeb0c61661e734a53bb356b806

Observation 4cd00ff3-e07a-4ac9-b454-5e03a94706d1 · outbound

This paper cites Moviedreamer: Hierarchical generation for coherent long visual sequence.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Moviedreamer: Hierarchical generation for coherent long visual sequence

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.643654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.643654Z digest=sha256:b30f490ed5a32ca436df3baa81a90ea4a403d4681fffc95f06257cd61c366142

Observation bfaca529-075e-4c00-81e6-caaaf624d3f9 · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.723287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.723287Z digest=sha256:63acbed5ec39207f9cb5bc0a7fc035b52f4282fae175ca913b27283ed2ba8c80

Observation 1a69e4fa-55ef-4a37-bc3e-75fd0c59ff2a · outbound

This paper cites Videogen-of-thought: A collaborative framework for multi-shot video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Videogen-of-thought: A collaborative framework for multi-shot video generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.804117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.804117Z digest=sha256:b48c7eab9541ad579e8710c37fddcd43710cba5e452a3d99fd2b90e877d24689

Observation 63da9f05-e6a4-45aa-a677-27469349e31c · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Towards automatic learning of procedures from web instructional videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:25.793585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:24.885503Z digest=sha256:b95460eef3e7d81ba78c367aabf8e32c726c6b0c1e6741bb0851a94c7244cac2

Observation 614f1169-0a76-4a97-82d3-70953de36695 · outbound

This paper cites Storydiffusion: Consistent self-attention for long-range image and video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Storydiffusion: Consistent self-attention for long-range image and video generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:25.576965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:52:24.986017Z digest=sha256:dcbb4bf525eb4ba9f5c74cfbd71a5e28c7966d5665371851ea583964aeb6fe9e

Pith citing papers

Observation 672da83a-bf59-448e-b9c2-ffe3a00488e4 · inbound

LoViC: Efficient Long Video Generation with Context Compression cites this paper.

LoViC: Efficient Long Video Generation with Context Compression Frame-Level Captions for Long Video Generation with Complex Multi Scenes

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:39:34.798842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:39:33.443014Z digest=sha256:f06d0be05c1c0543bbd73859d370916b48b15a63396c4ccf64f5285fdcbbb861