Pith. sign in

Paper Citation Record · LEDGER

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models

As of 22 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 3 inbound Pith citation observations for arXiv:2505.07652.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07652 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:17:38.923422Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:25.127558Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:30.438596Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13678533-e14c-4dc6-935e-74d8b7ff31ef · outbound

This paper cites Kling ai.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Kling ai

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.878159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.616855Z digest=sha256:e48bd299f69afa5545c66a31083a7edbf5403491a017d991c30e88bca282e2ef

Observation bff35877-1525-4efa-bada-1c9879acd1e3 · outbound

This paper cites Universal guidance for diffusion models.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Universal guidance for diffusion models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.622320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.622320Z digest=sha256:e56376e3ffdfbdb149946d7401071bb3ff278d3042af48dced71b4a78cb05c24

Observation 0927c5f7-50fd-4bd9-ba9e-cb74ca6b6720 · outbound

This paper cites Pyscenedetect: A cross-platform tool for video scene detection.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Pyscenedetect: A cross-platform tool for video scene detection

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.853431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.627275Z digest=sha256:f2cfe92afd87f404171dec02045377e2ff7d86f11cb437554fcae3b118ad9520

Observation 421f36e9-7aeb-42ed-93a8-69c5bd344957 · outbound

This paper cites Video generation models as world simulators, 2024.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Video generation models as world simulators, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.838322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.632443Z digest=sha256:586a879a24c25b96f013cf308306ebcbd0c1dba3fc02eee48fb41ce378a43f38

Observation daf16dce-778a-4038-b148-88eff605474e · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Emerg- ing properties in self-supervised vision transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.638533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.638533Z digest=sha256:6254b2954027652290ec5e40d022ff6c333d22a6b067c2f9d7863a1de833829f

Observation bac505e7-2db1-4e76-8819-3cbd58540c02 · outbound

This paper cites Scene detection in videos using shot clustering and sequence alignment.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Scene detection in videos using shot clustering and sequence alignment

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.814594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.643521Z digest=sha256:cf3a2926546c9b87eadce8718266a26d94f1de8fb42e0be2953274ef83d98894

Observation 150b4bf3-da9d-4f53-8171-b9184be0d4f2 · outbound

This paper cites Gentron: Diffusion trans- formers for image and video generation.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Gentron: Diffusion trans- formers for image and video generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.799607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.648009Z digest=sha256:9e3c4065db611dec29a118517200e0a083b94ef5db8a36bb945aafe71b814809

Observation 3147f4f3-22b1-40a4-af91-bed7f87fb9d1 · outbound

This paper cites Seine: Short-to-long video diffu- sion model for generative transition and prediction.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Seine: Short-to-long video diffu- sion model for generative transition and prediction

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.784972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.653103Z digest=sha256:0875da044573269663a165373cedda865f8d6536c39bb207d9a7b420fd01c2c1

Observation bb04a63b-a8c9-476a-a408-095a31f43532 · outbound

This paper cites Efficient video prediction via sparsely conditioned flow matching.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Efficient video prediction via sparsely conditioned flow matching

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.769009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.657861Z digest=sha256:fca9f357ee3c33dbc802f88d2e624a360431c6fa2179eda8a3961800ffc89edd

Observation fc3e2f14-21dc-40a1-b278-4312ac90f489 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Arcface: Additive angular margin loss for deep face recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.662373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.662373Z digest=sha256:1bfc6d58cf09246ea068aea4f26bb99f5678f1596b9edba357d5474bf407258b

Observation ebaf899a-4e70-4c44-bc48-84c366975214 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.667053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.667053Z digest=sha256:a73f08da6023eb2b083e3cef7a69571c0bb7d155726e6d5e3776eeaf80895a2f

Observation 3dfc15bb-947d-417e-852d-371e2d5cbf0e · outbound

This paper cites An image is worth one word: Personalizing text-to-image gener- ation using textual inversion.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models An image is worth one word: Personalizing text-to-image gener- ation using textual inversion

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.733344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.671638Z digest=sha256:5c9c2be349a5069fa99e5c054eddcfaeb36629fd7f10d1c7851e300f8ab5f588

Observation 67c0e663-dbd0-4206-b3a5-9dfaa0179164 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Preserve your own correlation: A noise prior for video diffusion models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.676268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.676268Z digest=sha256:271ef67d79e6937e98679275015eeed94806d685085d887272ee025b546098eb

Observation 3683d5e7-e7f6-419d-96d6-26702eda47e0 · outbound

This paper cites Tokenflow: Consistent diffusion features for consistent video editing.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Tokenflow: Consistent diffusion features for consistent video editing

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.709213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.681209Z digest=sha256:7a2304e457707eaed306f309a4e080d2f15bb2d989f5bac31decc26ce0046cec

Observation 5f047eb8-026d-4dd3-b7d0-09eb4260100f · outbound

This paper cites Photorealistic Video Generation with Diffusion Models.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Photorealistic Video Generation with Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.685991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.685991Z digest=sha256:f3f9d4a9aab9329a759f015de60e9a55086c3ba6b5013b6c16bc1cc1cbf13016

Observation e8042d70-fc6c-4ce7-99ae-5f167a73af98 · outbound

This paper cites DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.690991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.690991Z digest=sha256:2544d37ec25a27b11c8ee17bda2b757888b14f46f2d17bbad83e568a63a193ad

Observation 484d5403-2b65-4494-a126-619a26ae8f65 · outbound

This paper cites Style aligned image generation via shared atten- tion.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Style aligned image generation via shared atten- tion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.695687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.695687Z digest=sha256:e51fa5fe89a36897e7e3dc4220b32c1e07ac6bdfaac94f2ef71ec1720c818899

Observation 4e3abdbb-6f5d-47e6-90a2-46a8fc963656 · outbound

This paper cites Denoising dif- fusion probabilistic models.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Denoising dif- fusion probabilistic models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.700700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.700700Z digest=sha256:8db8e316cdf252cbd12f1c84da5249da9ed2e6cf91328b506303468ac9fc7508

Observation 62e4923c-24c5-4724-9a51-eb4f449bc3f3 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Vbench: Comprehensive bench- mark suite for video generative models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.705078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.705078Z digest=sha256:69c3c24854a5974bec0d7429682010c50dbeb27c5d4bc190aaa32c02dc0a2b0f

Observation 3db3973c-c3ff-417f-bf7d-a4645ef8a3dd · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Pyramidal flow matching for efficient video generative modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.709462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.709462Z digest=sha256:71128939ef297f51443224ba829d161bd8419ab1864871a0424afe28566df289

Observation c886c939-06cd-4e18-9767-10277bb7470e · outbound

This paper cites Rave: Randomized noise shuf- fling for fast and consistent video editing with diffusion mod- els.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Rave: Randomized noise shuf- fling for fast and consistent video editing with diffusion mod- els

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.714073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.714073Z digest=sha256:a5c732d8ded931ae266c311c1ac2d21452dca84ba669a8f02969dc9ffcbabab2

Observation ffaa91db-4b72-4f0b-82a6-8068d7d324ca · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.655431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.719798Z digest=sha256:63ddb254a0e6313df9649038fd10ec613ab6ee03b6d22456ce1aa4d3f3d4ba89

Observation 01e55a88-6fb6-4935-bb92-07feccdbf8a5 · outbound

This paper cites Open-sora-plan, 2024.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Open-sora-plan, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.724495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.724495Z digest=sha256:6bca643c00470609f0a8ddec3a5921a07e389ed9f9da4e1e1879a61751658970

Observation dd21cfb2-33d6-41fc-bd41-9e0708e3d24a · outbound

This paper cites Dream machine.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Dream machine

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.629387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.729932Z digest=sha256:cea3e1349f8150fda044d0d4e7dc16423bb972ef6e7ba4452ddd17976507669c

Observation b0cfaac7-6e24-4102-a408-e8f56f8f01bc · outbound

This paper cites Photomaker: Customizing re- 9 alistic human photos via stacked id embedding.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Photomaker: Customizing re- 9 alistic human photos via stacked id embedding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.613799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.734284Z digest=sha256:4e9bb7dfd0394b42107fb0bd9d7ec35d6bc84302cd76a350d315c87a57d837f4

Observation d250b7e7-d8bb-4f4f-87ab-97dc96d1378a · outbound

This paper cites Decoupled weight de- cay regularization.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Decoupled weight de- cay regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.739528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.739528Z digest=sha256:e9f1135949348614c05a13edd6838bf837caec336c66e46fc86645b751e14741

Observation e2d6245f-540e-4357-8b26-9f9587b02e38 · outbound

This paper cites Subject- diffusion: Open domain personalized text-to-image genera- tion without test-time fine-tuning.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Subject- diffusion: Open domain personalized text-to-image genera- tion without test-time fine-tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.744817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.744817Z digest=sha256:4d5ac8200829a0c4f9f7bdec4fa6fbc175fd64dec8519a90b593b21e85149ae1

Observation feac9aea-0d51-4339-b06f-727199b10015 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Latte: Latent Diffusion Transformer for Video Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.750394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.750394Z digest=sha256:30d4d94cbd55a9ab601e9f0dc07a7e76ce0cd78bffff574f813922d4147fc860

Observation 72bf8f92-7891-423a-b888-1e303b04b61f · outbound

This paper cites an unresolved cited work.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:17:39.580371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.755813Z digest=sha256:094bc4db2a917249baa093c2e8f4afb759440eacf6469cdf7c9175442d75caeb

Observation d1fcb8bf-6d48-41dd-94df-0fe893535e37 · outbound

This paper cites Mevg: Multi-event video generation with text-to-video models.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Mevg: Multi-event video generation with text-to-video models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.565443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.760280Z digest=sha256:c71444e7614641a0c5a6e1dfd393bb2f01177f021fb22c652b4d503c2622e981

Observation edadf137-e43e-4fb7-b4cb-1e06cbab4294 · outbound

This paper cites Chatgpt: A large language model.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Chatgpt: A large language model

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.550167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.765519Z digest=sha256:c09cd0e89dea955a8db166ff6e43ccdfa4e5f975a7e5334d749fd4944acd7e17

Observation 0fdf758a-16f1-48ad-aced-b2372144e029 · outbound

This paper cites Scalable diffusion models with transformers.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Scalable diffusion models with transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.770791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.770791Z digest=sha256:294895fbb6b0541cd5eae5c356ba86f773bdfeee9caf5fa081791123c514b736

Observation b488866c-0024-415c-bfe0-813538cae67b · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.525697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.776349Z digest=sha256:5f657e41fd538d9001daff2bc8006812a28f60328d5bfc86e2296f9ac36cd003

Observation b0bb7b0b-f70f-415b-9b0e-8c7fe2035584 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Movie Gen: A Cast of Media Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.780938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.780938Z digest=sha256:5e7a3f2bc4992dd303cbde59d328a836a5777d1ad6fcd2020156028aa8e508aa

Observation d019136f-e93c-43df-8e57-cb4d109af167 · outbound

This paper cites Prolific: Online participant recruitment for surveys and research.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Prolific: Online participant recruitment for surveys and research

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.510997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.785952Z digest=sha256:d4b1cade94b1a300985e0060828d0bdef7a3d426136e33c9a5591b968fa89c93

Observation 44b3f07b-f08d-4217-ac57-ed0ca612e5e5 · outbound

This paper cites Freenoise: Tuning- free longer video diffusion via noise rescheduling.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Freenoise: Tuning- free longer video diffusion via noise rescheduling

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.494760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.791618Z digest=sha256:54f4a5a0d32494f3a77d1fea72e17ff5640e6dc3a7beb098e275afc215df7222

Observation cee252ca-5c9b-46ca-a96e-d125a2a252bd · outbound

This paper cites Hierarchical text-conditional image gen- eration with clip latents.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Hierarchical text-conditional image gen- eration with clip latents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.478025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.796779Z digest=sha256:1804ef9a46454519eafb96eb5cf37175ba2264901a3316d1a0454a989f950f9e

Observation 6f801bd2-c66e-4cde-8115-4557a034eb76 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models High-resolution image synthesis with latent diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.801183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.801183Z digest=sha256:9dae31c0e2fddbd0a47761d67ca6f8f20be208fee3b91f21d928926729477f53

Observation adf21d14-3ea7-4613-a044-388d7225c89d · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.805545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.805545Z digest=sha256:f78100683cb0ef16c20407037f06440430fc726b8a1536141966e86f37b06703

Observation f2dd2e66-6f3c-42b4-9e36-36a9debae172 · outbound

This paper cites Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.444672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.810061Z digest=sha256:9f68554fd7ab430b22f81d7e728936aca6f2532e90b6d531b763f16fc9b5ef82

Observation 5b775e97-b644-47b1-b73d-2a409e39f3ff · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Make-a-video: Text-to-video generation without text-video data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.428520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.814505Z digest=sha256:a59fb3b0eb5c21349f8d2f90408d2e8a6177163e4edf635879eb0626dc66f635

Observation 3ed16200-56e0-48bb-8ddd-40a84f0125f4 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Deep unsupervised learning using nonequilibrium thermodynamics

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.819171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.819171Z digest=sha256:0e0ce1205070f18584e00b2acd5735d0e3407b5b91030462d3ff3d73445e511a

Observation 13cf3a9a-5757-45ec-acd0-c7ac268f91c3 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Raft: Recurrent all-pairs field transforms for optical flow

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.824135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.824135Z digest=sha256:4895318f8a6614910663609a5fc25ed93e30ca9abc262c4106006b4ff1464b49

Observation 61ea114c-61e2-4299-a2a1-ec159051aff1 · outbound

This paper cites Training-free consis- tent text-to-image generation.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Training-free consis- tent text-to-image generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.394818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.829386Z digest=sha256:990624f317e8d0f9cbc3d984c57a94165dfe8c8944f4993db0248f5f9ca1fdda

Observation 66e89b28-7241-4d84-a48f-f52dc5e841ec · outbound

This paper cites Film history: An introduction.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Film history: An introduction

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.379440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.833943Z digest=sha256:497f1b2601063759ab4431a9b32e21ca822e515fc633603fd65066e041b0db8b

Observation 9e45f1bc-2c20-4267-be71-32d6c7bf78b2 · outbound

This paper cites YOLOv5: A state-of-the-art real-time object de- tection system.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models YOLOv5: A state-of-the-art real-time object de- tection system

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.838486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.838486Z digest=sha256:c2782181b52a6749d81440d59f502c0aa72fffcf6051820ae89c2e318678b91b

Observation c52dfaaf-5e98-4ae7-86e6-73bb32b5ab12 · outbound

This paper cites FVD: A new metric for video generation, 2019.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models FVD: A new metric for video generation, 2019

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.336122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.847821Z digest=sha256:6998a33d6c5b86b1a60bf0036d768530747567dc326e86b046b62bf4648c681a

Observation 46ad2349-9303-40b9-9ecc-a72f9ab648bf · outbound

This paper cites Gen-l-video: Multi-text to long video generation via temporal co-denoising, 2023.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Gen-l-video: Multi-text to long video generation via temporal co-denoising, 2023

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.320519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.852124Z digest=sha256:8c1bd0c3c236e6e78eb6767f21277cfee11157ce3d6908e4f8c52f12d91b0590

Observation 1264697e-8fd4-46e2-8a1c-30a905398aa9 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.856667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.856667Z digest=sha256:4e45c2773a35e112363472fba3f2d42eb31cdfc9d596e7ac08bef34cc0b28e15

Observation 07d4b6c6-58bb-4996-9253-29a8e56b508d · outbound

This paper cites Internvid: A large-scale video-text dataset for mul- timodal understanding and generation.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Internvid: A large-scale video-text dataset for mul- timodal understanding and generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.305364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.861790Z digest=sha256:5556b845e92d16570ae58f5e64d3cc0192d80b703d95a750d16ac5dc49f2bdda

Observation 073c7195-f5bd-4f29-aaf0-3f39537a7587 · outbound

This paper cites Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Elite: Encoding visual con- cepts into textual embeddings for customized text-to-image generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.290291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.866843Z digest=sha256:764d2bc351199ec9681c4eb09caf043f1942a10229d9177ac33ac99c976904dd

Observation 0740253b-cf97-4a28-9174-c46137e68821 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.275256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.871418Z digest=sha256:4ef66e3c9a7d463f6f24f9771913420d1c5f85625b72a62d39116670bd866796

Observation ad1eeb89-1b14-46ad-9986-218bc092ba9d · outbound

This paper cites Fastcomposer: Tuning-free multi- subject image generation with localized attention.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Fastcomposer: Tuning-free multi- subject image generation with localized attention

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.875837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.875837Z digest=sha256:c401eb4d960c50327aa39a6341c0c00b973b186a87a775be995d33b143ee1b1e

Observation 79302690-86fc-43e8-acf0-d8f658a8e940 · outbound

This paper cites FaceStudio: Put Your Face Everywhere in Seconds.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models FaceStudio: Put Your Face Everywhere in Seconds

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.881049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.881049Z digest=sha256:9a96236d6754c301efd997345029d163321d52df4c3a3e8a3f933af2a9d00f81

Observation 14bf39c8-b05f-4c27-a6de-e0c5299436e6 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.891197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.891197Z digest=sha256:3592bbf0ec8e5da729e28dd6ae107df8042b8942bcbf214c7f1b8c6f44b574c0

Observation 7b2e5d4a-ef95-4327-abeb-dc682749d03f · outbound

This paper cites Automatic partitioning of full-motion video.Multimedia systems, 1:10–28, 1993.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Automatic partitioning of full-motion video.Multimedia systems, 1:10–28, 1993

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.250026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.895590Z digest=sha256:daaceabb036cc67f31874427ae0550f9953ab952740c2a51eb8bac8c2fb1554c

Observation d40f09fb-7899-40a4-a32d-11fa6831784e · outbound

This paper cites Llava-next: A strong zero-shot video understanding model.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Llava-next: A strong zero-shot video understanding model

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.235545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.899964Z digest=sha256:247962667c30d61a87088e8151528ca8217fc98d34ab9b672ac76e90153734aa

Observation 2f4eead7-6685-4040-84ba-bccb3d848609 · outbound

This paper cites Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.219905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.904280Z digest=sha256:2a3882a22fac7c9a1669270dc75cdd4f94e56a7aee1916311a9148fbe3cf05f5

Observation e2160dad-e90f-42e3-bf05-8d3954f2ed2f · outbound

This paper cites Open-sora: Democratizing efficient video production for all.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Open-sora: Democratizing efficient video production for all

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.204030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.908826Z digest=sha256:62e91006b47a769972c95699a16ef4d54cc90491b37e361a366d318009b3ee0f

Observation 17f2e3ad-8568-4fbf-b70d-01329b46c249 · outbound

This paper cites StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.913373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.913373Z digest=sha256:36cd6de85e9f9565fe63d075e9bd728dc7de719093c0723b1826d8937824dc54

Observation 8132c316-0c14-4c3f-936f-4124ee2c06df · outbound

This paper cites StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:38.918088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:38.918088Z digest=sha256:1154993eb6a8f2c87632a2d0521761d47c717723b334bfb2713ac6ea80dd17d7

Observation e4ed17f5-8423-4708-9276-6337c466ddb3 · outbound

This paper cites an unresolved cited work.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:17:39.351942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.843557Z digest=sha256:5d78e5846f950bb9ba6828b6d4f91cd9a479c37cebdb26ec3d15e4959fc1e32c

Observation 29608644-589f-48f6-afcf-c0b995dace0d · outbound

This paper cites a man reads a book under tree.

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models a man reads a book under tree

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:39.188383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:17:38.923422Z digest=sha256:3b673189e04a0a6f413c4e3a819e17a4021d689dd492e07dca5eb9352536e183

Pith citing papers

Observation 25fc22a4-82b8-4e38-a725-ab4140fa5bcc · inbound

LoViC: Efficient Long Video Generation with Context Compression cites this paper.

LoViC: Efficient Long Video Generation with Context Compression ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:25.127558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:39:25.127558Z digest=sha256:c71b4372762ec5a6c5fc7b40a87fd8afbbcbf62be1e9063e5d241bc2801dfe46

Observation 91725449-9114-4041-b4f5-ecb4f5dd7b41 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:19:30.440932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T18:15:51.862192Z digest=sha256:1c5d0c188b1cfb0cc6ccb5cc9ee3ff86482fd7eb83487579d209fa7a07f64d8a

Observation 392e4b67-0a10-4b7c-8b45-0e3d49bc1103 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:04.972588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:04.972588Z digest=sha256:39c18025bc479e5357801dfe1bc3213ed6b38d224c81e00ec0aa13e902c3318b