Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:24:07.779000Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2507.13753.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:24:07.779000Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d0aa1032-d329-4d63-8247-9bb45a30bd19 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Emerg- ing properties in self-supervised vision transformers
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fea59f5e-5851-4875-abd1-90ec627f9a3e · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Huang, and Niloy J
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3dfc697b-0294-40d4-abc9-6684c7e55d45 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Videocrafter2: Overcoming data limitations for high-quality video diffusion models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e1bcdd9-2d98-40f4-874d-dbd2d406b9e6 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a87cc984-5599-44c9-9934-d66f5bb13213 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Diffsynth: La- tent in-iteration deflickering for realistic video synthesis
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5741b480-aabe-4f75-800f-eb9fc30f5b5d · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9072bd9c-9ee6-4a21-9ce7-a2fabb2afd5d · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis TokenFlow: Consistent Diffusion Features for Consistent Video Editing
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 198ec7a1-a731-4f2f-b311-514e79b0cb7a · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8be94f5-25dc-4c24-b133-79ac6cb494f2 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c6848b5-8a91-4dd0-b625-b2ed812a968b · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Video dif- fusion models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 392234db-f6f1-463e-9773-724403ced627 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc5c1f7-20e0-4835-b18d-5dbda1226131 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vbench: Comprehensive benchmark suite for video generative models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9bd42f93-9578-4d15-8139-3352f5e9e531 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Text2video-zero: Text- to-image diffusion models are zero-shot video generators
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df7d391b-9067-43e4-85d2-c0e13ece3784 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9ef22ac-c314-4987-9511-ed90bea7f57e · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29fcdf9c-5ae6-4433-913d-484fa50812f1 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Amt: All-pairs multi-field transforms for efficient frame interpolation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ba76e2c-012c-4315-96b6-8d829a3a266e · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Flowvid: Taming imperfect op- tical flows for consistent video-to-video synthesis
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36f84924-0470-431f-8eb2-8773afda9bb5 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Open-Sora Plan: Open-Source Large Video Generation Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa1c850-ab78-40ae-a79a-833be2520ff8 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnimateDiff-Lightning: Cross-Model Diffusion Distillation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c554a101-d293-4e2f-bd43-d2de26e0588c · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Towards understanding cross and self-attention in stable diffusion for text-guided image editing
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9958775-c95f-4fb9-9431-f61f11ac0c70 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis At- tentive linguistic tracking in diffusion models for training- free text-guided image editing
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0d2149e-da78-48e9-9e14-de0bfcdc9884 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Video-p2p: Video editing with cross-attention control
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 918aafef-6849-436e-92b0-5db97238a645 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Evalcrafter: Benchmarking and evaluating large video generation models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4c8f94e-c403-4eb8-8480-bbe6e462190b · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72f10aa-8b15-455e-8fe0-e02e0c29693b · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Null-text inversion for editing real images using guided diffusion models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d940d20-b81a-42e8-97de-901cd1f70086 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Diffusion Models for Adversarial Purification
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6473f36f-99c7-45fe-bfb0-107d9b29ee7f · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15b108a6-e171-4e04-8766-8e853205d8b4 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Pika 1.0
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3e0ee66-6b11-45e7-99fe-3ba01e1f947e · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375ff990-6126-4afb-a6eb-1ae9bb0c884b · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Fatezero: Fus- ing attentions for zero-shot text-based video editing
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76492959-6da0-4ef4-9a55-911f3565b125 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis High-resolution image synthesis with latent diffusion models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 907c4bd1-90d0-4049-b627-4946d099121e · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3af56a6-1e74-46b2-95c1-85849e106bab · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 320ef7f0-0c21-4df9-940f-9781f83cfd51 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Bivdiff: A training-free framework for general-purpose video synthesis via bridging image and video diffusion models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aab08ec3-8f12-4e8a-9e98-8016e78cb9fa · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Edit-a-video: Single video editing with object-aware consistency
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4988830-1ec9-464e-b471-d6a70db25a75 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Make-A-Video: Text-to-Video Generation without Text-Video Data
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328a147e-05c2-41d5-81eb-ab35e98dfc4f · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Denoising Diffusion Implicit Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8657e463-d5c8-4a85-913c-bcf6d3112a97 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Plug-and-play diffusion features for text-driven image-to-image translation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 845dbd66-6d58-4869-8af5-3aa2b0ed9879 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis ModelScope Text-to-Video Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6e5d777-157d-4421-9480-31be1a41639a · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9a1aa62-1b7b-41d2-9abd-0de5e46b7720 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Exploring video quality assessment on user generated contents from aesthetic and technical perspectives
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7cf0cde8-aba3-4dc7-b04d-146c968e51e1 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Tune-a-video: One-shot tun- ing of image diffusion models for text-to-video generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffa0efa5-6e01-4291-81c3-745e8b7e73d2 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Rerender a video: Zero-shot text-guided video-to-video translation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a13928cf-a1bb-4530-9664-af5f101d3778 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Fresco: Spatial-temporal correspondence for zero-shot video translation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60f5666e-ec18-497c-99ac-9f1789767b36 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97fe92c6-dd19-48d1-8024-14ca4544c834 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f4480d7-b3f2-4914-bd8d-0eb6e57c9371 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis ControlVideo: Training-free Controllable Text-to-Video Generation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec343ff1-a59f-4dbf-ac32-4a867b6a9583 · outbound
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.