Pith. sign in

Paper Citation Record · LEDGER

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis

As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2507.13753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13753 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:24:07.779000Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0aa1032-d329-4d63-8247-9bb45a30bd19 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Emerg- ing properties in self-supervised vision transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.327857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:03.661367Z digest=sha256:ffe8b2d05b1709071f8e7e53af650f4a66f628bd7c7471da82ec4e7ef913e5c8

Observation fea59f5e-5851-4875-abd1-90ec627f9a3e · outbound

This paper cites Huang, and Niloy J.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Huang, and Niloy J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.308915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:03.740191Z digest=sha256:5367dd4d764a197e59fa0ebac3be77c226a17a4b5bfd07a9c4e319be38942a2e

Observation 3dfc697b-0294-40d4-abc9-6684c7e55d45 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.293233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:03.826011Z digest=sha256:7fbbcbef4ec159fa46fd448be22d946aef6cdc3f4c854f8e796508e4c458a530

Observation 0e1bcdd9-2d98-40f4-874d-dbd2d406b9e6 · outbound

This paper cites Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.276753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:03.932309Z digest=sha256:a95ec365e7b475531b7805d4c3fafd204383c577bdbac1ff1b93bbc84e1e1020

Observation a87cc984-5599-44c9-9934-d66f5bb13213 · outbound

This paper cites Diffsynth: La- tent in-iteration deflickering for realistic video synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Diffsynth: La- tent in-iteration deflickering for realistic video synthesis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.261382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:04.026929Z digest=sha256:4adb429078f58fff8857bf679e9dc242f7dae71624e77613e5e419c686dd458b

Observation 5741b480-aabe-4f75-800f-eb9fc30f5b5d · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.099250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.099250Z digest=sha256:3e00a6e5993541a082bf9f9fe6f5dc32b8c7866ffb47361aac37dc1dad6a1ed0

Observation 9072bd9c-9ee6-4a21-9ce7-a2fabb2afd5d · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.195327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.195327Z digest=sha256:bfbd91618d403dea3eb1150803fe9386c15c8ff2916f4778f3f19de3ce60f666

Observation 198ec7a1-a731-4f2f-b311-514e79b0cb7a · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.284760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.284760Z digest=sha256:bd6c2a44b6a3fd5c564a37c7a0f8a928f8c1224a55ffbf942cc176871f1d9f4b

Observation e8be94f5-25dc-4c24-b133-79ac6cb494f2 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.387504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.387504Z digest=sha256:e8a6901c87bfde4e4a47955f09c6a5e4be202f7de2331ec4ddefd799cb0ca7b9

Observation 4c6848b5-8a91-4dd0-b625-b2ed812a968b · outbound

This paper cites Video dif- fusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Video dif- fusion models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.473396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.473396Z digest=sha256:ca9cd6302e4e3d48d376e6d93c325074d463f1b0ddbf55b3fefbce80a419e305

Observation 392234db-f6f1-463e-9773-724403ced627 · outbound

This paper cites VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.571398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.571398Z digest=sha256:ca2fa707c0ee89eeaaccc495bbb0c0b9e5462b246676dc279bc958e977b97576

Observation 2dc5c1f7-20e0-4835-b18d-5dbda1226131 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vbench: Comprehensive benchmark suite for video generative models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.073040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:04.658315Z digest=sha256:6669f62e3bdea48729f2c391019fa0d155c3427f751b546e82a2f3ca8a6ee682

Observation 9bd42f93-9578-4d15-8139-3352f5e9e531 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.055169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:04.779140Z digest=sha256:8b380995a0caae2e7ddce01edfd9970db3721ace438da09327a574416a0bdd80

Observation df7d391b-9067-43e4-85d2-c0e13ece3784 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.867561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.867561Z digest=sha256:c4b2d34d9946e46cd7b92639e8ea47ae0097b72f326646e38a34041ac2da8ea0

Observation a9ef22ac-c314-4987-9511-ed90bea7f57e · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.956206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.956206Z digest=sha256:087102ab93b9b948d4c6dd4cf9c71b71e2dac1e90c1ce928b4c3d78263e88622

Observation 29fcdf9c-5ae6-4433-913d-484fa50812f1 · outbound

This paper cites Amt: All-pairs multi-field transforms for efficient frame interpolation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Amt: All-pairs multi-field transforms for efficient frame interpolation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.030716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:05.075194Z digest=sha256:76daa12e481d618076b8148bfa5ee87fec484e6c3d76f10a82c1e8d3db1b2d0c

Observation 6ba76e2c-012c-4315-96b6-8d829a3a266e · outbound

This paper cites Flowvid: Taming imperfect op- tical flows for consistent video-to-video synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Flowvid: Taming imperfect op- tical flows for consistent video-to-video synthesis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.001903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:05.139965Z digest=sha256:33c41933a61213efd293c9e596ab58f28532e0d464170fd5447d39b01fd64310

Observation 36f84924-0470-431f-8eb2-8773afda9bb5 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Open-Sora Plan: Open-Source Large Video Generation Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.217568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.217568Z digest=sha256:753b9559cbad9f28f210429db9e83d62a41e86b9134f2d9e290f9593436d6131

Observation 8fa1c850-ab78-40ae-a79a-833be2520ff8 · outbound

This paper cites AnimateDiff-Lightning: Cross-Model Diffusion Distillation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnimateDiff-Lightning: Cross-Model Diffusion Distillation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.312282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.312282Z digest=sha256:91f404b6f93c9f04cbff236b3d41f517edeb347640cef61d1717078f5a57f4fd

Observation c554a101-d293-4e2f-bd43-d2de26e0588c · outbound

This paper cites Towards understanding cross and self-attention in stable diffusion for text-guided image editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Towards understanding cross and self-attention in stable diffusion for text-guided image editing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.982327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:05.379869Z digest=sha256:1d5ab5a8835f04efb7d03b0e89f00c22d3de961d050ce128864cef54576c6872

Observation f9958775-c95f-4fb9-9431-f61f11ac0c70 · outbound

This paper cites At- tentive linguistic tracking in diffusion models for training- free text-guided image editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis At- tentive linguistic tracking in diffusion models for training- free text-guided image editing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.956632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:05.466037Z digest=sha256:79ab6ff5d439b8b7f29aef21ea8a96949fd1bd987d07d26cd9001a1decd4262f

Observation e0d2149e-da78-48e9-9e14-de0bfcdc9884 · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Video-p2p: Video editing with cross-attention control

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.939679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:05.555043Z digest=sha256:47cdc9e0d7db823edcb354c674e1b9e456b615fe119d9f8650ae5a715fb4135d

Observation 918aafef-6849-436e-92b0-5db97238a645 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Evalcrafter: Benchmarking and evaluating large video generation models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.920453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:05.625822Z digest=sha256:b826d5259522da28a596dfd5b13468842cb0c59108e4bef44fd5ddb247af536f

Observation a4c8f94e-c403-4eb8-8480-bbe6e462190b · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.715742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.715742Z digest=sha256:92adc4989c44421bc5ad86c19c681cbef58275abdf6a59d95c537ebadf3fb35f

Observation d72f10aa-8b15-455e-8fe0-e02e0c29693b · outbound

This paper cites Null-text inversion for editing real images using guided diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Null-text inversion for editing real images using guided diffusion models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.897595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:05.801114Z digest=sha256:90f6325bcb479d1629219159710942b6cae29c81b3a01c7db0b10fa6df44c2a2

Observation 6d940d20-b81a-42e8-97de-901cd1f70086 · outbound

This paper cites Diffusion Models for Adversarial Purification.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Diffusion Models for Adversarial Purification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.886229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.886229Z digest=sha256:5fec26e1f1bf82a1b075349dbd52501517733ede82d3e1abded476c94f70d6d1

Observation 6473f36f-99c7-45fe-bfb0-107d9b29ee7f · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:10.647955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:05.967881Z digest=sha256:02b00ceb061c93eefb2295ed120ab204200fea984ed195344cf037a72a81a89a

Observation 15b108a6-e171-4e04-8766-8e853205d8b4 · outbound

This paper cites Pika 1.0.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Pika 1.0

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.396746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:06.036782Z digest=sha256:7d18b9b1ea05ec1d348585ede2c4a6edb5e1056a87534a6d458eec674b8cbe59

Observation f3e0ee66-6b11-45e7-99fe-3ba01e1f947e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.099876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.099876Z digest=sha256:897ed5831b30055f27d89d6c1c1a8f57d5e1ac39a4df3369d201f9c689a4c2f8

Observation 375ff990-6126-4afb-a6eb-1ae9bb0c884b · outbound

This paper cites Fatezero: Fus- ing attentions for zero-shot text-based video editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Fatezero: Fus- ing attentions for zero-shot text-based video editing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.276468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:06.189085Z digest=sha256:86510235dcfc9047d954fa32ac1f16e9854fc9ee7426eb256c66f92eecc387a5

Observation 76492959-6da0-4ef4-9a55-911f3565b125 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.087093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:06.293597Z digest=sha256:62c2cf3cbf03d87846b89bda718f53c706024d8391d3020e621aab5b1800bfa4

Observation 907c4bd1-90d0-4049-b627-4946d099121e · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:09.833297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:06.383786Z digest=sha256:dee441bf00ef06470bc61ee2161198e48b6c00d5a5bb2954a7c33161285698aa

Observation a3af56a6-1e74-46b2-95c1-85849e106bab · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:09.717542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:06.484434Z digest=sha256:927e8d35ec7b2c922d5db36ae442be7d4664de7a55344a2b1ddf31a28e48d1db

Observation 320ef7f0-0c21-4df9-940f-9781f83cfd51 · outbound

This paper cites Bivdiff: A training-free framework for general-purpose video synthesis via bridging image and video diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Bivdiff: A training-free framework for general-purpose video synthesis via bridging image and video diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.612955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:06.570787Z digest=sha256:669ada40c63e74173b356baaa92da18a0a9551121b2499058ced0efdc3bd8b73

Observation aab08ec3-8f12-4e8a-9e98-8016e78cb9fa · outbound

This paper cites Edit-a-video: Single video editing with object-aware consistency.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Edit-a-video: Single video editing with object-aware consistency

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.422686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:06.666069Z digest=sha256:ffcb6448dee21fca0be82f163f320b18a2f434826968d95dc4c5a7c2ca0478f8

Observation b4988830-1ec9-464e-b471-d6a70db25a75 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.746047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.746047Z digest=sha256:97c35ba947a34b94b569ed393b42ff888986a2ee0f38cb3d2ad90fbc367a6687

Observation 328a147e-05c2-41d5-81eb-ab35e98dfc4f · outbound

This paper cites Denoising Diffusion Implicit Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Denoising Diffusion Implicit Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.829645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.829645Z digest=sha256:a66082b665f72f0ce06cd3a03519fd0aab903f7a93e1dcf62ef586a71afeb9f5

Observation 8657e463-d5c8-4a85-913c-bcf6d3112a97 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Plug-and-play diffusion features for text-driven image-to-image translation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.249295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:06.915722Z digest=sha256:e59c9debaf0a4995823483c1eb5476be376ead7589b05268f61453acb69da0aa

Observation 845dbd66-6d58-4869-8af5-3aa2b0ed9879 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis ModelScope Text-to-Video Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.980555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.980555Z digest=sha256:8872f4d9e503ab9085ff8377ce06b05fdb81f122b8952a352d4f2a548a1bf2e9

Observation f6e5d777-157d-4421-9480-31be1a41639a · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.087178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.087178Z digest=sha256:ed60c09664803cba1c6abfb03673f4f37dc93aa237e3238cba2ecc3c94527668

Observation a9a1aa62-1b7b-41d2-9abd-0de5e46b7720 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Exploring video quality assessment on user generated contents from aesthetic and technical perspectives

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.038369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:07.174284Z digest=sha256:d9eb0b7932c15460a8acc7b38dff8005469bfb82c7415eb822ccb74147530820

Observation 7cf0cde8-aba3-4dc7-b04d-146c968e51e1 · outbound

This paper cites Tune-a-video: One-shot tun- ing of image diffusion models for text-to-video generation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Tune-a-video: One-shot tun- ing of image diffusion models for text-to-video generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.865062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:07.254277Z digest=sha256:68c5cb2acb095be456d595c622e5706b36837ea042fed4fa9baf54a573542752

Observation ffa0efa5-6e01-4291-81c3-745e8b7e73d2 · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Rerender a video: Zero-shot text-guided video-to-video translation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.669898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:07.339187Z digest=sha256:5bc799fa69d945ce33289470d487aeb867ca838c70e641824bed80297f3c9fe7

Observation a13928cf-a1bb-4530-9664-af5f101d3778 · outbound

This paper cites Fresco: Spatial-temporal correspondence for zero-shot video translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Fresco: Spatial-temporal correspondence for zero-shot video translation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.509404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:07.427186Z digest=sha256:5a7162475d2f01d30b44504a3d044dac1c3cea745eddf7dd160f249668778fbd

Observation 60f5666e-ec18-497c-99ac-9f1789767b36 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.516251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.516251Z digest=sha256:18b76f0ddfded0c5c7a1bc77f704633b894daf60fc38d1b14429fb315ccb856b

Observation 97fe92c6-dd19-48d1-8024-14ca4544c834 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.602629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.602629Z digest=sha256:6a72aba75de61af276a790a176a7712b2f6bcdada0f7b7d259dca039df4aea8c

Observation 6f4480d7-b3f2-4914-bd8d-0eb6e57c9371 · outbound

This paper cites ControlVideo: Training-free Controllable Text-to-Video Generation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis ControlVideo: Training-free Controllable Text-to-Video Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.699049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.699049Z digest=sha256:29dd32afa8a73f10d532432fc3be8bb54d6418d7553851218ddf8fc2b892f0a5

Observation ec343ff1-a59f-4dbf-ac32-4a867b6a9583 · outbound

This paper cites VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:24:08.016744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:24:07.779000Z digest=sha256:53d05da5b5691775e731058ed5449ac2310734e9017865696805754089de4122

Pith citing papers

No inbound Pith citation observations are available.