Pith. sign in

Paper Citation Record · LEDGER

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis

As of 14 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2507.13753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13753 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:24:07.779000Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0aa1032-d329-4d63-8247-9bb45a30bd19 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Emerg- ing properties in self-supervised vision transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.327857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:03.661367Z digest=sha256:6360d14adeaf18fd08dd7806194c9061ef240dd7788a6d9b923d3e401065f15b

Observation fea59f5e-5851-4875-abd1-90ec627f9a3e · outbound

This paper cites Huang, and Niloy J.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Huang, and Niloy J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.308915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:03.740191Z digest=sha256:69f16015e5f74dc330f2cab278b87dd757a2c8034b4b0f61286dd850fe381a3f

Observation 3dfc697b-0294-40d4-abc9-6684c7e55d45 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.293233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:03.826011Z digest=sha256:5537cf467056feed40e629bc537c9d06302388fb6ab875da17bb2a8dce783d16

Observation 0e1bcdd9-2d98-40f4-874d-dbd2d406b9e6 · outbound

This paper cites Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.276753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:03.932309Z digest=sha256:24016627709d0780ce49e9af44eb13a26b585da2411c2e95dbceda7900b0545a

Observation a87cc984-5599-44c9-9934-d66f5bb13213 · outbound

This paper cites Diffsynth: La- tent in-iteration deflickering for realistic video synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Diffsynth: La- tent in-iteration deflickering for realistic video synthesis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.261382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:04.026929Z digest=sha256:4b92bd76655e23b2d7e8fc8dc6859ebf6e4960e4b599a017194af1d6b2ade400

Observation 5741b480-aabe-4f75-800f-eb9fc30f5b5d · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.099250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.099250Z digest=sha256:c8c3eb8da186a59dc5e6aebd9fbd750457a42b83d2f5e1137728a10ebb5e3bdb

Observation 9072bd9c-9ee6-4a21-9ce7-a2fabb2afd5d · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.195327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.195327Z digest=sha256:0a8eb6e7d2256fa6213e05985e8113e5d9f63b650c96b4fdd3a5a48e9e96ca5c

Observation 198ec7a1-a731-4f2f-b311-514e79b0cb7a · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.284760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.284760Z digest=sha256:540d8fa86425105a109eb387067e8eebd8e472e13841d3033e5e5d049aa17696

Observation e8be94f5-25dc-4c24-b133-79ac6cb494f2 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.387504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.387504Z digest=sha256:500eb76766fc26bb2a9c190c8558c7a64db086913de8ed575de5816eca0bd710

Observation 4c6848b5-8a91-4dd0-b625-b2ed812a968b · outbound

This paper cites Video dif- fusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Video dif- fusion models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.473396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.473396Z digest=sha256:3fa1307fc7a6a4492b96b22ae2d4e3671126373d2a98fac8b4cb974ae7120356

Observation 392234db-f6f1-463e-9773-724403ced627 · outbound

This paper cites VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.571398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.571398Z digest=sha256:2605a2478284b752eca2408964fe89d9b1a19ea4838571521856d0c55542315b

Observation 2dc5c1f7-20e0-4835-b18d-5dbda1226131 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vbench: Comprehensive benchmark suite for video generative models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.073040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:04.658315Z digest=sha256:367d2aad71bb768d288a142f4a39d9098e05a5dcd6f3b8482fcc7de360bfd2c0

Observation 9bd42f93-9578-4d15-8139-3352f5e9e531 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.055169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:04.779140Z digest=sha256:4fb40a35e2b7858935681988da5c4e95806d386f0e50cabf781db20144e8d0e3

Observation df7d391b-9067-43e4-85d2-c0e13ece3784 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.867561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.867561Z digest=sha256:abd87556b17cea58bb4df6aa845509544ab1b266cf18594e813cc0f660ffe053

Observation a9ef22ac-c314-4987-9511-ed90bea7f57e · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.956206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.956206Z digest=sha256:2b2ad084eec6bd4d8cb2ed7b0a4b67a36bd0afd65b5dd5e65d9c0629fb48b7e9

Observation 29fcdf9c-5ae6-4433-913d-484fa50812f1 · outbound

This paper cites Amt: All-pairs multi-field transforms for efficient frame interpolation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Amt: All-pairs multi-field transforms for efficient frame interpolation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.030716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:05.075194Z digest=sha256:a9d4c1d4167113793e3d61362d5b8fed2777da0f489ac92fbc435f2ce59d6569

Observation 6ba76e2c-012c-4315-96b6-8d829a3a266e · outbound

This paper cites Flowvid: Taming imperfect op- tical flows for consistent video-to-video synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Flowvid: Taming imperfect op- tical flows for consistent video-to-video synthesis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.001903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:05.139965Z digest=sha256:edfa7f6e0e4ebbd123aad3f50252a01269e55e1ddd5e56b0824cb290e675b421

Observation 36f84924-0470-431f-8eb2-8773afda9bb5 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Open-Sora Plan: Open-Source Large Video Generation Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.217568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.217568Z digest=sha256:3451380001096ed8b4442ebec8911d2f96ce7b0d5a5e76eea35b789e2020643e

Observation 8fa1c850-ab78-40ae-a79a-833be2520ff8 · outbound

This paper cites AnimateDiff-Lightning: Cross-Model Diffusion Distillation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnimateDiff-Lightning: Cross-Model Diffusion Distillation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.312282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.312282Z digest=sha256:8cb29e12f77f371832ac4bf477efee95439ce8d1b8f60ad03267c87e771015a3

Observation c554a101-d293-4e2f-bd43-d2de26e0588c · outbound

This paper cites Towards understanding cross and self-attention in stable diffusion for text-guided image editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Towards understanding cross and self-attention in stable diffusion for text-guided image editing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.982327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:05.379869Z digest=sha256:369c8d7536b61f8a7e0e999f0c9eca91bccce909f8bd9e9f947eb42baff371dd

Observation f9958775-c95f-4fb9-9431-f61f11ac0c70 · outbound

This paper cites At- tentive linguistic tracking in diffusion models for training- free text-guided image editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis At- tentive linguistic tracking in diffusion models for training- free text-guided image editing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.956632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:05.466037Z digest=sha256:dd5112cf668cb2c1fd358766049e6709726a847422a7aa694227806f2b986a00

Observation e0d2149e-da78-48e9-9e14-de0bfcdc9884 · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Video-p2p: Video editing with cross-attention control

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.939679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:05.555043Z digest=sha256:e97d8afcc1f31fc58623e732fb4cfd079618b8a6d6a048a466c96736993b29d8

Observation 918aafef-6849-436e-92b0-5db97238a645 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Evalcrafter: Benchmarking and evaluating large video generation models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.920453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:05.625822Z digest=sha256:e6a5e28910ea6fc953324f707934d3f7a0a5754478e5c153e9919a778866cf71

Observation a4c8f94e-c403-4eb8-8480-bbe6e462190b · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.715742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.715742Z digest=sha256:33855f4d29fec446837c73eb960b70dd256bdfac3a4ca496c7e0ca459cd509a0

Observation d72f10aa-8b15-455e-8fe0-e02e0c29693b · outbound

This paper cites Null-text inversion for editing real images using guided diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Null-text inversion for editing real images using guided diffusion models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.897595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:05.801114Z digest=sha256:b92b6190c79e194053aa269b5e701159ba1d1fe1bb712307850427218c6a5e30

Observation 6d940d20-b81a-42e8-97de-901cd1f70086 · outbound

This paper cites Diffusion Models for Adversarial Purification.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Diffusion Models for Adversarial Purification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.886229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.886229Z digest=sha256:c671a05f1b0153bdd8ae3a3462f97b2f5c810efeb180fe89a0653173d50ab969

Observation 6473f36f-99c7-45fe-bfb0-107d9b29ee7f · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:10.647955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:05.967881Z digest=sha256:6a84b86b1b59dd61204d72d95239342de56bfee87a84b105f1ad01608af7e1f5

Observation 15b108a6-e171-4e04-8766-8e853205d8b4 · outbound

This paper cites Pika 1.0.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Pika 1.0

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.396746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:06.036782Z digest=sha256:7bef18075bc3af035378ffb0ca2d8d8226008ca112721313a3066847a2b424b7

Observation f3e0ee66-6b11-45e7-99fe-3ba01e1f947e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.099876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.099876Z digest=sha256:eb5bcf56148a69d310ef4b84cf7837ffefe53bee80cee312c0f481b2dca5e33a

Observation 375ff990-6126-4afb-a6eb-1ae9bb0c884b · outbound

This paper cites Fatezero: Fus- ing attentions for zero-shot text-based video editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Fatezero: Fus- ing attentions for zero-shot text-based video editing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.276468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:06.189085Z digest=sha256:bf18e614fecbd52a30f695d1936e685d71fd6479a088150588c3b9444ce9422a

Observation 76492959-6da0-4ef4-9a55-911f3565b125 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.087093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:06.293597Z digest=sha256:2b0d7f665cd850ee4b22d6171c53e4d4dcafd3c2fde120b6a5f08e520a1c37d7

Observation 907c4bd1-90d0-4049-b627-4946d099121e · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:09.833297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:06.383786Z digest=sha256:4f7915b3be0622fefa7c117161f69a4e0640558e1ad49a2e3bdb89931b4bf99f

Observation a3af56a6-1e74-46b2-95c1-85849e106bab · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:09.717542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:06.484434Z digest=sha256:5c6c0e0915deb7b3d19dd46877d30312e5e19b6292e180041fd52b8271e16bb5

Observation 320ef7f0-0c21-4df9-940f-9781f83cfd51 · outbound

This paper cites Bivdiff: A training-free framework for general-purpose video synthesis via bridging image and video diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Bivdiff: A training-free framework for general-purpose video synthesis via bridging image and video diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.612955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:06.570787Z digest=sha256:d99d3acb1d50d305f9ef3cf1983b5dd6635f46e113f9eb9c708513c753d3b16c

Observation aab08ec3-8f12-4e8a-9e98-8016e78cb9fa · outbound

This paper cites Edit-a-video: Single video editing with object-aware consistency.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Edit-a-video: Single video editing with object-aware consistency

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.422686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:06.666069Z digest=sha256:0f6861ad15bf772e70f01d3b602783b79d8f2e19349bf2488c8e7f20ef5e99a1

Observation b4988830-1ec9-464e-b471-d6a70db25a75 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.746047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.746047Z digest=sha256:b89843c60f5e9270edeb58c3015409a90a08267096a4b1215062a88d3127bc9e

Observation 328a147e-05c2-41d5-81eb-ab35e98dfc4f · outbound

This paper cites Denoising Diffusion Implicit Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Denoising Diffusion Implicit Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.829645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.829645Z digest=sha256:7a57bf8324aca1f061d24f5bf0d2d3bbf31b6dd3bf3666e697677aabb5e75c5e

Observation 8657e463-d5c8-4a85-913c-bcf6d3112a97 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Plug-and-play diffusion features for text-driven image-to-image translation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.249295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:06.915722Z digest=sha256:d18cd8f087e809295a5646d687e93aa1e9ced7563619284d904a2c3e84152f86

Observation 845dbd66-6d58-4869-8af5-3aa2b0ed9879 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis ModelScope Text-to-Video Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.980555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.980555Z digest=sha256:e867d4da1edd48713ab13ed443d898d7b5f5d42639bb19b675770643f09953c8

Observation f6e5d777-157d-4421-9480-31be1a41639a · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.087178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.087178Z digest=sha256:d4709d45df0ae4e3253de5288aeece6278006ed03f6d8bad3664ef34ab69439f

Observation a9a1aa62-1b7b-41d2-9abd-0de5e46b7720 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Exploring video quality assessment on user generated contents from aesthetic and technical perspectives

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.038369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:07.174284Z digest=sha256:0cf24abf0cbac1d1c5df942f34c21ab40a1f163945619758052aa3d1efad92eb

Observation 7cf0cde8-aba3-4dc7-b04d-146c968e51e1 · outbound

This paper cites Tune-a-video: One-shot tun- ing of image diffusion models for text-to-video generation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Tune-a-video: One-shot tun- ing of image diffusion models for text-to-video generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.865062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:07.254277Z digest=sha256:92b20f67b8a35defb61d8225950dd38130066c45c2f94f57906d3d01ae549f89

Observation ffa0efa5-6e01-4291-81c3-745e8b7e73d2 · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Rerender a video: Zero-shot text-guided video-to-video translation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.669898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:07.339187Z digest=sha256:c1376fd89141fb4fa8bbd7a3a4b43a948d2fa35f073b57ca8dff273a118e6c8c

Observation a13928cf-a1bb-4530-9664-af5f101d3778 · outbound

This paper cites Fresco: Spatial-temporal correspondence for zero-shot video translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Fresco: Spatial-temporal correspondence for zero-shot video translation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.509404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:07.427186Z digest=sha256:9db9dc7f01240b26c717ce26831d166db5c2c80f9b2a0f960e988c4b531815d6

Observation 60f5666e-ec18-497c-99ac-9f1789767b36 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.516251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.516251Z digest=sha256:4c4143c5df8d5f8759f6e3f4a99e29175af5297128172520cd2e60698c0b05a2

Observation 97fe92c6-dd19-48d1-8024-14ca4544c834 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.602629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.602629Z digest=sha256:1898d305d30b442a4fabca7fce36c5bcd8d544cfdbb0b70ec4483b648e8b4fb3

Observation 6f4480d7-b3f2-4914-bd8d-0eb6e57c9371 · outbound

This paper cites ControlVideo: Training-free Controllable Text-to-Video Generation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis ControlVideo: Training-free Controllable Text-to-Video Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.699049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.699049Z digest=sha256:b48b7d780e526fd63ad8d7b9469a6ee0eeb3b2ae1cefd2a854d7c2676a9ed657

Observation ec343ff1-a59f-4dbf-ac32-4a867b6a9583 · outbound

This paper cites VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:24:08.016744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:24:07.779000Z digest=sha256:93da2627ce25761015a971e2401e428905324f3e7ad4d2ab4b71016acb6a8b23

Pith citing papers

No inbound Pith citation observations are available.