Pith. sign in

Paper Citation Record · LEDGER

VideoDirector: Precise Video Editing via Text-to-Video Models

As of 21 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 2 inbound Pith citation observations for arXiv:2411.17592.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17592 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:58:46.303484Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:42:45.195531Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:05:51.410933Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98e0d92e-b1a5-42b7-a3f8-08e6318333ed · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

VideoDirector: Precise Video Editing via Text-to-Video Models Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.194791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.194791Z digest=sha256:bd5e6cfb30ffc0e78605ec08cbbf9b5d377b5a051ad793a1f9d62abd55cff7e7

Observation 8174fa4c-cb1c-40c5-86ed-ecdb5361b68a · outbound

This paper cites Stable- video: Text-driven consistency-aware diffusion video edit- ing.

VideoDirector: Precise Video Editing via Text-to-Video Models Stable- video: Text-driven consistency-aware diffusion video edit- ing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.620404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.198632Z digest=sha256:63d4bb8588d2f1fa89272c7702897e3b14a95c253c4d2461fb6ec1b653a380a1

Observation 0213dacc-4fab-4df1-b362-dbbe4b388c06 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

VideoDirector: Precise Video Editing via Text-to-Video Models Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.202244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.202244Z digest=sha256:5b3fb7f035bd2037db184b939cdd8fb0111de6c14d6adf98469a5e32fd43eebe

Observation e1710b0e-3099-41b7-ab53-10e1d5f16af9 · outbound

This paper cites Slicedit: Zero-Shot Video Editing With Text-to-Image Diffusion Models Using Spatio-Temporal Slices.

VideoDirector: Precise Video Editing via Text-to-Video Models Slicedit: Zero-Shot Video Editing With Text-to-Image Diffusion Models Using Spatio-Temporal Slices

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.206188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.206188Z digest=sha256:5d7467163b9ee1e3f733a2312fe52f900a0560a3e4d6aaec55a310406c700f45

Observation d786962f-3d5e-40ff-9d71-bf678f8cc6f6 · outbound

This paper cites Flatten: optical flow-guided attention for consistent text-to-video editing.

VideoDirector: Precise Video Editing via Text-to-Video Models Flatten: optical flow-guided attention for consistent text-to-video editing

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.601177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.209830Z digest=sha256:d213b8a2c72792ed8c47931444f71ca38e77606642dd49a1fc4341ad5ce55ce4

Observation f5608a46-5623-447a-8e30-63a48127dd16 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

VideoDirector: Precise Video Editing via Text-to-Video Models TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.213335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.213335Z digest=sha256:0b57cae59ca7623ddd3380be224982583c469735634c09597f1ae6599f9bfd91

Observation 29d1d450-fd78-4dd9-b2b5-44b41ebc2aa4 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning.

VideoDirector: Precise Video Editing via Text-to-Video Models Animatediff: Animate your personalized text-to- image diffusion models without specific tuning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.591288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.217217Z digest=sha256:d44127aab1540dc152d62b914d2365b0b0feb1c33c6192dd3dcb9bdd71f11a1a

Observation 72f9431b-50a1-4f5b-8099-39a049503e81 · outbound

This paper cites Svdiff: Compact param- eter space for diffusion fine-tuning.

VideoDirector: Precise Video Editing via Text-to-Video Models Svdiff: Compact param- eter space for diffusion fine-tuning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.582024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.221317Z digest=sha256:900fb556b23bd8375e96be529a4ba1009e8398e01f7f1b165a6ba3fdb8e51ccd

Observation 801f69c5-9be5-464d-a768-c846e30b45df · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

VideoDirector: Precise Video Editing via Text-to-Video Models Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.224568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.224568Z digest=sha256:194d52924939142be621e719b368520f8665abd2bb84f2b2f1bd713fc3e3a1d7

Observation d8eb7dd6-ed18-4b1f-8b02-e4a9d1db2ddf · outbound

This paper cites Classifier-Free Diffusion Guidance.

VideoDirector: Precise Video Editing via Text-to-Video Models Classifier-Free Diffusion Guidance

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.228555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.228555Z digest=sha256:beec1cba1f92e2fdeef204ee3cfaab27a2c9f5bc98cdce383783daf6f672ca86

Observation 62324193-b51a-489b-bd9b-760b6d4fda90 · outbound

This paper cites Denoising diffu- sion probabilistic models.

VideoDirector: Precise Video Editing via Text-to-Video Models Denoising diffu- sion probabilistic models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.233005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.233005Z digest=sha256:cb87ff6afc697fd606f093a31452957a124bcd0713649be0bc5ee4d27281dfc4

Observation b712ab7e-462d-498c-8e04-72a5e6cd34a8 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

VideoDirector: Precise Video Editing via Text-to-Video Models Vbench: Comprehensive bench- mark suite for video generative models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.564477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.236972Z digest=sha256:7438725efd2e636a218cb48b3f37ac3ee0e96825cfbf006a5c924fc5f6a9144c

Observation de780c17-f048-41e2-8c57-1b1ea44e0309 · outbound

This paper cites Rave: Randomized noise shuf- fling for fast and consistent video editing with diffusion mod- els.

VideoDirector: Precise Video Editing via Text-to-Video Models Rave: Randomized noise shuf- fling for fast and consistent video editing with diffusion mod- els

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.552474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.240614Z digest=sha256:592ce4cc30e3796e7c05b8c6dc59d038048e7fa2a533ab89e22ac1e031f943c2

Observation 61fc56c2-1909-4500-9f28-477f42597e74 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

VideoDirector: Precise Video Editing via Text-to-Video Models Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.540925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.244522Z digest=sha256:dcbf20376290420c8db47cd1e34637c2df35c8925b8fcbde2485cd2f7a2f9360

Observation 6ffe684d-43e4-40af-8d4f-67ad0a822c9e · outbound

This paper cites Vidtome: Video token merging for zero-shot video editing.

VideoDirector: Precise Video Editing via Text-to-Video Models Vidtome: Video token merging for zero-shot video editing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.530405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.248023Z digest=sha256:2be02f273734445f02b765dc390b92f8d2de779915985e30f79011383ebfb858

Observation 5ec39047-2bae-40ad-b305-7d3efab1a2c6 · outbound

This paper cites MotionClone: Training-Free Motion Cloning for Controllable Video Generation.

VideoDirector: Precise Video Editing via Text-to-Video Models MotionClone: Training-Free Motion Cloning for Controllable Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.250950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.250950Z digest=sha256:5d068813c5b4b19b056e34d036176d2d31e6829752c1468e65bbcfab5e9a66ca

Observation 288e99ea-793a-4c49-ae84-840dd6b6370e · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

VideoDirector: Precise Video Editing via Text-to-Video Models Video-p2p: Video editing with cross-attention control

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.519352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.254495Z digest=sha256:ada81d199cd5d3376f47fd0ed23fbd89ee2a558a621a1868cbed1a32f6ecaa90

Observation 5b9720d6-e457-45f9-afc1-6a5ef03d7435 · outbound

This paper cites Understanding Diffusion Models: A Unified Perspective.

VideoDirector: Precise Video Editing via Text-to-Video Models Understanding Diffusion Models: A Unified Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.257677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.257677Z digest=sha256:4c9b2d3d15e49dbeae142f3fbc43e871853b67ff2b9336cca0b463f87c60d89b

Observation 815a2cf5-07d6-4664-b8de-e9c8c9d2e17c · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

VideoDirector: Precise Video Editing via Text-to-Video Models Latte: Latent Diffusion Transformer for Video Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.260847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.260847Z digest=sha256:f714acc92b26565aed714b6d9884869f0c992215f979c4d4ecb2c99feb31e84c

Observation aa8fd9e2-f6d9-42e2-a905-ff3cce37bf15 · outbound

This paper cites Adapedit: Spatio- temporal guided adaptive editing algorithm for text-based continuity-sensitive image editing.

VideoDirector: Precise Video Editing via Text-to-Video Models Adapedit: Spatio- temporal guided adaptive editing algorithm for text-based continuity-sensitive image editing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.507116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.264784Z digest=sha256:cec480fb191b8ab16b38bb8332437a47dc34c7f26200ee8b78e48d069719bced

Observation 83ec13b4-d6f4-4fa9-9015-05b05abd1a07 · outbound

This paper cites Null-text Inversion for Editing Real Images using Guided Diffusion Models.

VideoDirector: Precise Video Editing via Text-to-Video Models Null-text Inversion for Editing Real Images using Guided Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.268164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.268164Z digest=sha256:de930200b6069bd1dca3c0b6abeeb5b94456038d6dafeaed6933cd0da9ae8b3c

Observation 4d2bf4c9-efb6-4d28-81cf-860f9214219b · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

VideoDirector: Precise Video Editing via Text-to-Video Models A benchmark dataset and evaluation methodology for video object segmentation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.271933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.271933Z digest=sha256:17952f1153d95e2521eb683cdfbe349625f72fc88a14b3544174ed9822efad5e

Observation c229b18a-7299-4c03-858f-91ce0df9dba1 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

VideoDirector: Precise Video Editing via Text-to-Video Models SAM 2: Segment Anything in Images and Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.275301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.275301Z digest=sha256:7b692d820851fb49c13801601cc5db3cf2eb34496ba1b3ee25baf11d2678ffeb

Observation 873368d6-6b1a-4c00-b97b-28273ddc0b59 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

VideoDirector: Precise Video Editing via Text-to-Video Models High-resolution image syn- thesis with latent diffusion models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.278803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.278803Z digest=sha256:12feeb1a366c83e292cb42c9d1d409cd32dc8e60863f00ca07f8ae56eedd84c7

Observation 962767b4-e45c-4574-bf3f-f4ccd677bcc1 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

VideoDirector: Precise Video Editing via Text-to-Video Models Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.482469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.282256Z digest=sha256:654cd3c3a63fda7558481a67f102625f4abf2576f390907ac00ac8e9741c9ba3

Observation b457fbec-045d-48b7-b8b3-39191b3bac1a · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

VideoDirector: Precise Video Editing via Text-to-Video Models Photorealistic text-to-image diffusion models with deep language understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.285699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.285699Z digest=sha256:c6e2b7f6f0ee59b3f5d147afb6eb5325cc652b63f90756c2a67c2cb378c08c55

Observation 668ecfce-e69c-4d98-9a23-5b293bb9c210 · outbound

This paper cites 9 Dragdiffusion: Harnessing diffusion models for interactive point-based image editing.

VideoDirector: Precise Video Editing via Text-to-Video Models 9 Dragdiffusion: Harnessing diffusion models for interactive point-based image editing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.465684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.289016Z digest=sha256:ce5f2c95fafb7394e24f46435fb06022b2cdd65409a27a5973851da66f96a2c4

Observation 8cd46d38-c505-4422-8163-c74c0f023396 · outbound

This paper cites Denoising Diffusion Implicit Models.

VideoDirector: Precise Video Editing via Text-to-Video Models Denoising Diffusion Implicit Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.292519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.292519Z digest=sha256:a7415c678acfa9515f3529429bd7ac9d8cd3fb23d94176dece3567fb36da0d27

Observation 2a41b696-c44c-46a0-8673-f91274ac9a35 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

VideoDirector: Precise Video Editing via Text-to-Video Models Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:58:46.296444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:58:46.296444Z digest=sha256:64c511ae04b763c5cd97e08721a15c65479286c0744501b9cd15ca52443af553

Observation afca32cf-956d-462f-ba50-11272d8f6189 · outbound

This paper cites Turboedit: Instant text-based image editing.

VideoDirector: Precise Video Editing via Text-to-Video Models Turboedit: Instant text-based image editing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.448759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.300045Z digest=sha256:3f327c60b5381a2e777df9d54e542b075304dafb3d56b55ff131e0853fc408b2

Observation f945a147-0be0-4ba0-b226-a7c77d4e15f8 · outbound

This paper cites Camel: Causal motion enhance- ment tailored for lifting text-driven video editing.

VideoDirector: Precise Video Editing via Text-to-Video Models Camel: Causal motion enhance- ment tailored for lifting text-driven video editing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:58:46.437997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:58:46.303484Z digest=sha256:c12c452fa6150cdd00b0d055e7084294e3c000f2f629799dab1f8180d242974d

Pith citing papers

Observation 09d412d2-2378-4441-87c3-dba6d3c2a2d3 · inbound

DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing cites this paper.

DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing VideoDirector: Precise Video Editing via Text-to-Video Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:42:45.195531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:42:45.195531Z digest=sha256:4cf88c76d175710f92b690eba15d13ce3749c7ba900900c5259ec3f74e59c18c

Observation 49e79ce3-063b-46b9-a29a-c9cfec8ca42b · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations VideoDirector: Precise Video Editing via Text-to-Video Models

Reference 263

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.412791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:c4a67ac6f097e5f775fbd4822f26e44903b9fbd6c985a9bb61befea370276fec