Pith. sign in

Paper Citation Record · LEDGER

CoT-Edit: Let CoT Guide Instruction Video Editing

As of 16 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2608.01113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01113 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:14:36.146697Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 657b10a4-a0bb-42b0-8f78-5823b143d706 · outbound

This paper cites Scaling instruction-based video editing with a high-quality synthetic dataset.arXiv preprint arXiv:2510.15742, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Scaling instruction-based video editing with a high-quality synthetic dataset.arXiv preprint arXiv:2510.15742, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.975532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.975532Z digest=sha256:256fc04f343849044cf28d1c8d273ad8ef4f36952a8864c521c868e2cf63ea8e

Observation 1856cfa5-fccf-4d0a-825a-d496f3c5a391 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

CoT-Edit: Let CoT Guide Instruction Video Editing In- structpix2pix: Learning to follow image editing instructions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.987406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:35.979280Z digest=sha256:fbf7df7cca49cf5e71f7ff123e94ddc2c257967dc759cf33b9415af4a5fc5e2b

Observation 408e0bc6-58b0-4827-b1be-96613547c5e7 · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.982238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.982238Z digest=sha256:ff0ffc0ed97643ab1b6f9b08259d6b7cb47a9fcf767dfac40715dcb75aeb72ba

Observation 18f9807a-5d72-4114-a3c3-70770a3eedaf · outbound

This paper cites Consistent Video-to-Video Transfer Using Synthetic Dataset.

CoT-Edit: Let CoT Guide Instruction Video Editing Consistent Video-to-Video Transfer Using Synthetic Dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.986744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.986744Z digest=sha256:7e91510271bbc5106599d9ad2e070852dfea8193ba4357b0b6fc5a0412595760

Observation b75ff8c6-70e0-4f0d-9f67-4e751ab27f24 · outbound

This paper cites Rass: Improving denoising diffusion sam- plers with reinforced active sampling scheduler.

CoT-Edit: Let CoT Guide Instruction Video Editing Rass: Improving denoising diffusion sam- plers with reinforced active sampling scheduler

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.976630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:35.990411Z digest=sha256:914275999a83d52ab262f53d51b74e7468934317aa0a2290cc21c9fae61e539f

Observation db7eded1-6605-4037-9244-54308783348a · outbound

This paper cites Why compress what you can generate? when gpt-4o generation ushers in image compression fields.

CoT-Edit: Let CoT Guide Instruction Video Editing Why compress what you can generate? when gpt-4o generation ushers in image compression fields

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.966130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:35.994108Z digest=sha256:d6e4bc104f65d2d037c6ef0dd51439258b6d3a1a3ee0f8eabbeda2abb4152957

Observation c7e8d454-32fa-46ec-9c31-56a81201e8ce · outbound

This paper cites SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing.

CoT-Edit: Let CoT Guide Instruction Video Editing SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.997634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.997634Z digest=sha256:d513a12e63a1afc10d02266f26a2b1f172945c1d6df0fbdda455ae4b37df8c62

Observation eb512c56-8fb9-4959-b4af-05742885a328 · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

CoT-Edit: Let CoT Guide Instruction Video Editing Clipscore: A reference-free evaluation met- ric for image captioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.953342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.001663Z digest=sha256:1f3b5f18042807d52648e069b6658e4107aaccbee9dc22fe821f1eb60baae4dc

Observation 48d98d9a-6c21-44d9-86c0-995bf19d8587 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing Imagen Video: High Definition Video Generation with Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.005696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.005696Z digest=sha256:02ed8ba8ae87a5d342b5e783d3c1b99fa85ec84457193512ac55c4fb4442b71e

Observation f8cba0fc-3fa7-4c3d-93bf-6345952e11fa · outbound

This paper cites Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646, 2022.

CoT-Edit: Let CoT Guide Instruction Video Editing Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.009394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.009394Z digest=sha256:bccdb98cfc2897a79c210bf8996ff5e70cd3b7ad4a192c7bee1baffb26d1c987

Observation 25ef2ba6-7d42-469c-831d-7876d2fdfc32 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

CoT-Edit: Let CoT Guide Instruction Video Editing CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.012707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.012707Z digest=sha256:0307ae1f1ae055c27bb4fc6068b72f8f9dc7b6b543c55376c9ba6f119b1b9d7e

Observation c1b5049f-49ca-4103-864e-8cf6af35b0b1 · outbound

This paper cites Animate anyone: Consistent and controllable image- to-video synthesis for character animation.

CoT-Edit: Let CoT Guide Instruction Video Editing Animate anyone: Consistent and controllable image- to-video synthesis for character animation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.934410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.016317Z digest=sha256:3351fc66a25858034e89db5b053b5e55c30873c1831146f8cea9a7b285df0639

Observation 49884281-ad6b-42ba-a8a2-3f56dd32d611 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.019799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.019799Z digest=sha256:09084bd14cea7ef514025114bd859d7dbed1abacfb357169c768f8ce58310254

Observation 7cde7cc7-1d24-4b9c-aac6-7a87e4f26e5f · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

CoT-Edit: Let CoT Guide Instruction Video Editing Vbench: Comprehensive bench- mark suite for video generative models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.023233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.023233Z digest=sha256:aa7a327c6fde0e9d4c14f9ba023c454df92ed86be0cdd2af24f34d3a978263a2

Observation 4dd9d16c-bfc1-46db-894b-b7216878607e · outbound

This paper cites Anyedit: Edit any knowledge encoded in language models.arXiv preprint arXiv:2502.05628, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Anyedit: Edit any knowledge encoded in language models.arXiv preprint arXiv:2502.05628, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.026401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.026401Z digest=sha256:86509051c9fada43e7c5070607ed65b5fe6fce8c9555a79d10f3c5e1a339b931

Observation c4b2e429-8f5b-4a6f-b212-61d7d3bcaf88 · outbound

This paper cites Vace: All-in-one video creation and editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Vace: All-in-one video creation and editing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.029525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.029525Z digest=sha256:031498eb5c67959628f4f67c43bb82a2ae95a6e526a26aec3ebc94867eacb172

Observation e2c88d85-f4ca-499b-becf-0f87bb95e566 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

CoT-Edit: Let CoT Guide Instruction Video Editing Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.032244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.032244Z digest=sha256:de7511f732bc27db6aab012cdc589a20868a631ef9edb79162e2deb7c83c7a39

Observation 0242e969-b0d2-454e-8224-5ea3437cee65 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.035188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.035188Z digest=sha256:464b45d5a68dcf10bc90371c5d958650d084251b0c339fb5a67e2e7a751fab3f

Observation 05eb5eef-d43d-4443-912a-ad663df13610 · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

CoT-Edit: Let CoT Guide Instruction Video Editing AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.039061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.039061Z digest=sha256:e8f4fb89d05036ba598d55fba33ff67ccdad321c0797e9d6f4cff93003bace79

Observation e3f22532-7f1b-41d9-85ae-d90033f338a7 · outbound

This paper cites OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation.

CoT-Edit: Let CoT Guide Instruction Video Editing OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.042340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.042340Z digest=sha256:5fe6912b1d5c489ce93022ceef68a4568f9822edc9ed06e0e2adf2f8134f8f62

Observation 4c7772c1-8450-4621-a2fe-7362429c370a · outbound

This paper cites Grounding 3D Scene Affordance From Egocentric Interactions.

CoT-Edit: Let CoT Guide Instruction Video Editing Grounding 3D Scene Affordance From Egocentric Interactions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.046347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.046347Z digest=sha256:470120425f521bb86d44b6f56511262c2707121c0bc04f76672ddafc87bebeb2

Observation b1fa8719-768e-4bd2-bd97-7155ad412c6f · outbound

This paper cites Stablev2v: Stabilizing shape consistency in video-to- video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Stablev2v: Stabilizing shape consistency in video-to- video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.904108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.049901Z digest=sha256:88eefc0bc043321c8d94b61dff9f70ec901d3ceb628bbee0d2cdc4e9f74fdcd6

Observation 409ef355-6b9c-41e2-945a-6a7dd9d41699 · outbound

This paper cites The health- wealth gradient in labor markets: Integrating health, in- surance, and social metrics to predict employment density.

CoT-Edit: Let CoT Guide Instruction Video Editing The health- wealth gradient in labor markets: Integrating health, in- surance, and social metrics to predict employment density

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.893023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.052062Z digest=sha256:c242c69d72a8c766b7f1d7339f52c1a239be7134d6655b6f9fd1cea56af47c54

Observation 4f40e9e3-6641-4c18-a16c-9f04c1aa84a9 · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

CoT-Edit: Let CoT Guide Instruction Video Editing Video-p2p: Video editing with cross-attention control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.054760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.054760Z digest=sha256:b064d41ad6d381f8835f75981f1a1df28ead8d671ea19e2ca0af37277b45e1a9

Observation 34839ddd-5a94-46ac-b8ce-f8a987f2cca6 · outbound

This paper cites Follow your pose: Pose- guided text-to-video generation using pose-free videos.

CoT-Edit: Let CoT Guide Instruction Video Editing Follow your pose: Pose- guided text-to-video generation using pose-free videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.057103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.057103Z digest=sha256:2d60e6da877d6de430aba9ee6df72c7e14912126aefbe68f8f4125d435560b8e

Observation 5be0a624-c720-4b35-ab97-43354a0f784d · outbound

This paper cites Magic- stick: Controllable video editing via control handle transfor- mations.

CoT-Edit: Let CoT Guide Instruction Video Editing Magic- stick: Controllable video editing via control handle transfor- mations

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.868611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.059802Z digest=sha256:3d90346a79b72edb45fe595ea90901810e9db90cc30cd337c0889178f0cda43e

Observation 6d48c170-3a29-4bf3-98e0-67d67d71e691 · outbound

This paper cites In- structx: Towards unified visual editing with mllm guidance.

CoT-Edit: Let CoT Guide Instruction Video Editing In- structx: Towards unified visual editing with mllm guidance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.062312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.062312Z digest=sha256:4c337c95a26c6c5b5b25ac1d0c13f9a7d56951916d1d3a1a449f0e1e4feb617d

Observation d2585ad4-dec7-4090-9764-6c2391d8a21d · outbound

This paper cites Occluded video instance segmentation: A bench- mark.International Journal of Computer Vision, 130(8): 2022–2039, 2022.

CoT-Edit: Let CoT Guide Instruction Video Editing Occluded video instance segmentation: A bench- mark.International Journal of Computer Vision, 130(8): 2022–2039, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.855427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.064649Z digest=sha256:dbcfb549a72f95338f0a97049b57b2a06061c0b0e23d4a5335c6d516f64c51ac

Observation 98cde1d6-e308-4324-ab13-0f0a5d78d10a · outbound

This paper cites Instructvid2vid: Controllable video editing with natural language instructions.

CoT-Edit: Let CoT Guide Instruction Video Editing Instructvid2vid: Controllable video editing with natural language instructions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.843912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.067106Z digest=sha256:ddfb811ddc4ce7fe8660f723fd1fb489165dba8cf65a7c8ee2de584be0c678e1

Observation 64bf5aa5-8938-4b18-8474-807bb8e824c3 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

CoT-Edit: Let CoT Guide Instruction Video Editing Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.832657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.069273Z digest=sha256:ff610975e0952df19da475a03141e97884dfbc541f60c1c38f17f2e26114ddab

Observation 2141106f-c39d-43ae-bfa7-3cc3fee87fb2 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

CoT-Edit: Let CoT Guide Instruction Video Editing Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.071477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.071477Z digest=sha256:a9f6b00235e6a526412b5d3288534361e41b732ba48204ec391e460da18f776c

Observation 3fcd91b3-7f33-49e9-8c2f-baff1f40219b · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

CoT-Edit: Let CoT Guide Instruction Video Editing Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.821389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.075387Z digest=sha256:4f8cf9eea8bdc207c1601d5997a539d9cd13baf434f7724459f3e130d5437f08

Observation 83fff1b4-c3a0-4f43-bf40-4a6491434b31 · outbound

This paper cites Omni-video: Democratizing uni- fied video understanding and generation.arXiv preprint arXiv:2507.06119, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Omni-video: Democratizing uni- fied video understanding and generation.arXiv preprint arXiv:2507.06119, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.078800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.078800Z digest=sha256:31e0be22f3668f6338ffa35b7830578cb937c0e3cb6be32a7f4cc90fc8f808e2

Observation af04fc58-2e23-4c2a-bbe7-cc285382811b · outbound

This paper cites Lucy edit: Open-weight text-guided video editing, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Lucy edit: Open-weight text-guided video editing, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.811263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.081791Z digest=sha256:9d0b7e5b5a44086a8e2d909e24564bfaed56b20505885ff31d118956e62748be

Observation d0dc72c7-8972-4b1e-bc7e-45e9454f15ff · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CoT-Edit: Let CoT Guide Instruction Video Editing Gemini: A Family of Highly Capable Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.085011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.085011Z digest=sha256:30c9ff40c087e58abf0d18b921aa02355336b6613a992e944071a563913b12b7

Observation 57023dca-7e67-4dd7-8704-606aa4574292 · outbound

This paper cites Fvd: A new metric for video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Fvd: A new metric for video generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.800032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.088655Z digest=sha256:96beade541e17d4748aa99bdc3b82ee0a18a53b6e557bab625ebfc931d05e073

Observation 9ab29126-2f77-4800-ae0e-3ff4fbd845f6 · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

CoT-Edit: Let CoT Guide Instruction Video Editing Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.091525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.091525Z digest=sha256:f1b2b5dc79f2c0f09035a4fe0d91c0db0ba01033d78d3661e126b75595fbcf5e

Observation 3cb4d851-d6ff-425f-9239-27bf6f9efad8 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

CoT-Edit: Let CoT Guide Instruction Video Editing ModelScope Text-to-Video Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.096079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.096079Z digest=sha256:b53a5b31328ebe617a725d492a687c70cd586e44b7cc5f0e81f5158b489a25b8

Observation e663e69f-bf22-4cbb-9f2b-755e9db47ae8 · outbound

This paper cites Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content.

CoT-Edit: Let CoT Guide Instruction Video Editing Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.788680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.099979Z digest=sha256:d4d191e46b8eb953f50fddbe5fd536cd5a59bf66f076567ee5eaeb9f1cb41abf

Observation 8fb1b249-ffac-4e93-a955-c98c5a7a2aa0 · outbound

This paper cites Tiv-diffusion: Towards object-centric movement for text-driven image to video gen- eration.

CoT-Edit: Let CoT Guide Instruction Video Editing Tiv-diffusion: Towards object-centric movement for text-driven image to video gen- eration

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.778051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.103092Z digest=sha256:a003bc3fe5616cc53456b7a45df567728125298db8d69528ff273811a01db442

Observation f013a9ff-17b8-40ca-9547-f6c71f7c7fc6 · outbound

This paper cites Re-attentional con- trollable video diffusion editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Re-attentional con- trollable video diffusion editing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.765446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.106568Z digest=sha256:cbab9b1e568c7da48d3a4911b5434003561d9f6e7f2329c903369bff8b08fa70

Observation 43379ae1-bdc5-4091-8111-ffe361a8a947 · outbound

This paper cites Training-free controllable text-guided video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2026.

CoT-Edit: Let CoT Guide Instruction Video Editing Training-free controllable text-guided video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2026

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.755803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.109921Z digest=sha256:45717fa7f69630a0078ef077bf5f0646697e76b3e57347afc85665ad0473b51e

Observation d7a7aec6-83e1-4638-adb3-63faecad4f77 · outbound

This paper cites Phrasecut: Language-based image segmen- tation in the wild.

CoT-Edit: Let CoT Guide Instruction Video Editing Phrasecut: Language-based image segmen- tation in the wild

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.745577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.112795Z digest=sha256:c71c8d555e24a45fd0ebdd2320ce15096b9a41052e58012b70e3a649ee54e963

Observation c685c6db-e0dd-4fca-b8ee-7b27984456af · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.733642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.115970Z digest=sha256:04a723a70cc2fd4fa4a7809b467eb703506a9556f37cc080f284ca678365aa60

Observation 2f70e7ea-a7e9-4628-b819-5d9dc8fb09d1 · outbound

This paper cites Insvie-1m: Effective instruction-based video editing with elaborate dataset construction.

CoT-Edit: Let CoT Guide Instruction Video Editing Insvie-1m: Effective instruction-based video editing with elaborate dataset construction

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.721670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.119760Z digest=sha256:b33356e8b3364c83df04c7bcc025d26db1cfe15a569f6a1bded262d964b14d0d

Observation 67b4273c-b3c0-4edc-8f4b-aa0bce3e75e6 · outbound

This paper cites Veg- gie: Instructional editing and reasoning video concepts with grounded generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Veg- gie: Instructional editing and reasoning video concepts with grounded generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.122798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.122798Z digest=sha256:f20f80f5ca3423271023fbd7e32601e7023b7e32e5f489a342536d0e045c0c15

Observation 222f1b9e-0d29-495a-98a4-845715a0059e · outbound

This paper cites Editworld: Simulating world dynamics for instruction- following image editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Editworld: Simulating world dynamics for instruction- following image editing

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.704257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.125635Z digest=sha256:042a58ea59e6ffa60de49ad65b4418ec64acc50f1d8bd276f8b825ffadb4c74a

Observation 8ba8b5ea-f2e3-400f-bd20-5da9728d9eab · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.128527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.128527Z digest=sha256:76bda0ac10da2ddf6fab2e23dca644476c73bb42e55862c0976019e87da3560a

Observation 89c0ee69-c1c6-411c-8e1c-6e5af7d15788 · outbound

This paper cites EffiVED:Efficient Video Editing via Text-instruction Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing EffiVED:Efficient Video Editing via Text-instruction Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.132218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.132218Z digest=sha256:1efb771c0980e9e6a6f93d44e34606f8ca566a948980232da07eeb30987a8682

Observation 106458a7-e319-4faf-9c0d-7a77c611b3f4 · outbound

This paper cites Motionpro: A precise mo- tion controller for image-to-video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Motionpro: A precise mo- tion controller for image-to-video generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.693440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.135815Z digest=sha256:13ddbef2b13c6f9bb593c8b525cc8a32832b0bca43ceec1407d60d38a9f57e4a

Observation 892ca003-1808-4665-9ca2-76f4c6b06c01 · outbound

This paper cites Ultraedit: Instruction-based fine-grained im- age editing at scale.Advances in Neural Information Pro- cessing Systems, 37:3058–3093, 2024.

CoT-Edit: Let CoT Guide Instruction Video Editing Ultraedit: Instruction-based fine-grained im- age editing at scale.Advances in Neural Information Pro- cessing Systems, 37:3058–3093, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.681969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.139111Z digest=sha256:1f6b87a75b79a8427df82c82d4bc95ed120ad5d2273fcb7be46a7fd312dac60d

Observation 6937549a-7058-4ae0-a562-31ce580c7c73 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.International journal of computer vision, 127(3):302–321, 2019.

CoT-Edit: Let CoT Guide Instruction Video Editing Semantic under- standing of scenes through the ade20k dataset.International journal of computer vision, 127(3):302–321, 2019

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.670892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:14:36.142912Z digest=sha256:de308c9e1f7375c62a052e707a0f4c9f776c3b479c9f844496122a9cac67cb3c

Observation f2bbfc9a-9c3a-44b6-b89a-7f1cfbfb6813 · outbound

This paper cites Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists.

CoT-Edit: Let CoT Guide Instruction Video Editing Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.146697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.146697Z digest=sha256:fb67fc291727317525f391b0082c7d05cb88c5c0cba36cbeb954989b94b8bd46

Pith citing papers

No inbound Pith citation observations are available.