Pith. sign in

Paper Citation Record · LEDGER

CoT-Edit: Let CoT Guide Instruction Video Editing

As of 16 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2608.01113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01113 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:14:36.146697Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 657b10a4-a0bb-42b0-8f78-5823b143d706 · outbound

This paper cites Scaling instruction-based video editing with a high-quality synthetic dataset.arXiv preprint arXiv:2510.15742, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Scaling instruction-based video editing with a high-quality synthetic dataset.arXiv preprint arXiv:2510.15742, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.975532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.975532Z digest=sha256:256fc04f343849044cf28d1c8d273ad8ef4f36952a8864c521c868e2cf63ea8e

Observation 1856cfa5-fccf-4d0a-825a-d496f3c5a391 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

CoT-Edit: Let CoT Guide Instruction Video Editing In- structpix2pix: Learning to follow image editing instructions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.987406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:35.979280Z digest=sha256:092af8b44556500ff1e18cc8c3a8fc78797db8a7bf734bab8d5f79d06c3f7c41

Observation 408e0bc6-58b0-4827-b1be-96613547c5e7 · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.982238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.982238Z digest=sha256:ff0ffc0ed97643ab1b6f9b08259d6b7cb47a9fcf767dfac40715dcb75aeb72ba

Observation 18f9807a-5d72-4114-a3c3-70770a3eedaf · outbound

This paper cites Consistent Video-to-Video Transfer Using Synthetic Dataset.

CoT-Edit: Let CoT Guide Instruction Video Editing Consistent Video-to-Video Transfer Using Synthetic Dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.986744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.986744Z digest=sha256:7e91510271bbc5106599d9ad2e070852dfea8193ba4357b0b6fc5a0412595760

Observation b75ff8c6-70e0-4f0d-9f67-4e751ab27f24 · outbound

This paper cites Rass: Improving denoising diffusion sam- plers with reinforced active sampling scheduler.

CoT-Edit: Let CoT Guide Instruction Video Editing Rass: Improving denoising diffusion sam- plers with reinforced active sampling scheduler

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.976630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:35.990411Z digest=sha256:72f08a09b5acdaca0e44bef3aeb6ea831e3706b61c0582bcd2d3cb0fcb89577e

Observation db7eded1-6605-4037-9244-54308783348a · outbound

This paper cites Why compress what you can generate? when gpt-4o generation ushers in image compression fields.

CoT-Edit: Let CoT Guide Instruction Video Editing Why compress what you can generate? when gpt-4o generation ushers in image compression fields

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.966130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:35.994108Z digest=sha256:b6022301bbf4b155e57c01bb8806f20864b217c414c525ddfde29d441a0b7c92

Observation c7e8d454-32fa-46ec-9c31-56a81201e8ce · outbound

This paper cites SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing.

CoT-Edit: Let CoT Guide Instruction Video Editing SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.997634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.997634Z digest=sha256:d513a12e63a1afc10d02266f26a2b1f172945c1d6df0fbdda455ae4b37df8c62

Observation eb512c56-8fb9-4959-b4af-05742885a328 · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

CoT-Edit: Let CoT Guide Instruction Video Editing Clipscore: A reference-free evaluation met- ric for image captioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.953342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.001663Z digest=sha256:068a3d21084d80046f94841eeb22407583bfa665e7fc3284ffcdc70af6b29fb0

Observation 48d98d9a-6c21-44d9-86c0-995bf19d8587 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing Imagen Video: High Definition Video Generation with Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.005696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.005696Z digest=sha256:02ed8ba8ae87a5d342b5e783d3c1b99fa85ec84457193512ac55c4fb4442b71e

Observation f8cba0fc-3fa7-4c3d-93bf-6345952e11fa · outbound

This paper cites Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646, 2022.

CoT-Edit: Let CoT Guide Instruction Video Editing Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.009394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.009394Z digest=sha256:bccdb98cfc2897a79c210bf8996ff5e70cd3b7ad4a192c7bee1baffb26d1c987

Observation 25ef2ba6-7d42-469c-831d-7876d2fdfc32 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

CoT-Edit: Let CoT Guide Instruction Video Editing CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.012707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.012707Z digest=sha256:0307ae1f1ae055c27bb4fc6068b72f8f9dc7b6b543c55376c9ba6f119b1b9d7e

Observation c1b5049f-49ca-4103-864e-8cf6af35b0b1 · outbound

This paper cites Animate anyone: Consistent and controllable image- to-video synthesis for character animation.

CoT-Edit: Let CoT Guide Instruction Video Editing Animate anyone: Consistent and controllable image- to-video synthesis for character animation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.934410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.016317Z digest=sha256:9b5127e539d915869266f16e08b0e0af6d99840b91a08281af5d932558e06fc3

Observation 49884281-ad6b-42ba-a8a2-3f56dd32d611 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.019799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.019799Z digest=sha256:09084bd14cea7ef514025114bd859d7dbed1abacfb357169c768f8ce58310254

Observation 7cde7cc7-1d24-4b9c-aac6-7a87e4f26e5f · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

CoT-Edit: Let CoT Guide Instruction Video Editing Vbench: Comprehensive bench- mark suite for video generative models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.023233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.023233Z digest=sha256:aa7a327c6fde0e9d4c14f9ba023c454df92ed86be0cdd2af24f34d3a978263a2

Observation 4dd9d16c-bfc1-46db-894b-b7216878607e · outbound

This paper cites Anyedit: Edit any knowledge encoded in language models.arXiv preprint arXiv:2502.05628, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Anyedit: Edit any knowledge encoded in language models.arXiv preprint arXiv:2502.05628, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.026401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.026401Z digest=sha256:86509051c9fada43e7c5070607ed65b5fe6fce8c9555a79d10f3c5e1a339b931

Observation c4b2e429-8f5b-4a6f-b212-61d7d3bcaf88 · outbound

This paper cites Vace: All-in-one video creation and editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Vace: All-in-one video creation and editing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.029525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.029525Z digest=sha256:031498eb5c67959628f4f67c43bb82a2ae95a6e526a26aec3ebc94867eacb172

Observation e2c88d85-f4ca-499b-becf-0f87bb95e566 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

CoT-Edit: Let CoT Guide Instruction Video Editing Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.032244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.032244Z digest=sha256:de7511f732bc27db6aab012cdc589a20868a631ef9edb79162e2deb7c83c7a39

Observation 0242e969-b0d2-454e-8224-5ea3437cee65 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.035188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.035188Z digest=sha256:464b45d5a68dcf10bc90371c5d958650d084251b0c339fb5a67e2e7a751fab3f

Observation 05eb5eef-d43d-4443-912a-ad663df13610 · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

CoT-Edit: Let CoT Guide Instruction Video Editing AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.039061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.039061Z digest=sha256:e8f4fb89d05036ba598d55fba33ff67ccdad321c0797e9d6f4cff93003bace79

Observation e3f22532-7f1b-41d9-85ae-d90033f338a7 · outbound

This paper cites OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation.

CoT-Edit: Let CoT Guide Instruction Video Editing OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.042340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.042340Z digest=sha256:5fe6912b1d5c489ce93022ceef68a4568f9822edc9ed06e0e2adf2f8134f8f62

Observation 4c7772c1-8450-4621-a2fe-7362429c370a · outbound

This paper cites Grounding 3D Scene Affordance From Egocentric Interactions.

CoT-Edit: Let CoT Guide Instruction Video Editing Grounding 3D Scene Affordance From Egocentric Interactions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.046347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.046347Z digest=sha256:470120425f521bb86d44b6f56511262c2707121c0bc04f76672ddafc87bebeb2

Observation b1fa8719-768e-4bd2-bd97-7155ad412c6f · outbound

This paper cites Stablev2v: Stabilizing shape consistency in video-to- video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Stablev2v: Stabilizing shape consistency in video-to- video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.904108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.049901Z digest=sha256:11573166d8c659178de88fffccfdcb13983a0910ba5b1117d595623b195dc876

Observation 409ef355-6b9c-41e2-945a-6a7dd9d41699 · outbound

This paper cites The health- wealth gradient in labor markets: Integrating health, in- surance, and social metrics to predict employment density.

CoT-Edit: Let CoT Guide Instruction Video Editing The health- wealth gradient in labor markets: Integrating health, in- surance, and social metrics to predict employment density

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.893023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.052062Z digest=sha256:1bc55a14647c765221770348709538da2d002b0118c1b5a0818026dfe92f1cc6

Observation 4f40e9e3-6641-4c18-a16c-9f04c1aa84a9 · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

CoT-Edit: Let CoT Guide Instruction Video Editing Video-p2p: Video editing with cross-attention control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.054760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.054760Z digest=sha256:b064d41ad6d381f8835f75981f1a1df28ead8d671ea19e2ca0af37277b45e1a9

Observation 34839ddd-5a94-46ac-b8ce-f8a987f2cca6 · outbound

This paper cites Follow your pose: Pose- guided text-to-video generation using pose-free videos.

CoT-Edit: Let CoT Guide Instruction Video Editing Follow your pose: Pose- guided text-to-video generation using pose-free videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.057103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.057103Z digest=sha256:2d60e6da877d6de430aba9ee6df72c7e14912126aefbe68f8f4125d435560b8e

Observation 5be0a624-c720-4b35-ab97-43354a0f784d · outbound

This paper cites Magic- stick: Controllable video editing via control handle transfor- mations.

CoT-Edit: Let CoT Guide Instruction Video Editing Magic- stick: Controllable video editing via control handle transfor- mations

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.868611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.059802Z digest=sha256:6c847a816dcf41308351268e96cd9e9867b12e920c67a42b8875ae4ee8724de1

Observation 6d48c170-3a29-4bf3-98e0-67d67d71e691 · outbound

This paper cites In- structx: Towards unified visual editing with mllm guidance.

CoT-Edit: Let CoT Guide Instruction Video Editing In- structx: Towards unified visual editing with mllm guidance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.062312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.062312Z digest=sha256:4c337c95a26c6c5b5b25ac1d0c13f9a7d56951916d1d3a1a449f0e1e4feb617d

Observation d2585ad4-dec7-4090-9764-6c2391d8a21d · outbound

This paper cites Occluded video instance segmentation: A bench- mark.International Journal of Computer Vision, 130(8): 2022–2039, 2022.

CoT-Edit: Let CoT Guide Instruction Video Editing Occluded video instance segmentation: A bench- mark.International Journal of Computer Vision, 130(8): 2022–2039, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.855427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.064649Z digest=sha256:ec9ce13b9e237f9659c365989804c1def5bac2c2dd8a7b6215698598a8783039

Observation 98cde1d6-e308-4324-ab13-0f0a5d78d10a · outbound

This paper cites Instructvid2vid: Controllable video editing with natural language instructions.

CoT-Edit: Let CoT Guide Instruction Video Editing Instructvid2vid: Controllable video editing with natural language instructions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.843912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.067106Z digest=sha256:8d3854bf089a3ed72d7cba06d6a0c6082ca94999efc3781db045eb584ffff224

Observation 64bf5aa5-8938-4b18-8474-807bb8e824c3 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

CoT-Edit: Let CoT Guide Instruction Video Editing Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.832657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.069273Z digest=sha256:6839bcd5cbba73cdcadfd5014dcf12188d600fae0ef7e656ab7bd37b46462253

Observation 2141106f-c39d-43ae-bfa7-3cc3fee87fb2 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

CoT-Edit: Let CoT Guide Instruction Video Editing Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.071477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.071477Z digest=sha256:a9f6b00235e6a526412b5d3288534361e41b732ba48204ec391e460da18f776c

Observation 3fcd91b3-7f33-49e9-8c2f-baff1f40219b · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

CoT-Edit: Let CoT Guide Instruction Video Editing Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.821389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.075387Z digest=sha256:c0ff764be8fb6f8c807a76b6fdfe52c52b9c7248034220f8d5dc7e4574d25080

Observation 83fff1b4-c3a0-4f43-bf40-4a6491434b31 · outbound

This paper cites Omni-video: Democratizing uni- fied video understanding and generation.arXiv preprint arXiv:2507.06119, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Omni-video: Democratizing uni- fied video understanding and generation.arXiv preprint arXiv:2507.06119, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.078800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.078800Z digest=sha256:31e0be22f3668f6338ffa35b7830578cb937c0e3cb6be32a7f4cc90fc8f808e2

Observation af04fc58-2e23-4c2a-bbe7-cc285382811b · outbound

This paper cites Lucy edit: Open-weight text-guided video editing, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Lucy edit: Open-weight text-guided video editing, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.811263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.081791Z digest=sha256:694e3ca1506d51486f14eac56a9d572647d29cf2f93e67a748599fec85580a16

Observation d0dc72c7-8972-4b1e-bc7e-45e9454f15ff · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CoT-Edit: Let CoT Guide Instruction Video Editing Gemini: A Family of Highly Capable Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.085011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.085011Z digest=sha256:30c9ff40c087e58abf0d18b921aa02355336b6613a992e944071a563913b12b7

Observation 57023dca-7e67-4dd7-8704-606aa4574292 · outbound

This paper cites Fvd: A new metric for video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Fvd: A new metric for video generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.800032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.088655Z digest=sha256:f21b73ebd3901a4318de75811b60e11c582df47ec3e3d91e9801d891f5d31e07

Observation 9ab29126-2f77-4800-ae0e-3ff4fbd845f6 · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

CoT-Edit: Let CoT Guide Instruction Video Editing Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.091525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.091525Z digest=sha256:f1b2b5dc79f2c0f09035a4fe0d91c0db0ba01033d78d3661e126b75595fbcf5e

Observation 3cb4d851-d6ff-425f-9239-27bf6f9efad8 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

CoT-Edit: Let CoT Guide Instruction Video Editing ModelScope Text-to-Video Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.096079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.096079Z digest=sha256:b53a5b31328ebe617a725d492a687c70cd586e44b7cc5f0e81f5158b489a25b8

Observation e663e69f-bf22-4cbb-9f2b-755e9db47ae8 · outbound

This paper cites Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content.

CoT-Edit: Let CoT Guide Instruction Video Editing Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.788680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.099979Z digest=sha256:1dd99ff85aef44268e1659128f2793bd5cf28ca3c40a0e52f9868ea4ca15b0ac

Observation 8fb1b249-ffac-4e93-a955-c98c5a7a2aa0 · outbound

This paper cites Tiv-diffusion: Towards object-centric movement for text-driven image to video gen- eration.

CoT-Edit: Let CoT Guide Instruction Video Editing Tiv-diffusion: Towards object-centric movement for text-driven image to video gen- eration

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.778051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.103092Z digest=sha256:111c464dd257244e6f301b746a7705af32e09147615d591c94af0502c6dee303

Observation f013a9ff-17b8-40ca-9547-f6c71f7c7fc6 · outbound

This paper cites Re-attentional con- trollable video diffusion editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Re-attentional con- trollable video diffusion editing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.765446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.106568Z digest=sha256:e63dddd0a305e871c40adfbd393320064e505855b5ef791d7177013cf42ba386

Observation 43379ae1-bdc5-4091-8111-ffe361a8a947 · outbound

This paper cites Training-free controllable text-guided video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2026.

CoT-Edit: Let CoT Guide Instruction Video Editing Training-free controllable text-guided video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2026

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.755803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.109921Z digest=sha256:06ce7049c4023f7fa79de7e3a670406fc47dbdda20799f4c04ec5338c15ee025

Observation d7a7aec6-83e1-4638-adb3-63faecad4f77 · outbound

This paper cites Phrasecut: Language-based image segmen- tation in the wild.

CoT-Edit: Let CoT Guide Instruction Video Editing Phrasecut: Language-based image segmen- tation in the wild

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.745577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.112795Z digest=sha256:e42907694596c8d1cfa9be3105fa3f453d7d18794c3a7bc2964ddb051762448e

Observation c685c6db-e0dd-4fca-b8ee-7b27984456af · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.733642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.115970Z digest=sha256:8366cb7a03dc02b56b89e393a10477b9647c7208dac17ca59f72e5cce4aba096

Observation 2f70e7ea-a7e9-4628-b819-5d9dc8fb09d1 · outbound

This paper cites Insvie-1m: Effective instruction-based video editing with elaborate dataset construction.

CoT-Edit: Let CoT Guide Instruction Video Editing Insvie-1m: Effective instruction-based video editing with elaborate dataset construction

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.721670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.119760Z digest=sha256:2e220f0e1e1d8d7c839f199a96c47f5bd8ed9014dcee8c03cce10fd151fcb316

Observation 67b4273c-b3c0-4edc-8f4b-aa0bce3e75e6 · outbound

This paper cites Veg- gie: Instructional editing and reasoning video concepts with grounded generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Veg- gie: Instructional editing and reasoning video concepts with grounded generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.122798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.122798Z digest=sha256:f20f80f5ca3423271023fbd7e32601e7023b7e32e5f489a342536d0e045c0c15

Observation 222f1b9e-0d29-495a-98a4-845715a0059e · outbound

This paper cites Editworld: Simulating world dynamics for instruction- following image editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Editworld: Simulating world dynamics for instruction- following image editing

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.704257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.125635Z digest=sha256:9565f51575eac2d963c29f79285e7b445d573390a6f60dbcf3f7bc6fad76d750

Observation 8ba8b5ea-f2e3-400f-bd20-5da9728d9eab · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.128527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.128527Z digest=sha256:76bda0ac10da2ddf6fab2e23dca644476c73bb42e55862c0976019e87da3560a

Observation 89c0ee69-c1c6-411c-8e1c-6e5af7d15788 · outbound

This paper cites EffiVED:Efficient Video Editing via Text-instruction Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing EffiVED:Efficient Video Editing via Text-instruction Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.132218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.132218Z digest=sha256:1efb771c0980e9e6a6f93d44e34606f8ca566a948980232da07eeb30987a8682

Observation 106458a7-e319-4faf-9c0d-7a77c611b3f4 · outbound

This paper cites Motionpro: A precise mo- tion controller for image-to-video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Motionpro: A precise mo- tion controller for image-to-video generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.693440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.135815Z digest=sha256:ef169acbc646ad7351728aa66561a4f2d5a5ee1a5c3d856e1fe660d28c7bb9fd

Observation 892ca003-1808-4665-9ca2-76f4c6b06c01 · outbound

This paper cites Ultraedit: Instruction-based fine-grained im- age editing at scale.Advances in Neural Information Pro- cessing Systems, 37:3058–3093, 2024.

CoT-Edit: Let CoT Guide Instruction Video Editing Ultraedit: Instruction-based fine-grained im- age editing at scale.Advances in Neural Information Pro- cessing Systems, 37:3058–3093, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.681969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.139111Z digest=sha256:caaba0854f99bb37619c6299177ad3e7da3a64132c197b4e260f430453d54120

Observation 6937549a-7058-4ae0-a562-31ce580c7c73 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.International journal of computer vision, 127(3):302–321, 2019.

CoT-Edit: Let CoT Guide Instruction Video Editing Semantic under- standing of scenes through the ade20k dataset.International journal of computer vision, 127(3):302–321, 2019

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.670892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:14:36.142912Z digest=sha256:9a8724588a4a0e835359cfc70f4c4bb0a41e6dd5e190f31b51c7ae74199286fa

Observation f2bbfc9a-9c3a-44b6-b89a-7f1cfbfb6813 · outbound

This paper cites Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists.

CoT-Edit: Let CoT Guide Instruction Video Editing Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.146697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.146697Z digest=sha256:fb67fc291727317525f391b0082c7d05cb88c5c0cba36cbeb954989b94b8bd46

Pith citing papers

No inbound Pith citation observations are available.