Pith. sign in

Paper Citation Record · LEDGER

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

As of 23 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 7 inbound Pith citation observations for arXiv:2411.15260.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15260 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:54:50.730336Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:34:02.014040Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 62149c36-1396-46fa-be3d-f351b3b3fd13 · outbound

This paper cites UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.532030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.532030Z digest=sha256:fcd0c36df2777d891f128d031b864ac5682898611bbeefb320d31ba686998452

Observation 4d36eb92-57b5-49c6-9315-31c00ba25d08 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing In- structpix2pix: Learning to follow image editing instructions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.538129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.538129Z digest=sha256:120c678fa5b6aa30d55c9e09627b16c967739499c91272774c8f6e0ae8092498

Observation 94542935-7e18-49e2-87c4-7e496b158163 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.501693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.543042Z digest=sha256:e8efb7c1b74dfdbdea0225465a8a34146c0ba226ecd7f692cf5d611399e09567

Observation 4fbbe647-dc53-47a4-bda1-cfe749bf49e9 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.548353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.548353Z digest=sha256:c6a1d3b47c36cf01b04f2b4c16889bf501fc5020cb768cb99c8868ec7cd883fc

Observation 5836d2c1-1aaa-4e73-9409-f2bdb81c67f6 · outbound

This paper cites Consistent Video-to-Video Transfer Using Synthetic Dataset.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Consistent Video-to-Video Transfer Using Synthetic Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.553671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.553671Z digest=sha256:53b9367f2c89ec8ca9647f9b896037da92403721113187bbd601ab77db9c12cb

Observation 114506a5-c85e-43de-8afc-e30ae7627ec7 · outbound

This paper cites FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.558767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.558767Z digest=sha256:f532947e08285dcd50d646527e10c2b87d45a8d137b979ea6239580e01c0f043

Observation 665a9d35-87f1-45b4-bb38-25ddbc5db4ef · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.563810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.563810Z digest=sha256:b471472d6d31a98b6ec80ce5757ac76946c91700394b74024e42a4e6b7c7bd7e

Observation aaef9768-ff54-4f77-a6c6-e45304547d5e · outbound

This paper cites Proxedit: Improv- ing tuning-free real image editing with proximal guidance.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Proxedit: Improv- ing tuning-free real image editing with proximal guidance

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.484947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.569181Z digest=sha256:2f4b1c46c5194b7c677dbcd0b5be37515e3d773e30cfa9c957ac78f2fcbdf7a7

Observation d8e24bf8-e1f8-4eeb-a099-dc76917b1dcf · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.573649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.573649Z digest=sha256:29775a7c2f969f3fc69d4ed46c7f99f477ab42689231853d5a896e4eeb9339d7

Observation 859634e8-52eb-4ac9-8481-ebef907adf3a · outbound

This paper cites Clipscore, 2023.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Clipscore, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.468120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.578615Z digest=sha256:2d87cc73af31e55bd4ce5ea4fb9ddf795b63ab272cc29c810d9a47f1e1872eb3

Observation 6cc3bebc-6611-4f07-8ca8-8ec953ad1cb2 · outbound

This paper cites Denoising dif- fusion probabilistic models.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Denoising dif- fusion probabilistic models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.584509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.584509Z digest=sha256:2fa4e6a5dae04c7d939e26a9f2b2f97a4b1f7fb086538f3cde6f0180ad0cb08d

Observation 5a7ba036-dc30-479c-a36f-025025707ea4 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing LoRA: Low-Rank Adaptation of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.589452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.589452Z digest=sha256:29ea8af8b3f4d43f17266b093c7650909367d3fe61d11927e1bd146677e946bd

Observation 0e548be8-9872-4d22-9b7f-16694643a7a1 · outbound

This paper cites HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.594607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.594607Z digest=sha256:6147d032995366049bd13ea9d239a10778db3caeece798fccbbab53ccb0a2d5d

Observation 9b2ca79a-840a-4c4a-94e3-483756f29acb · outbound

This paper cites BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.599316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.599316Z digest=sha256:54656a1b142bfe78e598630c1c8683738999dd5563919e839b2e28259f2176cf

Observation 00d43c3d-d6f4-4320-b479-b48b66dc0b79 · outbound

This paper cites Rave: Randomized noise shuf- fling for fast and consistent video editing with diffusion models.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Rave: Randomized noise shuf- fling for fast and consistent video editing with diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.440238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.604303Z digest=sha256:7216b1529bb59341a1d1c87bd7e3aef70641265fb48c7cd24e5d92b8a1c5ac1a

Observation be02da15-4048-4341-867c-eec419927692 · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Imagic: Text-based real image editing with diffusion models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.423352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.608817Z digest=sha256:196a0158eb372f80487992256666dcecf80745e89c62c4103942afe4f11ed473

Observation 77e2f4e7-c3fe-49d5-8388-fd53b95f9406 · outbound

This paper cites Segment anything.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Segment anything

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.613537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.613537Z digest=sha256:0d75789354ec3092cf410a854a7f0b574864f3493a5f08c9fd6bbb52435aeeaa

Observation b1930388-2a57-4864-a374-ec69ee444aa6 · outbound

This paper cites DiffuEraser: A Diffusion Model for Video Inpainting.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing DiffuEraser: A Diffusion Model for Video Inpainting

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.618245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.618245Z digest=sha256:592528acd56778e1899f88d2cf51f1f17d6a07b08b1afd18cb9eb715614a9c74

Observation 07cab2b5-7a2b-47fe-b7e4-adf775ce8fac · outbound

This paper cites MagicEdit: High-Fidelity and Temporally Coherent Video Editing.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing MagicEdit: High-Fidelity and Temporally Coherent Video Editing

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.623129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.623129Z digest=sha256:50fe4391f11ff9ddd536cd74115e7afba9170c51e4217410dc8a336f963545f8

Observation 87d1ffc8-3f2c-416b-94a1-fe3b69f66a87 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.629211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.629211Z digest=sha256:54b052fc4d248e7ecec4936805f1941fef7a17ca577ac8325418ab4c01be8b54

Observation 277c82f2-6777-4973-ad20-670ba86b135d · outbound

This paper cites Negative-prompt Inversion: Fast Image Inversion for Editing with Text-guided Diffusion Models.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Negative-prompt Inversion: Fast Image Inversion for Editing with Text-guided Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.633722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.633722Z digest=sha256:020a1544d813dae871e7fcd5042dce903ecd643818eed7f5d4b880ba91b81bba

Observation 90b0720f-535d-4d51-bcd4-8e62691cc139 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Movie Gen: A Cast of Media Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.638492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.638492Z digest=sha256:0331f877c44a0037830a2e51f3d56d7aa588312aafd4d9c3b0d8e62964053342

Observation 6331c421-1df8-4d5f-a2c9-4369366c68d6 · outbound

This paper cites Fatezero: Fus- ing attentions for zero-shot text-based video editing.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Fatezero: Fus- ing attentions for zero-shot text-based video editing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.395430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.643220Z digest=sha256:eca19ab0e72b05c2f0d9983870fdf2b1f47de494273565d245d1eb0c61426814

Observation 4a0e1232-93bf-4b81-bc11-d5574636f6b2 · outbound

This paper cites Instructvid2vid: Controllable video editing with natural language instructions.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Instructvid2vid: Controllable video editing with natural language instructions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.378368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.647773Z digest=sha256:27984396466fc8c38439dfe2058f693b13c0c616c406f39cf5e2594b4cb89cbe

Observation c6543ee4-77f0-44ee-8d17-c97c31d0b24f · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing SAM 2: Segment Anything in Images and Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.652203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.652203Z digest=sha256:0fd7bcee895415c00b5a5da1f29b9f65c80c392195b112cd92db5afc1fbdccea

Observation 5e89716b-b9fa-401c-bd11-7b73da3b9ebc · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing High-resolution image synthesis with latent diffusion models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.657077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.657077Z digest=sha256:33f572bdfc6e571b74c851b137b57f7a0d9e9652793057c4deb1e4e0da309ec4

Observation 9ccaa253-d46b-4eec-8ccf-157a4d3c6e09 · outbound

This paper cites Denoising Diffusion Implicit Models.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Denoising Diffusion Implicit Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.666571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.666571Z digest=sha256:34101072bd9b35dcf940dfc3fdcae4d88e0a4824e1da714483dfd0af293dd1eb

Observation 6f504d10-287e-4123-94db-4b8537504bc4 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Score-Based Generative Modeling through Stochastic Differential Equations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.671098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.671098Z digest=sha256:8a1979c2c6bc4b660e0ba800229b8ca6901b38d00aadfca247f1d978bac1de1c

Observation 1351d430-0af4-4b62-9aca-6522e1386352 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Raft: Recurrent all-pairs field transforms for optical flow

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.349632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.676050Z digest=sha256:a47d8ab9295dfd1ee6934e671b7b9db03f84fab1bf21e9c0ac0cb37d9a3c5e0f

Observation 72b25aaa-3c4b-4035-9a44-30702ee6e8c5 · outbound

This paper cites Motioneditor: Editing video motion via content-aware diffusion.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Motioneditor: Editing video motion via content-aware diffusion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.332310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.680861Z digest=sha256:e34db054ab71586a27527a192ff1dfc7c7b021d2928764da0dbab7c85d585d8e

Observation 963f2830-2b84-40f2-b85e-61a66c7e24a5 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Videocomposer: Compositional video synthesis with motion controllability

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.311298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.685771Z digest=sha256:7e8b438f7ca767ba203a6b1814d22d9cf874a63f81964a3db0434acd54539cf5

Observation d3c40406-79ea-4103-b41e-7e7646dbdd5e · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.691096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.691096Z digest=sha256:24f4b7299b233f67afbe387b88ab03a8d7d5511e165bd51e9335335f603d5265

Observation 1bd7bafb-16a8-4de3-9a7e-4334f90c069f · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.696188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.696188Z digest=sha256:4b5f16db3d90e17d028add276cbf50a18f935043e3b85b81086fb6597ea310c6

Observation e7e130e7-c401-4687-b6eb-4adf3c0c5d97 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Magicbrush: A manually annotated dataset for instruction- guided image editing

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.280350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.701052Z digest=sha256:6f93445e6eb2cc9c5a8bbab1ae5a1e84763865800377d96549e32dfa33bf1958

Observation e7624576-2d55-4e12-84c9-d2c569ddc349 · outbound

This paper cites Recognize anything: A strong image tagging model.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Recognize anything: A strong image tagging model

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.262638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.705798Z digest=sha256:b60fd05bfc344ba903d44245f89a5a63f40ac9a0eec8d4af8cbf151d4444925b

Observation 72b657ac-a1c4-4731-a1a8-ff593f1bf62a · outbound

This paper cites Avid: Any-length video inpainting with dif- fusion model.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Avid: Any-length video inpainting with dif- fusion model

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.245604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.710513Z digest=sha256:056c373c328e844062e2d194357a94e16bc08554e4e6a34e7f76b512a977deb2

Observation 257e56d3-1cdc-444c-8a1a-3fa838cdcddb · outbound

This paper cites UltraEdit: Instruction-based Fine-Grained Image Editing at Scale.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing UltraEdit: Instruction-based Fine-Grained Image Editing at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.714962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.714962Z digest=sha256:5b537bc4d64f0e0a3ffff1487a85ea8388307ae5da29e03221f1ca157c2505f6

Observation b8d4b6c7-0f9b-4a58-b784-a52d5cbf4392 · outbound

This paper cites Propainter: Improving propagation and transformer for video inpainting.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Propainter: Improving propagation and transformer for video inpainting

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:54:51.228225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T14:54:50.719738Z digest=sha256:a174b0d932efd7e11ea7bfd69ea392ca857235dc115899fcb8f7a4fef3bb203e

Observation 2468b0f2-92b6-4eed-b9d9-c6d7a0e3865e · outbound

This paper cites A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.724863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.724863Z digest=sha256:17e81e1ae6e9ce7d612e8791fd2817a00f07bccbf9342172282956fa658ee507

Observation eedd5048-4ad9-480b-b520-247437be7ac5 · outbound

This paper cites CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility.

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:54:50.730336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:54:50.730336Z digest=sha256:9278d4a2f33349f3b159afc076116cb8d55614f23de51b1e457094fc6bf8692f

Pith citing papers

Observation 56ceb27b-0a23-4396-b94a-19cb1c46bd9f · inbound

Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists cites this paper.

Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T14:34:02.014040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:34:02.014040Z digest=sha256:86fc9d1622e8c36cf01ae90b7de4d733497efe2d0d8a7676f54e82939f53f3c6

Observation 2a3962f7-d4c2-454d-ae84-c82e39445a20 · inbound

MiniMax-Remover: Taming Bad Noise Helps Video Object Removal cites this paper.

MiniMax-Remover: Taming Bad Noise Helps Video Object Removal VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:21:38.845137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:21:38.845137Z digest=sha256:0e6c30f7964508673d9f1c8fc7a1c220be58a92bc56b107fb59bf128b74444f8

Observation 746b9c8e-e72b-4789-9a36-0f9d428cafa3 · inbound

EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning cites this paper.

EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:46:25.767860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T13:46:05.547669Z digest=sha256:b3946ca9b2c745f45d5d94a1d3a430f6c057fd115d6e83eccb7e3bf423b3cac8

Observation 1107d8e1-f21d-4457-9c6d-3a88dec2943f · inbound

From Ideal to Real: Stable Video Object Removal under Imperfect Conditions cites this paper.

From Ideal to Real: Stable Video Object Removal under Imperfect Conditions VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T13:35:51.545452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T13:34:45.782304Z digest=sha256:9c849ea07f3e1f4f63c4ca66707f5ca2e5a7a83384e42f72065472d2e4500fad

Observation 9afc401a-2586-47b4-b45c-87c8195dae2f · inbound

MiVE: Multiscale Vision-language features for reference-guided video Editing cites this paper.

MiVE: Multiscale Vision-language features for reference-guided video Editing VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:25:03.032960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T05:23:20.179387Z digest=sha256:d7a0f1dfd700551ab6d1eed59775b7e0a97032322742baa9a797d222680c99da

Observation 61cc3b24-4d13-4bc2-8044-9cd2058a8b71 · inbound

Diffusing in the Right Space: A Systematic Study of Latent Diffusability cites this paper.

Diffusing in the Right Space: A Systematic Study of Latent Diffusability VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:46:27.672332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T10:44:24.318786Z digest=sha256:ead9c7467ff6eb0acf2c0ceb586a91ed360b647195da4ebee0cb5b329cc6a16f

Observation de744b06-fdfd-4b1f-a811-e9516906f7d2 · inbound

SteerVTE: Seamless Video Text Editing with Style and Glyph Control cites this paper.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.520621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:d92bf810ba436a1e492169ace0ebe17a58a1152dfe066fedaa362541cdb93969