Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:54:50.730336Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 7 inbound Pith citation observations for arXiv:2411.15260.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:54:50.730336Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:34:02.014040Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
40 of 40 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 62149c36-1396-46fa-be3d-f351b3b3fd13 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d36eb92-57b5-49c6-9315-31c00ba25d08 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing In- structpix2pix: Learning to follow image editing instructions
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94542935-7e18-49e2-87c4-7e496b158163 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4fbbe647-dc53-47a4-bda1-cfe749bf49e9 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5836d2c1-1aaa-4e73-9409-f2bdb81c67f6 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Consistent Video-to-Video Transfer Using Synthetic Dataset
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 114506a5-c85e-43de-8afc-e30ae7627ec7 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 665a9d35-87f1-45b4-bb38-25ddbc5db4ef · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaef9768-ff54-4f77-a6c6-e45304547d5e · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Proxedit: Improv- ing tuning-free real image editing with proximal guidance
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d8e24bf8-e1f8-4eeb-a099-dc76917b1dcf · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 859634e8-52eb-4ac9-8481-ebef907adf3a · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Clipscore, 2023
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6cc3bebc-6611-4f07-8ca8-8ec953ad1cb2 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Denoising dif- fusion probabilistic models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a7ba036-dc30-479c-a36f-025025707ea4 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing LoRA: Low-Rank Adaptation of Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e548be8-9872-4d22-9b7f-16694643a7a1 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b2ca79a-840a-4c4a-94e3-483756f29acb · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d43c3d-d6f4-4320-b479-b48b66dc0b79 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Rave: Randomized noise shuf- fling for fast and consistent video editing with diffusion models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation be02da15-4048-4341-867c-eec419927692 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Imagic: Text-based real image editing with diffusion models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 77e2f4e7-c3fe-49d5-8388-fd53b95f9406 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Segment anything
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1930388-2a57-4864-a374-ec69ee444aa6 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing DiffuEraser: A Diffusion Model for Video Inpainting
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07cab2b5-7a2b-47fe-b7e4-adf775ce8fac · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing MagicEdit: High-Fidelity and Temporally Coherent Video Editing
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87d1ffc8-3f2c-416b-94a1-fe3b69f66a87 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 277c82f2-6777-4973-ad20-670ba86b135d · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Negative-prompt Inversion: Fast Image Inversion for Editing with Text-guided Diffusion Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b0720f-535d-4d51-bcd4-8e62691cc139 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Movie Gen: A Cast of Media Foundation Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6331c421-1df8-4d5f-a2c9-4369366c68d6 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Fatezero: Fus- ing attentions for zero-shot text-based video editing
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4a0e1232-93bf-4b81-bc11-d5574636f6b2 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Instructvid2vid: Controllable video editing with natural language instructions
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c6543ee4-77f0-44ee-8d17-c97c31d0b24f · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing SAM 2: Segment Anything in Images and Videos
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e89716b-b9fa-401c-bd11-7b73da3b9ebc · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing High-resolution image synthesis with latent diffusion models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ccaa253-d46b-4eec-8ccf-157a4d3c6e09 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Denoising Diffusion Implicit Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f504d10-287e-4123-94db-4b8537504bc4 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Score-Based Generative Modeling through Stochastic Differential Equations
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1351d430-0af4-4b62-9aca-6522e1386352 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Raft: Recurrent all-pairs field transforms for optical flow
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 72b25aaa-3c4b-4035-9a44-30702ee6e8c5 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Motioneditor: Editing video motion via content-aware diffusion
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 963f2830-2b84-40f2-b85e-61a66c7e24a5 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Videocomposer: Compositional video synthesis with motion controllability
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d3c40406-79ea-4103-b41e-7e7646dbdd5e · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bd7bafb-16a8-4de3-9a7e-4334f90c069f · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e130e7-c401-4687-b6eb-4adf3c0c5d97 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Magicbrush: A manually annotated dataset for instruction- guided image editing
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e7624576-2d55-4e12-84c9-d2c569ddc349 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Recognize anything: A strong image tagging model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 72b657ac-a1c4-4731-a1a8-ff593f1bf62a · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Avid: Any-length video inpainting with dif- fusion model
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 257e56d3-1cdc-444c-8a1a-3fa838cdcddb · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d4b6c7-0f9b-4a58-b784-a52d5cbf4392 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing Propainter: Improving propagation and transformer for video inpainting
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2468b0f2-92b6-4eed-b9d9-c6d7a0e3865e · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eedd5048-4ad9-480b-b520-247437be7ac5 · outbound
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ceb27b-0a23-4396-b94a-19cb1c46bd9f · inbound
Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a3962f7-d4c2-454d-ae84-c82e39445a20 · inbound
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 746b9c8e-e72b-4789-9a36-0f9d428cafa3 · inbound
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1107d8e1-f21d-4457-9c6d-3a88dec2943f · inbound
From Ideal to Real: Stable Video Object Removal under Imperfect Conditions VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9afc401a-2586-47b4-b45c-87c8195dae2f · inbound
MiVE: Multiscale Vision-language features for reference-guided video Editing VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 61cc3b24-4d13-4bc2-8044-9cd2058a8b71 · inbound
Diffusing in the Right Space: A Systematic Study of Latent Diffusability VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation de744b06-fdfd-4b1f-a811-e9516906f7d2 · inbound
SteerVTE: Seamless Video Text Editing with Style and Glyph Control VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.