Pith. sign in

REVIEW 11 cited by

Consistent Video-to-Video Transfer Using Synthetic Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.00213 v3 pith:BHXKI46B submitted 2023-11-01 cs.CV cs.AI

classification cs.CVcs.AI
keywords videovideo-to-videoeditingtransferapproachconsistentdatasetintroduce
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce a novel and efficient approach for text-based video-to-video editing that eliminates the need for resource-intensive per-video-per-model finetuning. At the core of our approach is a synthetic paired video dataset tailored for video-to-video transfer tasks. Inspired by Instruct Pix2Pix's image transfer via editing instruction, we adapt this paradigm to the video domain. Extending the Prompt-to-Prompt to videos, we efficiently generate paired samples, each with an input video and its edited counterpart. Alongside this, we introduce the Long Video Sampling Correction during sampling, ensuring consistent long videos across batches. Our method surpasses current methods like Tune-A-Video, heralding substantial progress in text-based video-to-video editing and suggesting exciting avenues for further exploration and deployment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

    cs.CV 2026-06 unverdicted novelty 6.5 of 10

    LiveEdit distills a bidirectional video foundation model into a unidirectional streaming editor via three-stage training plus mask caching to reach 12.66 FPS with stable edits.

  2. CoT-Edit: Let CoT Guide Instruction Video Editing

    cs.CV 2026-08 conditional novelty 6.0 of 10

    CoT-Edit achieves state-of-the-art instruction-based video editing by generating bounding boxes and enriched instructions with a CoT-enhanced multimodal planner, which guide mask-based diffusion editing.

  3. FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A single video-diffusion framework composites both static images and dynamic footage along user-defined trajectories by transporting canonical foreground latents directly into the background latent sequence.

  4. Under One Sun: Multi-Object Generative Perception of Materials and Illumination

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Factorizing video editing into semantic-token anchoring and motion-restoration pre-training produces strong zero-shot and SOTA open-source instruction-guided video edits without heavy external structural priors.

  5. IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation

    cs.CV 2025-06 reject novelty 6.0 of 10

    A diffusion video model that jointly uses HDR lighting, relit frames, and 3D point tracks to relight videos from text prompts.

  6. AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection

    cs.CV 2025-02 conditional novelty 6.0 of 10

    AdaFlow demonstrates a training-free method to edit more than 1,000 video frames in one inference on a single A800 GPU via adaptive attention token slimming and content-aware keyframe selection.

  7. RelightVid: Temporal-Consistent Diffusion Model for Video Relighting

    cs.CV 2025-01 conditional novelty 6.0 of 10

    RelightVid lifts IC-Light from image relighting to temporally consistent video relighting, trained on a new synthetic and in-the-wild paired dataset called LightAtlas.

  8. Generative Video Propagation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GenProp propagates first-frame edits through video with a single generative model, unifying removal, insertion, replacement, and tracking tasks.

  9. Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    UES adds a self-supervised video condition to text-to-video diffusion models, enabling them to edit videos from delta prompts without paired supervision.

  10. VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A large open-source hybrid image-video dataset and a LoRA-based diffusion baseline for interactive local video editing.

  11. Dense-Face: Personalized Face Generation Model via Dense Annotation Prediction

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Dense-Face is a personalized face generation model that adds a pose-controllable adapter and dense face annotation prediction to Stable Diffusion, improving identity preservation and text alignment.

Pith tools