Pith. sign in

REVIEW 4 cited by

DreamMotion: Space-Time Self-Similar Score Distillation for Zero-Shot Video Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12002 v2 pith:HEZMMESE submitted 2024-03-18 cs.CV cs.AI

classification cs.CVcs.AI
keywords videodistillationscoreeditingmotionapproachdiffusionoriginal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-driven diffusion-based video editing presents a unique challenge not encountered in image editing literature: establishing real-world motion. Unlike existing video editing approaches, here we focus on score distillation sampling to circumvent the standard reverse diffusion process and initiate optimization from videos that already exhibit natural motion. Our analysis reveals that while video score distillation can effectively introduce new content indicated by target text, it can also cause significant structure and motion deviation. To counteract this, we propose to match space-time self-similarities of the original video and the edited video during the score distillation. Thanks to the use of score distillation, our approach is model-agnostic, which can be applied for both cascaded and non-cascaded video diffusion frameworks. Through extensive comparisons with leading methods, our approach demonstrates its superiority in altering appearances while accurately preserving the original structure and motion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  2. Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.

  3. TransPixeler: Advancing Text-to-Video Generation with Transparency

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A LoRA-based adaptation of DiT video generators that jointly outputs aligned RGB and alpha channels via extra tokens, shared positions, and attention masking.

  4. SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SnapGen-V prunes, searches, and adversarially distills a video diffusion model down to 0.6B parameters that generates a five-second, 512x512 video on an iPhone 16 Pro Max in under five seconds.

Pith tools