Pith. sign in

REVIEW 2 cited by

Spectral Motion Alignment for Video Motion Transfer using Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.15249 v2 pith:VLLKTNZP submitted 2024-03-22 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords motionvideodiffusionmodelsalignmentcustomizationglobalspectral
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The evolution of diffusion models has greatly impacted video generation and understanding. Particularly, text-to-video diffusion models (VDMs) have significantly facilitated the customization of input video with target appearance, motion, etc. Despite these advances, challenges persist in accurately distilling motion information from video frames. While existing works leverage the consecutive frame residual as the target motion vector, they inherently lack global motion context and are vulnerable to frame-wise distortions. To address this, we present Spectral Motion Alignment (SMA), a novel framework that refines and aligns motion vectors using Fourier and wavelet transforms. SMA learns motion patterns by incorporating frequency-domain regularization, facilitating the learning of whole-frame global motion dynamics, and mitigating spatial artifacts. Extensive experiments demonstrate SMA's efficacy in improving motion transfer while maintaining computational efficiency and compatibility across various video customization frameworks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors

    cs.CV 2025-01 conditional novelty 7.0 of 10

    Interactive image editing can be cast as image-to-video generation: initializing from Stable Video Diffusion plus a new matching attention mechanism yields high-quality sketch, drag, and coarse-edit results with far l...

  2. Video Creation by Demonstration

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A self-supervised diffusion approach, δ-Diffusion, transfers action concepts from a demonstration video to a new context image using appearance-bottlenecked action latents.

Pith tools