Pith. sign in

Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Text-driven video editing utilizing generative diffusion models has garnered significant attention due to their potential applications. However, existing approaches are constrained by the limited word embeddings provided in pre-training, which hinders nuanced editing targeting open concepts with specific attributes. Directly altering the keywords in target prompts often results in unintended disruptions to the attention mechanisms. To achieve more flexible editing easily, this work proposes an improved concept-augmented video editing approach that generates diverse and stable target videos flexibly by devising abstract conceptual pairs. Specifically, the framework involves concept-augmented textual inversion and a dual prior supervision mechanism. The former enables plug-and-play guidance of stable diffusion for video editing, effectively capturing target attributes for more stylized results. The dual prior supervision mechanism significantly enhances video stability and fidelity. Comprehensive evaluations demonstrate that our approach generates more stable and lifelike videos, outperforming state-of-the-art methods.

fields

cs.CV 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production cs.CV · 2024-12-18 · conditional · none · ref 13 · internal anchor

    Sign-IDD generates sign language poses by disentangling joint coordinates into bone direction and length attributes inside a gloss-conditioned diffusion model, reporting state-of-the-art scores on PHOENIX14T and USTC-CSL.