KeySync applies a keyframe-interpolated diffusion model with a lower-face mask to generate 512x512 lip-synced video with reduced expression leakage and SAM2-based occlusion handling.
Speech driven video editing via an audio-conditioned diffusion model
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
KeySync applies a keyframe-interpolated diffusion model with a lower-face mask to generate 512x512 lip-synced video with reduced expression leakage and SAM2-based occlusion handling.