A jointly trained audio-video VAE with segment-level contrastive alignment and semantic distillation improves downstream text-to-audio-video synchronization and quality.
Out of time: Automated lip sync in the wild
1 Pith paper cite this work, alongside 603 external citations. Polarity classification is still indexing.
1
Pith paper citing it
603
external citations · OpenAlex
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
A jointly trained audio-video VAE with segment-level contrastive alignment and semantic distillation improves downstream text-to-audio-video synchronization and quality.