Video-to-audio synthesis can be made incremental by subtracting an audio-conditioned prediction from the text-guided prediction, producing complementary layers without extra multi-reference training data.
Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
Video-to-audio synthesis can be made incremental by subtracting an audio-conditioned prediction from the text-guided prediction, producing complementary layers without extra multi-reference training data.