Video-to-audio synthesis can be made incremental by subtracting an audio-conditioned prediction from the text-guided prediction, producing complementary layers without extra multi-reference training data.
Action2sound: Ambient-aware generation of action sounds from egocentric videos
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
Video-to-audio synthesis can be made incremental by subtracting an audio-conditioned prediction from the text-guided prediction, producing complementary layers without extra multi-reference training data.