A sound-source-aware image-to-audio generator that detects objects, disambiguates their audio semantics in a learned cross-modal manifold, and mixes them to synthesize audio.
VGGSound : A large-scale audio-visual dataset
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Gotta Hear Them All: Towards Sound Source Aware Audio Generation
A sound-source-aware image-to-audio generator that detects objects, disambiguates their audio semantics in a learned cross-modal manifold, and mixes them to synthesize audio.