A U-Net conditioned on target direction and a text embedding extracts the target ambisonic sound field from mixtures, outperforming beamforming baselines and showing the largest semantic benefit when a secondary source is near the target.
SoundSculpt consis- tently outperformed traditional signal processing baselines
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
baseline 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
baseline 1polarities
baseline 1representative citing papers
citing papers explorer
-
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
A U-Net conditioned on target direction and a text embedding extracts the target ambisonic sound field from mixtures, outperforming beamforming baselines and showing the largest semantic benefit when a secondary source is near the target.