A U-Net conditioned on target direction and a text embedding extracts the target ambisonic sound field from mixtures, outperforming beamforming baselines and showing the largest semantic benefit when a secondary source is near the target.
We present SI-SDRi (averaged across channels) for the baseline algorithms and SoundSculpt models that use various types of conditioning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
A U-Net conditioned on target direction and a text embedding extracts the target ambisonic sound field from mixtures, outperforming beamforming baselines and showing the largest semantic benefit when a secondary source is near the target.