A U-Net conditioned on target direction and a text embedding extracts the target ambisonic sound field from mixtures, outperforming beamforming baselines and showing the largest semantic benefit when a secondary source is near the target.
Problem Setup Let a(t, θ, ϕ) be the plane wave amplitude distribution describ- ing an incident sound field
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
A U-Net conditioned on target direction and a text embedding extracts the target ambisonic sound field from mixtures, outperforming beamforming baselines and showing the largest semantic benefit when a secondary source is near the target.