Pith. sign in

SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

This paper introduces SoundSculpt, a neural network designed to extract target sound fields from ambisonic recordings. SoundSculpt employs an ambisonic-in-ambisonic-out architecture and is conditioned on both spatial information (e.g., target direction obtained by pointing at an immersive video) and semantic embeddings (e.g., derived from image segmentation and captioning). Trained and evaluated on synthetic and real ambisonic mixtures, SoundSculpt demonstrates superior performance compared to various signal processing baselines. Our results further reveal that while spatial conditioning alone can be effective, the combination of spatial and semantic information is beneficial in scenarios where there are secondary sound sources spatially close to the target. Additionally, we compare two different semantic embeddings derived from a text description of the target sound using text encoders.

citation-role summary

baseline 1

citation-polarity summary

fields

eess.AS 1

years

2025 1

verdicts

CONDITIONAL 1

roles

baseline 1

polarities

baseline 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction eess.AS · 2025-05-30 · conditional · none · ref 6 · internal anchor

    A U-Net conditioned on target direction and a text embedding extracts the target ambisonic sound field from mixtures, outperforming beamforming baselines and showing the largest semantic benefit when a secondary source is near the target.