TWNM framework equips audio-language models with spatial scene analysis via FOA simulation and metadata-grounded training, reaching 70.8% accuracy on a new ASA benchmark.
Fsd50k: An open dataset of human-labeled sound events
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.SD 3representative citing papers
A FiLM-conditioned transformer masker on DAC codec latents performs text-guided sound separation with claimed efficiency, but the main comparison against AudioSep is confounded by asymmetric input processing.
Causal probing of attention in audio separation transformers identifies dual pathways and asynchronous convergence, enabling a training-free Layer-Selective Attention Caching method that reduces self-attention computation by ~25% with negligible quality loss.
citing papers explorer
-
The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models
TWNM framework equips audio-language models with spatial scene analysis via FOA simulation and metadata-grounded training, reaching 70.8% accuracy on a new ASA benchmark.
-
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
A FiLM-conditioned transformer masker on DAC codec latents performs text-guided sound separation with claimed efficiency, but the main comparison against AudioSep is confounded by asymmetric input processing.
-
Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models
Causal probing of attention in audio separation transformers identifies dual pathways and asynchronous convergence, enabling a training-free Layer-Selective Attention Caching method that reduces self-attention computation by ~25% with negligible quality loss.