A hybrid two-stage framework pairs a discriminative front-end for interference suppression with a generative decoder-only LM back-end to improve perceptual quality and speaker consistency in target speaker extraction and speech enhancement.
An efficient encoder-decoder architec- ture with top-down attention for speech separation
3 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
A FiLM-conditioned transformer masker on DAC codec latents performs text-guided sound separation with claimed efficiency, but the main comparison against AudioSep is confounded by asymmetric input processing.
TF-MoE uses dynamic per-frame and per-mel-band expert selection in time and frequency dimensions to improve speech separation performance at comparable compute cost to prior models.
citing papers explorer
-
Discriminative-Generative Target Speaker Extraction with Decoder-Only Language Models
A hybrid two-stage framework pairs a discriminative front-end for interference suppression with a generative decoder-only LM back-end to improve perceptual quality and speaker consistency in target speaker extraction and speech enhancement.
-
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
A FiLM-conditioned transformer masker on DAC codec latents performs text-guided sound separation with claimed efficiency, but the main comparison against AudioSep is confounded by asymmetric input processing.
-
TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation
TF-MoE uses dynamic per-frame and per-mel-band expert selection in time and frequency dimensions to improve speech separation performance at comparable compute cost to prior models.