UAF is the first unified audio front-end LLM that turns multiple front-end tasks into one sequence prediction model processing streaming audio chunks and reference prompts to output semantic and control tokens for full-duplex interaction.
Dccrn: Deep complex convolution recurrent network for phase-aware speech enhancement.arXiv preprint arXiv:2008.00264
4 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
A FiLM-conditioned transformer masker on DAC codec latents performs text-guided sound separation with claimed efficiency, but the main comparison against AudioSep is confounded by asymmetric input processing.
BioSEN enhances animal vocalizations with multi-scale attention, harmonic modeling, and energy-adaptive gating, matching speech models at lower compute on three bioacoustic datasets.
citing papers explorer
-
UAF: A Unified Audio Front-end LLM for Full-Duplex Speech Interaction
UAF is the first unified audio front-end LLM that turns multiple front-end tasks into one sequence prediction model processing streaming audio chunks and reference prompts to output semantic and control tokens for full-duplex interaction.
-
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
A FiLM-conditioned transformer masker on DAC codec latents performs text-guided sound separation with claimed efficiency, but the main comparison against AudioSep is confounded by asymmetric input processing.
-
BioSEN: A Bio-acoustic Signal Enhancement Network for Animal Vocalizations
BioSEN enhances animal vocalizations with multi-scale attention, harmonic modeling, and energy-adaptive gating, matching speech models at lower compute on three bioacoustic datasets.
- A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models