BirdDiff combines a multi-band enhancement stage with a DiffWave-based diffusion generator and multimodal conditioning, reporting substantially better bird-call synthesis metrics than DiffWave on a 12-species proprietary dataset.
Vision Transformer Segmentation for Visual Bird Sound Denoising
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Audio denoising, especially in the context of bird sounds, remains a challenging task due to persistent residual noise. Traditional and deep learning methods often struggle with artificial or low-frequency noise. In this work, we propose ViTVS, a novel approach that leverages the power of the vision transformer (ViT) architecture. ViTVS adeptly combines segmentation techniques to disentangle clean audio from complex signal mixtures. Our key contributions encompass the development of ViTVS, introducing comprehensive, long-range, and multi-scale representations. These contributions directly tackle the limitations inherent in conventional approaches. Extensive experiments demonstrate that ViTVS outperforms state-of-the-art methods, positioning it as a benchmark solution for real-world bird sound denoising applications. Source code is available at: https://github.com/aiai-4/ViVTS.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Towards High-Fidelity and Controllable Bioacoustic Generation via Enhanced Diffusion Learning
BirdDiff combines a multi-band enhancement stage with a DiffWave-based diffusion generator and multimodal conditioning, reporting substantially better bird-call synthesis metrics than DiffWave on a 12-species proprietary dataset.