SwiftAudio performs caption-only distillation of a one-step TTA diffusion model by adapting VSD to audio with temporal smoothness regularization, achieving SOTA among one-step methods on AudioCaps and Clotho using ~45K captions.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4representative citing papers
A coarse-to-fine hybrid MMDiT/DiT audio editor trained with rectified flow matching improves fidelity and cuts edit time versus prior instruction-guided baselines on synthetic overlapping-event tasks.
Meta-ensemble learning on diverse ICBHI data splits reaches 66.49% Score and improves generalization on two external datasets.
Integrated gradients on a 10-class domestic sound classifier yields 0.39 mean IoU, 0.52 frame F1 and 82.6% Pointing Game accuracy for temporal event detection, approaching weakly and strongly supervised framewise CNN baselines.
citing papers explorer
-
SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation
SwiftAudio performs caption-only distillation of a one-step TTA diffusion model by adapting VSD to audio with temporal smoothness regularization, achieving SOTA among one-step methods on AudioCaps and Clotho using ~45K captions.
-
RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers
A coarse-to-fine hybrid MMDiT/DiT audio editor trained with rectified flow matching improves fidelity and cuts edit time versus prior instruction-guided baselines on synthetic overlapping-event tasks.
-
Meta-Ensemble Learning with Diverse Data Splits for Improved Respiratory Sound Classification
Meta-ensemble learning on diverse ICBHI data splits reaches 66.49% Score and improves generalization on two external datasets.
-
Evaluating the Temporal Detection Capability of Integrated Gradients Applied on Sound Classifier
Integrated gradients on a 10-class domestic sound classifier yields 0.39 mean IoU, 0.52 frame F1 and 82.6% Pointing Game accuracy for temporal event detection, approaching weakly and strongly supervised framewise CNN baselines.