SwiftAudio performs caption-only distillation of a one-step TTA diffusion model by adapting VSD to audio with temporal smoothness regularization, achieving SOTA among one-step methods on AudioCaps and Clotho using ~45K captions.
Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.SD 2years
2026 2representative citing papers
A coarse-to-fine hybrid MMDiT/DiT audio editor trained with rectified flow matching improves fidelity and cuts edit time versus prior instruction-guided baselines on synthetic overlapping-event tasks.
citing papers explorer
-
SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation
SwiftAudio performs caption-only distillation of a one-step TTA diffusion model by adapting VSD to audio with temporal smoothness regularization, achieving SOTA among one-step methods on AudioCaps and Clotho using ~45K captions.
-
RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers
A coarse-to-fine hybrid MMDiT/DiT audio editor trained with rectified flow matching improves fidelity and cuts edit time versus prior instruction-guided baselines on synthetic overlapping-event tasks.