Inference energy of seven text-to-audio diffusion models grows linearly with denoising steps, while quality saturates, so the best quality-per-energy settings use 10 to 50 steps.
Dong, Generative AI for Music and Audio
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
Inference energy of seven text-to-audio diffusion models grows linearly with denoising steps, while quality saturates, so the best quality-per-energy settings use 10 to 50 steps.