Lightweight networks trained only on autoencoder latent codes can do bandwidth extension and mono-to-stereo upmixing at a fraction of the FLOPS of raw-audio models, but match those models only when the baselines are also degraded by the same autoencoder.
Stable audio open,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning to Upsample and Upmix Audio in the Latent Domain
Lightweight networks trained only on autoencoder latent codes can do bandwidth extension and mono-to-stereo upmixing at a fraction of the FLOPS of raw-audio models, but match those models only when the baselines are also degraded by the same autoencoder.