A learned audio codec represents 2.98 seconds of audio as 32 sparse events, each decoded as a noise burst convolved with decaying resonances and a room impulse response.
Pytorch: An imperative style, high- performance deep learning library,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Toward a Sparse and Interpretable Audio Codec
A learned audio codec represents 2.98 seconds of audio as 32 sparse events, each decoded as a noise burst convolved with decaying resonances and a room impulse response.