Pith. sign in

REVIEW 1 cited by

From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.02560 v2 pith:3JPWFH4E submitted 2023-08-02 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords audioconditionedhigh-fidelitymodelsrepresentationsapproachbeendiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep generative models can generate high-fidelity audio conditioned on various types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients (MFCC)). Recently, such models have been used to synthesize audio waveforms conditioned on highly compressed representations. Although such methods produce impressive results, they are prone to generate audible artifacts when the conditioning is flawed or imperfect. An alternative modeling approach is to use diffusion models. However, these have mainly been used as speech vocoders (i.e., conditioned on mel-spectrograms) or generating relatively low sampling rate signals. In this work, we propose a high-fidelity multi-band diffusion-based framework that generates any type of audio modality (e.g., speech, music, environmental sounds) from low-bitrate discrete representations. At equal bit rate, the proposed approach outperforms state-of-the-art generative techniques in terms of perceptual quality. Training and, evaluation code, along with audio samples, are available on the facebookresearch/audiocraft Github page.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Compression of Higher Order Ambisonics with Multichannel RVQGAN

    cs.SD 2024-11 conditional novelty 5.0 of 10

    A multichannel RVQGAN with a covariance loss compresses 16-channel third-order Ambisonics to 16 kbps and outperforms Opus at 160 kbps in a MUSHRA listening test on ambient scenes.

Pith tools