Pith. sign in

REVIEW 2 cited by

WSRGlow: A Glow-based Waveform Generative Model for Audio Super-Resolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.08507 v1 pith:TUJCQPP6 submitted 2021-06-16 cs.SD eess.AS

classification cs.SDeess.AS
keywords audiomodelgenerativeinformationsuper-resolutionwsrglowdomainencoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio super-resolution is the task of constructing a high-resolution (HR) audio from a low-resolution (LR) audio by adding the missing band. Previous methods based on convolutional neural networks and mean squared error training objective have relatively low performance, while adversarial generative models are difficult to train and tune. Recently, normalizing flow has attracted a lot of attention for its high performance, simple training and fast inference. In this paper, we propose WSRGlow, a Glow-based waveform generative model to perform audio super-resolution. Specifically, 1) we integrate WaveNet and Glow to directly maximize the exact likelihood of the target HR audio conditioned on LR information; and 2) to exploit the audio information from low-resolution audio, we propose an LR audio encoder and an STFT encoder, which encode the LR information from the time domain and frequency domain respectively. The experimental results show that the proposed model is easier to train and outperforms the previous works in terms of both objective and perceptual quality. WSRGlow is also the first model to produce 48kHz waveforms from 12kHz LR audio.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution

    cs.SD 2025-01 conditional novelty 6.0 of 10

    HiFi-SR is a single end-to-end transformer-convolutional GAN that upscales speech from 4 to 32 kHz inputs to 48 kHz with slightly better spectral distance and listener preference than existing two-stage systems.

  2. Bridge-SR: Schr\"odinger Bridge for Efficient SR

    cs.SD 2025-01 conditional novelty 6.0 of 10

    Bridge-SR applies tractable Schrödinger bridge models to waveform-domain speech super-resolution, and with 1.7M parameters reports the lowest log-spectral distance on VCTK while matching diffusion quality at 4 sampling steps.

Pith tools