Pith. sign in

REVIEW 2 cited by

NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.08545 v2 pith:CJKLWWKP submitted 2022-06-17 eess.AS cs.LG

classification eess.AScs.LG
keywords audionu-wavesamplingmodelratesinputsmodelsneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Conventionally, audio super-resolution models fixed the initial and the target sampling rates, which necessitate the model to be trained for each pair of sampling rates. We introduce NU-Wave 2, a diffusion model for neural audio upsampling that enables the generation of 48 kHz audio signals from inputs of various sampling rates with a single model. Based on the architecture of NU-Wave, NU-Wave 2 uses short-time Fourier convolution (STFC) to generate harmonics to resolve the main failure modes of NU-Wave, and incorporates bandwidth spectral feature transform (BSFT) to condition the bandwidths of inputs in the frequency domain. We experimentally demonstrate that NU-Wave 2 produces high-resolution audio regardless of the sampling rate of input while requiring fewer parameters than other models. The official code and the audio samples are available at https://mindslab-ai.github.io/nuwave2.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic Sampling via Learnable Observation Front Ends

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Learnable acoustic filterbanks plus constrained mixing and temporal readout produce more informative low-rate observations for speech reconstruction than fixed waveform sampling at the same budget.

  2. Inference-time Scaling for Diffusion-based Audio Super-resolution

    cs.SD 2025-08 conditional novelty 4.0 of 10

    Generating 120 candidate super-resolved audios and choosing the best by task-specific verifiers improves speech, music, and sound effects over single-sample diffusion output, at 120x compute.

Pith tools