Pith. sign in

FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Audio super-resolution is challenging owing to its ill-posed nature. Recently, the application of diffusion models in audio super-resolution has shown promising results in alleviating this challenge. However, diffusion-based models have limitations, primarily the necessity for numerous sampling steps, which causes significantly increased latency when synthesizing high-quality audio samples. In this paper, we propose FLowHigh, a novel approach that integrates flow matching, a highly efficient generative model, into audio super-resolution. We also explore probability paths specially tailored for audio super-resolution, which effectively capture high-resolution audio distributions, thereby enhancing reconstruction quality. The proposed method generates high-fidelity, high-resolution audio through a single-step sampling process across various input sampling rates. The experimental results on the VCTK benchmark dataset demonstrate that FLowHigh achieves state-of-the-art performance in audio super-resolution, as evaluated by log-spectral distance and ViSQOL while maintaining computational efficiency with only a single-step sampling process.

fields

cs.SD 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

A2SB: Audio-to-Audio Schrodinger Bridges

cs.SD · 2025-01-20 · conditional · novelty 6.0

A2SB applies Schrödinger bridges to music restoration, achieving state-of-the-art bandwidth extension and inpainting at 44.1kHz in a single vocoder-free model.

citing papers explorer

Showing 1 of 1 citing paper.

  • A2SB: Audio-to-Audio Schrodinger Bridges cs.SD · 2025-01-20 · conditional · none · ref 79 · internal anchor

    A2SB applies Schrödinger bridges to music restoration, achieving state-of-the-art bandwidth extension and inpainting at 44.1kHz in a single vocoder-free model.