Pith. sign in

REVIEW 3 cited by

HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.09493 v1 pith:OQKIOO3I submitted 2023-09-18 eess.AS cs.AIcs.SD

classification eess.AScs.AIcs.SD
keywords hiftnetistftnetachievingbigvganfastfilterfourierharmonic-plus-noise
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Recent advancements in speech synthesis have leveraged GAN-based networks like HiFi-GAN and BigVGAN to produce high-fidelity waveforms from mel-spectrograms. However, these networks are computationally expensive and parameter-heavy. iSTFTNet addresses these limitations by integrating inverse short-time Fourier transform (iSTFT) into the network, achieving both speed and parameter efficiency. In this paper, we introduce an extension to iSTFTNet, termed HiFTNet, which incorporates a harmonic-plus-noise source filter in the time-frequency domain that uses a sinusoidal source from the fundamental frequency (F0) inferred via a pre-trained F0 estimation network for fast inference speed. Subjective evaluations on LJSpeech show that our model significantly outperforms both iSTFTNet and HiFi-GAN, achieving ground-truth-level performance. HiFTNet also outperforms BigVGAN-base on LibriTTS for unseen speakers and achieves comparable performance to BigVGAN while being four times faster with only $1/6$ of the parameters. Our work sets a new benchmark for efficient, high-quality neural vocoding, paving the way for real-time applications that demand high quality speech synthesis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AnyBand: Unified Multi-Bandwidth Speech Extension via Frequency-Aware In-Context Spectral Infilling

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A single flow-matching model performs speech bandwidth extension across continuously varying cutoff frequencies by treating the observed low-band spectrum as an in-context prompt and infilling the masked high band.

  2. TokAN: Accent Normalization Using Self-Supervised Speech Tokens

    cs.SD 2026-07 accept novelty 6.0 of 10

    TokAN maps L2 speech tokens to L1-like tokens via an autoregressive converter plus GRPO rewards, cutting WER to 9.23% on seven English accents without natural parallel L1-L2 recordings.

  3. Controllable Accent Normalization via Discrete Diffusion

    eess.AS 2026-03 conditional novelty 6.0 of 10

    Masked discrete diffusion over SSL speech tokens plus a Common Token Predictor yields the lowest WER among compared accent-normalization systems and continuous accent-strength control via source-token reuse.

Pith tools