Pith. sign in

REVIEW 2 cited by

QuickVC: Any-to-many Voice Conversion Using Inverse Short-time Fourier Transform for Faster Conversion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.08296 v4 pith:YYFJ6H3Y submitted 2023-02-16 cs.SD eess.AS

classification cs.SDeess.AS
keywords modelinformationconversioncompetitivecontentfourierinverseresults
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the development of automatic speech recognition (ASR) and text-to-speech (TTS) technology, high-quality voice conversion (VC) can be achieved by extracting source content information and target speaker information to reconstruct waveforms. However, current methods still require improvement in terms of inference speed. In this study, we propose a lightweight VITS-based VC model that uses the HuBERT-Soft model to extract content information features without speaker information. Through subjective and objective experiments on synthesized speech, the proposed model demonstrates competitive results in terms of naturalness and similarity. Importantly, unlike the original VITS model, we use the inverse short-time Fourier transform (iSTFT) to replace the most computationally expensive part. Experimental results show that our model can generate samples at over 5000 kHz on the 3090 GPU and over 250 kHz on the i9-10900K CPU, achieving competitive speed for the same hardware configuration.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment

    eess.AS 2025-07 conditional novelty 6.0 of 10

    SemAlignVC strips source-speaker timbre by aligning a speech semantic encoder to BERT text embeddings, then resynthesizes the content conditioned only on a target voice reference.

  2. Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Conan achieves chunkwise online zero-shot voice conversion, preserving source content while adopting the reference speaker's timbre and style, with a latency as low as 37 milliseconds.

Pith tools