Pith. sign in

REVIEW 1 cited by

Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.11528 v1 pith:772HQGIZ submitted 2024-08-21 cs.SD eess.AS

classification cs.SDeess.AS
keywords conversionspeakervoicespeechzero-shotwhispereddomainsgenerated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Zero-shot voice conversion aims to transfer the voice of a source speaker to that of a speaker unseen during training, while preserving the content information. Although various methods have been proposed to reconstruct speaker information in generated speech, there is still room for improvement in achieving high similarity between generated and ground truth recordings. Furthermore, zero-shot voice conversion for speech in specific domains, such as whispered, remains an unexplored area. To address this problem, we propose a SpeakerVC model that can effectively perform zero-shot speech conversion in both voiced and whispered domains, while being lightweight and capable of running in streaming mode without significant quality degradation. In addition, we explore methods to improve the quality of speaker identity transfer and demonstrate their effectiveness for a variety of voice conversion systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion

    eess.AS 2025-11 conditional novelty 5.0 of 10

    WhisperVC converts whispered Mandarin to natural speech by decoupling domain alignment from synthesis, reporting DNSMOS 3.11, UTMOS 2.52, CER 18.67%, and speaker cosine 0.76.

Pith tools