Pith. sign in

REVIEW 1 cited by

Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07280 v1 pith:YN5FFANY submitted 2024-06-11 cs.SD eess.AS

classification cs.SDeess.AS
keywords trainingspeechqualityrecordingenvironmentlatentmodelsource
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose noise-robust voice conversion (VC) which takes into account the recording quality and environment of noisy source speech. Conventional denoising training improves the noise robustness of a VC model by learning noisy-to-clean VC process. However, the naturalness of the converted speech is limited when the noise of the source speech is unseen during the training. To this end, our proposed training conditions a VC model on two latent variables representing the recording quality and environment of the source speech. These latent variables are derived from deep neural networks pre-trained on recording quality assessment and acoustic scene classification and calculated in an utterance-wise or frame-wise manner. As a result, the trained VC model can explicitly learn information about speech degradation during the training. Objective and subjective evaluations show that our training improves the quality of the converted speech compared to the conventional training.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Vo-Ve, a 44-dimension vector of voice-attribute probabilities, offers attribute-level explanations for speaker similarity, but its discrimination accuracy and listener-above-chance validation are modest.

Pith tools