Pith. sign in

REVIEW 2 cited by

HLTCOE JHU Submission to the Voice Privacy Challenge 2024

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.08913 v2 pith:EAUDONAR submitted 2024-09-13 eess.AS cs.LG

classification eess.AScs.LG
keywords systemsvoiceconversionbetterchallengeincludingmethodprivacy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-to-speech (TTS) based systems including Whisper-VITS. We found that while voice conversion systems better preserve emotional content, they struggle to conceal speaker identity in semi-white-box attack scenarios; conversely, TTS methods perform better at anonymization and worse at emotion preservation. Finally, we propose a random admixture system which seeks to balance out the strengths and weaknesses of the two category of systems, achieving a strong EER of over 40% while maintaining UAR at a respectable 47%.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SegReConcat: A Data Augmentation Method for Voice Anonymization Attack

    cs.SD 2025-08 conditional novelty 6.0 of 10

    SegReConcat, a word-shuffle-and-concatenate augmentation, improves attacker speaker verification against five of seven voice anonymization systems in the VPAC 2024 benchmark.

  2. Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A context-dependent encoding of phoneme durations identifies speakers far better than average-duration vectors and remains effective on anonymized speech without retraining on anonymized data.

Pith tools