Pith. sign in

REVIEW 3 cited by

NPU-NTU System for Voice Privacy 2024 Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.04173 v2 pith:5MN3CYMS submitted 2024-09-06 eess.AS

classification eess.AS
keywords speakeridentitydistillationsystemanonymizationcontentemotionlinguistic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speaker anonymization is an effective privacy protection solution that conceals the speaker's identity while preserving the linguistic content and paralinguistic information of the original speech. To establish a fair benchmark and facilitate comparison of speaker anonymization systems, the VoicePrivacy Challenge (VPC) was held in 2020 and 2022, with a new edition planned for 2024. In this paper, we describe our proposed speaker anonymization system for VPC 2024. Our system employs a disentangled neural codec architecture and a serial disentanglement strategy to gradually disentangle the global speaker identity and time-variant linguistic content and paralinguistic information. We introduce multiple distillation methods to disentangle linguistic content, speaker identity, and emotion. These methods include semantic distillation, supervised speaker distillation, and frame-level emotion distillation. Based on these distillations, we anonymize the original speaker identity using a weighted sum of a set of candidate speaker identities and a randomly generated speaker identity. Our system achieves the best trade-off of privacy protection and emotion preservation in VPC 2024.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SegReConcat: A Data Augmentation Method for Voice Anonymization Attack

    cs.SD 2025-08 conditional novelty 6.0 of 10

    SegReConcat, a word-shuffle-and-concatenate augmentation, improves attacker speaker verification against five of seven voice anonymization systems in the VPAC 2024 benchmark.

  2. Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A context-dependent encoding of phoneme durations identifies speakers far better than average-duration vectors and remains effective on anonymized speech without retraining on anonymized data.

  3. EASY: Emotion-aware Speaker Anonymization via Factorized Distillation

    eess.AS 2025-05 conditional novelty 5.0 of 10

    EASY separates speaker identity, linguistic content, and emotion through sequential factorized distillation, and reports better privacy and emotion preservation than prior VoicePrivacy 2024 systems.

Pith tools