Pith. sign in

REVIEW 3 cited by

MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.11629 v1 pith:DX435LSJ submitted 2024-07-16 eess.AS

classification eess.AS
keywords speakeranonymizationdisentanglementidentitymulti-lingualmusaserialstrategy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speaker anonymization is an effective privacy protection solution designed to conceal the speaker's identity while preserving the linguistic content and para-linguistic information of the original speech. While most prior studies focus solely on a single language, an ideal speaker anonymization system should be capable of handling multiple languages. This paper proposes MUSA, a Multi-lingual Speaker Anonymization approach that employs a serial disentanglement strategy to perform a step-by-step disentanglement from a global time-invariant representation to a temporal time-variant representation. By utilizing semantic distillation and self-supervised speaker distillation, the serial disentanglement strategy can avoid strong inductive biases and exhibit superior generalization performance across different languages. Meanwhile, we propose a straightforward anonymization strategy that employs empty embedding with zero values to simulate the speaker identity concealment process, eliminating the need for conversion to a pseudo-speaker identity and thereby reducing the complexity of speaker anonymization process. Experimental results on VoicePrivacy official datasets and multi-lingual datasets demonstrate that MUSA can effectively protect speaker privacy while preserving linguistic content and para-linguistic information.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SegReConcat: A Data Augmentation Method for Voice Anonymization Attack

    cs.SD 2025-08 conditional novelty 6.0 of 10

    SegReConcat, a word-shuffle-and-concatenate augmentation, improves attacker speaker verification against five of seven voice anonymization systems in the VPAC 2024 benchmark.

  2. Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A context-dependent encoding of phoneme durations identifies speakers far better than average-duration vectors and remains effective on anonymized speech without retraining on anonymized data.

  3. Mitigating Language Mismatch in SSL-Based Speaker Anonymization

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Fine-tuning an SSL content encoder on Japanese, especially when the encoder is pre-trained multilingually, makes anonymized Japanese and Mandarin speech much more intelligible while keeping speaker privacy at usable levels.

Pith tools