Pith. sign in

REVIEW 1 cited by

A Speech Representation Anonymization Framework via Selective Noise Perturbation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.14171 v2 pith:HM2FACF4 submitted 2022-03-26 eess.AS cs.CRcs.SD

classification eess.AScs.CRcs.SD
keywords speechanonymizationframeworkprivacyperturbationapproachautomaticnoise
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Privacy and security are major concerns when communicating speech signals to cloud services such as automatic speech recognition (ASR) and speech emotion recognition (SER). Existing solutions for speech anonymization mainly focus on voice conversion or voice modification to convert a raw utterance into another one with similar content but different, or no, identity-related information. However, an alternative approach to share speech data under the form of privacy-preserving representation has been largely under-explored. In this paper, we propose a speech anonymization framework that achieves privacy via noise perturbation to a selected subset of the high-utility representations extracted using a pre-trained speech encoder. The subset is chosen with a Transformer-based privacy-risk saliency estimator. We validate our framework on four tasks, namely, Automatic Speaker Verification (ASV), ASR, SER and Intent Classification (IC) for privacy and utility assessment. Experimental results show that our approach is able to achieve a competitive, or even better, utility compared to the speech anonymization baselines from the VoicePrivacy2022 Challenges, providing the same level of privacy. Moreover, the easily-controlled amount of perturbation allows our framework to have a flexible range of privacy-utility trade-offs without re-training any component.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages

    cs.CL 2024-12 reject novelty 3.0 of 10

    Fine-tuning Wav2Vec2-xlsr-53 on Common Voice audio augmented with pitch shift, Gaussian noise, and band-stop filtering lowers WER and CER in Arabic, Russian, and Portuguese, though the Whisper comparison is overstated.

Pith tools