REVIEW 2 major objections 2 minor 39 references
A Large-Scale Per-Speaker Analysis of Re-identification Risk in Speech Anonymization
T0 review · 2 major / 2 minor · reviewed 2026-06-27 · grok-4.3
Pith's one-line read Re-identification risk after speech anonymization stems from the interplay of attacker, anonymizer, and available speech rather than from fixed speaker properties.
desk verdict The paper's main finding is that speaker re-identification risk in anonymization comes from interactions between attacker, anonymizer, and speech amount rather than fixed speaker properties, shown at scale across nearly 5000 speakers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Per-speaker linkability metric under a worst-case attacker scenario, applied across multiple anonymizers and speech amounts
What would settle it
A demonstration that the same fixed set of speakers remains easy or hard to re-identify across every combination of anonymizer, attacker, and speech length would falsify the interaction claim.
Extended reading notes
Core claim
The paper establishes that linkability scores are highly polarized at the speaker level but that the sets of easy-to-re-identify and hard-to-re-identify speakers vary substantially across different anonymization systems, attacker architectures, and conversation lengths. No single factor explains speaker vulnerability; instead, the re-identification risk emerges from the interaction between the attacker, the anonymizer, and the amount of available speech. These findings challenge the notion of intrinsic speaker-level privacy risks and emphasize the need for evaluation protocols that are explicitly conditioned on the attacker and anonymizer.
Load-bearing premise
The chosen linkability-based metric under a worst-case attacker scenario provides a faithful measure of real-world re-identification risk across the tested systems and speaker populations.
Editorial extensions
If this is right
- Anonymization systems must be evaluated against multiple attackers rather than relying on average metrics alone.
- The amount of speech data available to an attacker determines which speakers face elevated risk in any given setup.
- Evaluation protocols need to condition results explicitly on both the attacker model and the anonymizer chosen.
- There is no universal ranking of speaker privacy that holds independently of the configuration.
Reading between the lines
- This interaction view of risk may apply to other biometric anonymization tasks such as face or gait data.
- Designers could explore adaptive anonymizers tuned to anticipated attacker profiles rather than fixed transformations.
- Real-world deployments would benefit from testing privacy under a range of attacker capabilities instead of single worst-case assumptions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that re-identification risk in speech anonymization is not an intrinsic speaker property but emerges from interactions among the attacker, anonymizer, and available speech amount. This is based on a large-scale per-speaker analysis of nearly 5,000 speakers using a linkability metric under a worst-case attacker scenario, across multiple anonymization systems, attacker architectures, and conversation lengths. Results show highly polarized per-speaker linkability scores whose easy/hard sets vary substantially by configuration, with no single factor explaining vulnerability, leading to a call for attacker- and anonymizer-conditioned evaluation protocols.
Significance. If the empirical findings hold, the work would meaningfully advance privacy evaluation in speech processing by moving beyond aggregate metrics such as EER to demonstrate context-dependent speaker vulnerability. It supplies concrete evidence that challenges fixed speaker-level privacy assumptions and supports more realistic, interaction-aware assessment protocols with potential impact on system design and deployment.
major comments (2)
- [Evaluation methodology and abstract] The central claim that risk is interaction-driven rather than intrinsic rests on the linkability metric under worst-case attacker being a faithful proxy for re-identification risk. The manuscript reports polarized scores and configuration-dependent sets but provides no validation (e.g., correlation with attack success under partial anonymizer knowledge) that this worst-case metric tracks practical re-identification outcomes; without such grounding the observed variation could be protocol-specific rather than evidence against intrinsic risks.
- [Results] The assertion that 'no single factor explains speaker vulnerability' is load-bearing for the interaction conclusion, yet the results section does not report controls, regression analysis, or ablation over speaker attributes (age, accent, duration statistics) to substantiate that claim; the polarization alone does not rule out latent speaker factors interacting with the tested configurations.
minor comments (2)
- The exact speaker count, dataset splits, and precise definitions of the linkability metric and worst-case attacker should be stated in the methods section rather than summarized only in the abstract.
- Figure legends for polarization plots should explicitly define the linkability score range, threshold used for easy/hard classification, and how sets are compared across configurations.
Simulated Author's Rebuttal
We thank the referee for their constructive comments, which help clarify the scope and grounding of our claims. We address each major comment below, proposing targeted revisions to strengthen the manuscript while preserving its core empirical contributions on configuration-dependent speaker vulnerability.
read point-by-point responses
-
Referee: [Evaluation methodology and abstract] The central claim that risk is interaction-driven rather than intrinsic rests on the linkability metric under worst-case attacker being a faithful proxy for re-identification risk. The manuscript reports polarized scores and configuration-dependent sets but provides no validation (e.g., correlation with attack success under partial anonymizer knowledge) that this worst-case metric tracks practical re-identification outcomes; without such grounding the observed variation could be protocol-specific rather than evidence against intrinsic risks.
Authors: We agree that explicit validation correlating worst-case linkability with re-identification success under partial anonymizer knowledge would provide stronger grounding. The worst-case linkability metric is nevertheless a standard upper-bound proxy in the speech privacy literature, chosen here precisely to expose maximum risk. The central observation—that easy/hard speaker sets shift markedly across anonymizer-attacker-length configurations—holds within a fixed evaluation protocol and therefore cannot be dismissed as a protocol artifact. In revision we will add a dedicated limitations paragraph discussing the metric's relation to practical attacks and citing prior linkability validations, constituting a partial revision. revision: partial
-
Referee: [Results] The assertion that 'no single factor explains speaker vulnerability' is load-bearing for the interaction conclusion, yet the results section does not report controls, regression analysis, or ablation over speaker attributes (age, accent, duration statistics) to substantiate that claim; the polarization alone does not rule out latent speaker factors interacting with the tested configurations.
Authors: We acknowledge that the manuscript does not present formal regression or ablation analyses over speaker attributes. The claim rests on the empirical finding that speaker-level linkability rankings are highly unstable across the tested configurations; if vulnerability were driven by fixed intrinsic speaker properties, the easy/hard partitions would remain largely invariant. To address the referee's point directly, the revised version will include a new subsection reporting Spearman correlations and simple regression models between linkability scores and available speaker metadata (utterance duration, speaker sex where annotated). Preliminary checks indicate no single attribute explains the observed variance, but we will report the full results. This constitutes a partial revision. revision: partial
Circularity Check
No circularity: empirical measurement study with no derivations or fitted predictions
full rationale
This is a large-scale empirical analysis of per-speaker linkability scores under different anonymization and attacker configurations. The central claims rest on direct measurements of nearly 5,000 speakers across systems and conversation lengths, with observations that easy/hard speaker sets vary by configuration. No equations, parameter fitting, self-citations used as load-bearing uniqueness theorems, or renamings of known results are described that would reduce any result to its inputs by construction. The work is self-contained against external benchmarks as a measurement study.
Assumptions & free parameters
Cite this review
Pith. "Pith review of A Large-Scale Per-Speaker Analysis of Re-identification Risk in Speech Anonymization." pith.science (2026). https://pith.science/paper/D5QXQB7D
@misc{pith2026260607210,
author = {Pith},
title = {Pith review of: A Large-Scale Per-Speaker Analysis of Re-identification Risk in Speech Anonymization},
year = {2026},
howpublished = {\url{https://pith.science/paper/D5QXQB7D}},
note = {Machine review of arXiv:2606.07210}
}
read the original abstract
Speech anonymization is commonly evaluated using averagecase metrics such as the equal error rate, which can hide large disparities in re-identification risks across individuals. In this paper, we conduct a large-scale per-speaker privacy analysis using a linkability-based metric under a worst-case scenario. Nearly 5,000 speakers are evaluated across multiple anonymization systems, attacker architectures, and conversation lengths. While linkability scores are highly polarized at the speaker level, the sets of easy to re-identify and hard to re-identify speakers vary substantially across configurations. We show that no single factor explains speaker vulnerability. Instead, the re-identification risk emerges from the interaction between the attacker, the anonymizer, and the amount of available speech. These results challenge the notion of intrinsic speaker-level privacy risks and emphasize the need for evaluation protocols that are explicitly conditioned on the attacker and anonymizer.
Figures
Reference graph
Works this paper leans on
-
[1]
A Large-Scale Per-Speaker Analysis of Re-identification Risk in Speech Anonymization
Introduction Speech conveys a multitude of personal information about the speaker such as biometric identity, age, gender, health condi- tion, or emotional state [1]. As such, the storage and pro- cessing of large speech datasets puts privacy protection at risk [1, 2]. Existing data protection regulations, such as those outlined in the General Data Protec...
work page Pith review arXiv 2020
-
[2]
Attacker strategy Within the speech anonymization framework, the attacker is modeled as an ASV system
Methodology 2.1. Attacker strategy Within the speech anonymization framework, the attacker is modeled as an ASV system. This deep neural network, trained on a speaker identification task, learns to extract speaker- discriminative acoustic features from speech. The last hidden layer is used as speaker representation (speaker embedding) of the utterance. Du...
-
[3]
Our code is available for a reproductibility purpose2
Experimental setup This section presents our experimental protocol. Our code is available for a reproductibility purpose2. 3.1. Datasets Table 1 describes the datasets used in our experiments. We use LibriSpeech [16] to train the ASV models (more specifically we use the train-clean-360 split), and CommonV oice’s 11th English release (CV 11.0) [17] as the ...
2025
-
[4]
Per-speaker score distributions Figure 1 shows the 18 speaker score distributions (2 anonymiza- tion systems×3 attackers×3 conversation lengthL)
Results 4.1. Per-speaker score distributions Figure 1 shows the 18 speaker score distributions (2 anonymiza- tion systems×3 attackers×3 conversation lengthL). As shown in the figure, the WavLM ECAPA attacker consistently outperforms ResNet and ECAPA, and B3 is systematically eas- ier to attack than B5. WhenL= 1most speakers obtain a linkability close to 0...
-
[5]
Conclusion We conducted the first large-scale per-speaker privacy risk anal- ysis across multiple anonymization and attacker architectures and showed that the privacy risk does not depend solely on speakers. Instead, vulnerability to re-identification emerges from the interaction of the attacker’s ASV architecture, the anonymization system, and the amount...
-
[6]
It was provided with computer and storage resources by GENCI at IDRIS thanks to the grant 2025-AD011015838R1 on the supercomputer Jean Zay’s V100 partition
Acknowledgments This work was supported by the French Agence Nationale de la Recherche via the SpeechPrivacy project (ANR-23-CE23- 0022). It was provided with computer and storage resources by GENCI at IDRIS thanks to the grant 2025-AD011015838R1 on the supercomputer Jean Zay’s V100 partition
2025
-
[7]
The authors declare that generative AI tools have been used solely to correct grammar
Use of Generative AI Disclosure We acknowledge the ISCA policy regarding the use of genera- tive AI tools. The authors declare that generative AI tools have been used solely to correct grammar. No such tools were used to write significant parts of this manuscript
-
[8]
Privacy implica- tions of voice and speech analysis – information disclosure by inference,
J. L. Kr ¨oger, O. H.-M. Lutz, and P. Raschke, “Privacy implica- tions of voice and speech analysis – information disclosure by inference,” inPrivacy and Identity Management. Data for Better Living: AI and Privacy: 14th IFIP WG 9.2, 9.6/11.7, 11.6/SIG 9.2.2 International Summer School, 2020, pp. 242–258
2020
Show all 39 references
-
[9]
Preserving privacy in speaker and speech characterisation,
A. Nautsch, A. Jim ´enez, A. Treiber, J. Kolberg, C. Jasserand, E. Kindt, H. Delgado, M. Todisco, M. A. Hmani, A. Mtibaa, M. A. Abdelraheem, A. Abad, F. Teixeira, D. Matrouf, M. Gomez- Barrero, D. Petrovska-Delacr ´etaz, G. Chollet, N. Evans, T. Schneider, J.-F. Bonastre, B. R...
2019
-
[10]
“Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 april 2016 on the protection of natural per- sons with regard to the processing of personal data and on the free movement of such data, and repealing directive 95/46/EC (General Data Protection Regul...
2016
-
[11]
The V oicePrivacy 2020 Challenge: Results and findings,
N. Tomashenko, X. Wang, E. Vincent, J. Patino, B. M. L. Srivas- tava, P.-G. No´e, A. Nautsch, N. Evans, J. Yamagishi, B. OBrien, A. Chanclu, J.-F. Bonastre, M. Todisco, and M. Maouche, “The V oicePrivacy 2020 Challenge: Results and findings,”Computer Speech & Language, vol. 74...
2020
-
[12]
Speaker recognition based on deep learning: An overview,
Z. Bai and X.-L. Zhang, “Speaker recognition based on deep learning: An overview,”Neural Networks, vol. 140, pp. 65–99, 2021
2021
-
[13]
The first V oicePrivacy Attacker Challenge,
N. Tomashenko, X. Miao, E. Vincent, and J. Yamagishi, “The first V oicePrivacy Attacker Challenge,” inIEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1–2
2025
-
[14]
The V oicePrivacy 2024 Challenge evaluation plan,
N. Tomashenko, X. Miao, P. Champion, S. Meyer, X. Wang, E. Vincent, M. Panariello, N. Evans, J. Yamagishi, and M. Todisco, “The V oicePrivacy 2024 Challenge evaluation plan,” 2024, arXiv preprint arxiv:2404.02677
2024
-
[15]
The privacy ZEBRA: Zero evidence biometric recognition assessment,
A. Nautsch, J. Patino, N. Tomashenko, J. Yamagishi, P.-G. No ´e, J.-F. Bonastre, M. Todisco, and N. Evans, “The privacy ZEBRA: Zero evidence biometric recognition assessment,” inINTER- SPEECH, 21st Annual Conference of the International Speech Communication Association, Shangh...
2020
-
[16]
Legally validated evaluation framework for voice anonymiza- tion,
N. Vauquier, B. M. L. Srivastava, S. A. Hosseini, and E. Vincent, “Legally validated evaluation framework for voice anonymiza- tion,” inInterspeech, 2025, pp. 3229–3233
2025
-
[17]
Anonymiz- ing speaker voices: Easy to imitate, difficult to recognize?
J. Williams, K. Pizzi, N. Tomashenko, and S. Das, “Anonymiz- ing speaker voices: Easy to imitate, difficult to recognize?” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 12 491–12 495
2024
-
[18]
V oice anonymization for all-bias evaluation of the voice privacy chal- lenge baseline systems,
A. Leschanowsky, . E. Gaznepoglu, and N. Peters, “V oice anonymization for all-bias evaluation of the voice privacy chal- lenge baseline systems,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 4785–4789
2024
-
[19]
A comprehensive evalua- tion framework for speaker anonymization systems,
C. Franzreb, T. Polzehl, and S. M ¨oller, “A comprehensive evalua- tion framework for speaker anonymization systems,” in3rd Sym- posium on Security and Privacy in Speech Communication, 2023, pp. 65–72
2023
-
[20]
You are what you say: Exploiting linguistic content for V oicePrivacy attacks,
U. E. Gaznepoglu, A. Leschanowsky, A. Aloradi, P. Singh, D. Tenbrinck, E. A. P. Habets, and N. Peters, “You are what you say: Exploiting linguistic content for V oicePrivacy attacks,” inIn- terspeech, 2025, pp. 4238–4242
2025
-
[21]
Speaker anonymization : representation, evaluation and formal guarantees,
B. M. L. Srivastava, “Speaker anonymization : representation, evaluation and formal guarantees,” Ph.D. dissertation, Universit´e de Lille, 2021. [Online]. Available: https://theses.hal.science/ tel-03674540
2021
-
[22]
Optimizing the dataset for the privacy evaluation of speaker anonymizers,
C. Franzreb, A. Das, T. Polzehl, and S. M ¨oller, “Optimizing the dataset for the privacy evaluation of speaker anonymizers,” in5th Symposium on Security and Privacy in Speech Communication, 2025, pp. 18–26
2025
-
[23]
Lib- rispeech: An ASR corpus based on public domain audio books,
V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- rispeech: An ASR corpus based on public domain audio books,” in2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, pp. 5206–5210
2015
-
[24]
Common V oice: A massively-multilingual speech corpus,
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common V oice: A massively-multilingual speech corpus,” in 12th Language Resources and Evaluation Conference (LREC), 2020, pp. 4218–4222
2020
-
[25]
Speaker anonymization using neural audio codec language models,
M. Panariello, F. Nespoli, M. Todisco, and N. Evans, “Speaker anonymization using neural audio codec language models,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 4725–4729
2024
-
[26]
Prosody is not identity: A speaker anonymization approach using prosody cloning,
S. Meyer, F. Lux, J. Koch, P. Denisov, P. Tilli, and T. Vu, “Prosody is not identity: A speaker anonymization approach using prosody cloning,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5
2023
-
[27]
Anonymizing speech: Evaluating and designing speaker anonymization techniques,
P. Champion, “Anonymizing speech: Evaluating and designing speaker anonymization techniques,” Ph.D. dissertation, Universit´e de Lorraine, 2024. [Online]. Available: https://hal.univ-lorraine. fr/tel-04218098v1
2024
-
[28]
Speaker anonymization with phonetic intermediate representa- tions,
S. Meyer, F. Lux, P. Denisov, J. Koch, P. Tilli, and N. T. Vu, “Speaker anonymization with phonetic intermediate representa- tions,” inInterspeech, 2022, pp. 4925–4929
2022
-
[29]
Anonymizing speech with generative adversarial networks to preserve speaker privacy,
S. Meyer, P. Tilli, P. Denisov, F. Lux, J. Koch, and N. T. Vu, “Anonymizing speech with generative adversarial networks to preserve speaker privacy,” inIEEE Spoken Language Technology, 2022, pp. 912–919
2022
-
[30]
wav2vec 2.0: A framework for self-supervised learning of speech representations,
A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” pp. 12 449–12 460, 2020. [Online]. Avail- able: https://proceedings.neurips.cc/paper files/paper/2020/file/ 92d1e1eb1cd6f9fba3227870bb6d7f07-Paper.pdf
2020
-
[31]
Improving the speaker anonymization evaluation’s robustness to target speakers with adversarial learning,
C. Franzreb, A. Das, T. Polzehl, and S. M ¨oller, “Improving the speaker anonymization evaluation’s robustness to target speakers with adversarial learning,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
2026
-
[32]
ECAPA- TDNN: Emphasized channel attention, propagation and aggrega- tion in TDNN based speaker verification,
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA- TDNN: Emphasized channel attention, propagation and aggrega- tion in TDNN based speaker verification,” inInterspeech, 2020, pp. 3830–3834
2020
-
[33]
WavLM: Large-scale self-supervised pre-training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selecte...
2022
-
[34]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778
2016
-
[35]
Kiwano: A Cutting-Edge Open- Source Toolkit for Speaker Verification,
M. Rouvier and P.-M. Bousquet, “Kiwano: A Cutting-Edge Open- Source Toolkit for Speaker Verification,” inOdyssey 2026, 2026
2026
-
[36]
DAST: A dual-stream voice anonymization attacker with staged training,
R. Arefeen, X. Miao, R. Tong, A. B. Ng, S. See, and T. Liu, “DAST: A dual-stream voice anonymization attacker with staged training,” 2026. [Online]. Available: https: //arxiv.org/abs/2603.12840
2026
-
[37]
Reshape dimensions network for speaker recognition,
I. Yakovlev, R. Makarov, A. Balykin, P. Malov, A. Okhotnikov, and N. Torgashov, “Reshape dimensions network for speaker recognition,” inInterspeech, 2024, p. 32353239
2024
-
[38]
Fast adaptation of pre- trained speaker verification system for source speaker tracking,
X. Lyu, Y . Wang, T. Zhao, and H. Liu, “Fast adaptation of pre- trained speaker verification system for source speaker tracking,” inIEEE International Conference on Acoustics, Speech and Sig- nal Processing (ICASSP), 2025, pp. 1–2
2025
-
[39]
Reliability crite- rion based on learning-phase entropy for speaker recognition with neural network,
P.-M. Bousquet, M. Rouvier, and J.-F. Bonastre, “Reliability crite- rion based on learning-phase entropy for speaker recognition with neural network,” inInterspeech, 2022, pp. 281–285
2022
Reviewed June 27, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.