REVIEW 2 major objections 5 minor 104 references
This paper shows that fixed-threshold voice-clone attribution is unreliable and unfair in professional voice actors.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 01:37 UTC pith:X7PMKNWV
load-bearing objection A careful, unusually transparent large-scale measurement of a real misidentification floor; the 'geometry-limited' label overreaches, but the empirical core deserves a serious referee. the 2 major comments →
A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that in this high-density professional-speaker domain, the embedding space retains a geometry-limited misidentification floor and segment-level hubness that high headline rank-1 accuracy masks. The floor is not a bad threshold: monotone recalibration cannot change a rank-1 error, and discriminative re-rankers that genuinely reorder neighbors, including LDA, WCCN, two-covariance PLDA, neural PLDA, and a pair-feature MLP, reduce but do not remove it. The residual is attributed to the arrangement of embeddings themselves, driven by inter-speaker proximity and intra-speaker style range. A separate real-versus-synthetic covariate shift pushes clones off the real-speaker manif
What carries the argument
The load-bearing object is the closed-set misidentification floor, defined as the fraction of queries whose nearest-neighbor speaker differs from the true speaker. Because no monotone calibration can move an argmax decision, the paper tests whether re-ranking that genuinely reorders neighbors, namely LDA, WCCN, two-covariance PLDA log-likelihood ratio, neural PLDA, and a pair-feature MLP, lowers the floor. The floor's persistence across this back-end suite is taken as evidence that the limit is geometric. The synthetic-versus-real shift is diagnosed with a linear separability probe and controlled by codec, vocoder copy-synthesis, and content-matched experiments, supporting a covariate-shift
Load-bearing premise
The 'geometry-limited' label rests on the premise that the tested suite of scoring methods (cosine, AS-norm, LDA, WCCN, two-covariance PLDA, neural PLDA, pair-MLP) is representative of what can be done with a fixed embedding space, with encoder fine-tuning left untested.
What would settle it
Train a strong re-ranker or fine-tune a domain-matched encoder on consented voice-actor audio and show that the closed-set misidentification floor drops to the level of the matched control populations, or find a single operating point where both wrongful accusation and missed attribution fall below the paper's reported rates; either would break the claimed geometry limit.
If this is right
- Any fixed-threshold clone-attribution system in this domain carries a minimum unavoidable misidentification rate; adjusting the threshold cannot resolve the false-accusation versus missed-detection trade-off.
- A domain-matched encoder trained on the target speaker population is a necessary mitigation, cutting wrongful accusation from roughly 12 to 19 percent down to 1.5 to 10 percent and removing a four-fold gender gap, but it does not eliminate the floor.
- Detection thresholds calibrated on real speech do not transfer to synthetic audio: even high-fidelity clones that cluster cleanly by target speaker can fall below the real-calibrated operating point.
- False attributions are systematic, not random: confusion graphs show community structure and reciprocity, meaning the wrong actor is structurally determined and can be anticipated.
- Similarity scores in this domain should support screening and human review, not autonomous enforcement or legal penalties, unless an anti-spoofing gate and per-speaker calibration are added.
Where Pith is reading between the lines
- If the floor is truly geometric relative to the standard back-end family, then encoder fine-tuning on in-domain data is the key untested lever; the paper's own domain-matched encoder comparison suggests representation choice can move the floor, so a consent-based, domain-matched model might push it further down than the tested generic back-ends.
- The synthetic-versus-real shift implies a concrete benchmark: measuring threshold-transfer failure across synthesizers as a calibration-transfer metric, which the paper approximates but does not formalize as a standalone evaluation.
- The reciprocity and community structure of confusions suggest that per-voice-type or per-cluster calibration might outperform per-speaker calibration, a testable variant that the paper does not explore.
- The finding that roughly half of non-enrolled clones falsely accuse an enrolled actor on a generic encoder implies that any policy or platform relying on similarity scores alone will systematically implicate innocent enrolled voice actors; this policy consequence follows from the paper's data but is not stated as a recommendation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies speaker-embedding identification in a corpus of 1,168 professional Japanese voice actors (56,568 segments). Across eight encoders and two ensembles, it reports a closed-set rank-1 misidentification floor that survives calibration, AS-norm, and a broad discriminative re-ranking suite (LDA, WCCN, two-covariance PLDA with moment and EM fits, neural PLDA, pair-MLP) on speaker-disjoint splits. The best same-session configuration leaves roughly 2.6% misID (1.4% for the domain-matched animeva encoder), several-fold above matched controls; under session-disjoint trials the floor is about 13%. The same geometry is shown to drive clone-attribution failures: on a generic English encoder, roughly half of clones of non-enrolled speakers falsely accuse an enrolled actor, while a real-vs-synthetic covariate shift causes misses; a domain-matched encoder mitigates but does not remove the problem. The paper includes extensive confound controls, a non-mated probe, a designed-voice probe, and a detailed limitations section.
Significance. If the results hold, this is a valuable negative result for voice-clone attribution: it shows that naive similarity-threshold defenses are unreliable and unfair in a high-density professional-speaker domain, and that the failure persists across standard scoring back-ends. The paper's strengths are its unusually comprehensive control suite (speaker-disjoint splits, bootstrapped CIs, EM refit cross-check, matched JVS/Common Voice populations, codec/channel/vocoder/content controls, session-disjoint and non-mated probes), reproducible code release, and honest reporting of fragile results, including the retraction of in-degree hubness as a segment-count artifact. The animeva comparison is a constructive mitigation direction. The main weakness is that the 'geometry-limited' label in the title and abstract is stronger than the body's own carefully qualified claim, which explicitly leaves encoder fine-tuning untested.
major comments (2)
- [Title, Abstract, Section IV-B, Section VIII] The phrase 'geometry-limited' overstates the evidence. Section IV-B concedes 'the only remaining untested lever is encoder fine-tuning, and the animeva-vs-generic gap shows representation choice can move the floor,' and Table III shows the floor drops from 9.0% (ECAPA-TDNN) to 1.4% (animeva) at the PLDA stage. Thus the floor is a property of the specific encoder/back-end combination tested, not an established domain invariant. The abstract's 'residual is a limit of the embedding geometry' and the title 'Geometry-Limited Identification Floor' should be qualified (e.g., 'for the tested fixed embeddings and re-ranking suite') or the encoder-fine-tuning caveat should be moved into the abstract. The measured error rates are not in question; the label is.
- [Section IV-C and Table VII] The 'several-fold above matched controls' ratios at the PLDA stage are based on very small control training sets: style-varied JVS has 34 train-side speakers (LDA/PCA dimension capped at 33) and the CV condition is reported as unstable at 8 segments/speaker. The text acknowledges this, but the ratios (1.9–5.4×) are presented without confidence intervals or a sensitivity analysis. Since these ratios are a load-bearing part of the 'genuinely harder' claim, please add speaker-level bootstrap CIs for the control ratios, or a sensitivity sweep over train-speaker counts, to show the several-fold gap is not a small-sample artifact.
minor comments (5)
- [Throughout] The word 'Voice' is rendered as 'V oice' in many places ('V oice-Clone', 'V oice Actor', 'V oicePrivacy'). Fix the typography/rendering issue.
- [Abstract] The phrase 'session-disjoint, re-ranking lowers the floor only to 13.0%' is ambiguous; clarify that 13.0% is the post-re-ranking PLDA-stage misID under session-disjoint trials, not the reduction achieved by re-ranking.
- [Table III caption] State explicitly that all values are on the N=497 eval-half gallery (not the full ~1,100-speaker gallery of Table II), to avoid confusion between tables with different gallery sizes.
- [Figure 2] The critical-margin band (|margin|<0.02) is described in the text but not visibly marked in the figure; consider adding a shaded band or explicit annotation in each panel.
- [Throughout] The '1 :NEER' notation with a space before the colon is unconventional; use '1:N EER' for readability.
Circularity Check
No significant circularity: the headline rates are measured on speaker-disjoint splits with held-out thresholds; the 'geometry-limited' label is conditional, not definitionally forced.
full rationale
The derivation chain is measurement-based rather than reduction-by-construction. The misidentification floor is computed directly on speaker-disjoint evaluation halves: the paper states 'All columns are averaged over three speaker-disjoint splits' and 'train and test speakers are disjoint.' Clone-probe thresholds are set at the real-vs-real EER operating point and then applied to clones, so attribution recall and wrongful-accusation rates are held-out applications, not fitted predictions. The paper explicitly discloses the one in-sample optimistic metric: 'both are in-sample optimistic bounds' (actCllr), and the real-vs-synthetic separator is labeled a diagnostic probe, 'an in-distribution, optimistic stand-in for such a gate,' not a predicted deployment result. No parameter is fitted to the target result, and no load-bearing self-citation appears: the encoders and back-ends are external standard methods or public artifacts. The central 'geometry-limited' label is representation-relative — the paper itself concedes 'the only remaining untested lever is encoder fine-tuning, and the animeva-vs-generic gap shows representation choice can move the floor.' That is a scope limitation or an over-generalization risk, not circularity: the measured misID rates, control comparisons, and re-ranking results do not reduce to their own inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (5)
- critical-margin band width =
0.02 cosine
- AS-norm cohort size =
top-300 impostor cohort from eval set
- enrollment centroid size =
first 8 segments per speaker
- LDA/PLDA dimensionality caps =
LDA min(150,#train−1); PLDA PCA min(200,#train−1)
- clone-probe operating threshold τ =
real-vs-real EER point (per corpus); cross-file point as sensitivity
axioms (5)
- domain assumption Cosine similarity in centered, L2-normalized speaker-embedding space is an appropriate score for identity comparison
- domain assumption The speaker_id label spanning a voice actor's full performed range (narration, dialogue, character voices) is the operationally correct identity definition for attribution
- ad hoc to paper The tested back-end family (cosine, AS-norm, LDA, WCCN, two-covariance PLDA, neural PLDA, pair-MLP) is representative of the space of scoring approaches for fixed embeddings
- domain assumption A linear separator in speaker-embedding space is an appropriate in-distribution stand-in for what a dedicated anti-spoofing countermeasure could detect
- domain assumption JVS and Common Voice JA are suitable matched controls for the voice-actor domain
read the original abstract
A voice actor's voice is their asset, and AI cloning directly threatens it. The natural defense flags the enrolled actor whose embedding similarity to a suspect recording crosses a threshold. We show it fails where it is most needed: trained voices crowd the embedding space, and each actor performs many styles. On 1,168 Japanese voice actors (56,568 segments, ~63 h), a misidentification floor survives calibration, score normalization, and discriminative re-ranking (linear and nonlinear, including PLDA): the residual is a limit of the embedding geometry, not of the back-ends we evaluate. The best ensemble still leaves ~2.6% closed-set misidentification, several-fold above matched controls; session-disjoint, re-ranking lowers the floor only to 13.0%. The same crowding drives false attribution: on a generic English encoder, roughly half the clones of non-enrolled people falsely accuse an enrolled actor, while -- by a separate real-vs-synthetic shift -- 32% of Seed-VC clones of enrolled targets are missed at the same threshold; one operating point couples the two, and none escapes both. A domain-matched, voice-actor-trained encoder mitigates substantially (a four-fold gender gap vanishes; wrongful misattribution falls to 1.5-10%), but does not remove the floor. Controls (codec, channel, vocoder, content) support reading the miss rate as a real-versus-synthetic covariate shift, not missing speaker information. Fixed-threshold clone attribution is thus unreliable here, and on a generic encoder unfair. Robust attribution must extend spoofing-aware speaker verification to open-set 1:N (anti-spoofing gate, domain-matched encoder, per-speaker calibration, abstain option), and even then supports detection, not autonomous enforcement.
Figures
Reference graph
Works this paper leans on
-
[1]
Neural codec language models are zero-shot text to speech synthesizers,
C. Wang, S. Chen, Y . Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y . Liu, H. Wang, J. Li, L. He, S. Zhao, and F. Wei, “Neural codec language models are zero-shot text to speech synthesizers,” arXiv:2301.02111, 2023
Pith/arXiv arXiv 2023
-
[2]
Labor, power, and belonging: The work of voice in the age of AI reproduction,
S. G. Almeda, R. Netzorg, I. Li, E. Tam, S. Ma, and B. T. Wei, “Labor, power, and belonging: The work of voice in the age of AI reproduction,” inProc. ACM Conf. Fairness, Accountability, and Transparency (FAccT), 2025, pp. 1238–1249
2025
-
[3]
NOMORE無断生成AI (No More Unauthorized Generative AI) campaign,
NOMORE Mudan Seisei AI Campaign, “NOMORE無断生成AI (No More Unauthorized Generative AI) campaign,” https://nomore-mudan. com/, 2024, launched 2024-10-15 by 26 professional Japanese voice actors; accessed 2026-07
2024
-
[4]
Copyright and artificial intelligence, part 1: Digital replicas,
U.S. Copyright Office, “Copyright and artificial intelligence, part 1: Digital replicas,” U.S. Copyright Office, Tech. Rep., 2024, report of the Register of Copyrights, July 2024
2024
-
[5]
Ensuring likeness, voice, and image security (ELVIS) act of 2024,
State of Tennessee, “Ensuring likeness, voice, and image security (ELVIS) act of 2024,” Tennessee Public Chapter No. 588 (H.B. 2091, 113th General Assembly), 2024, signed March 21, 2024; effective July 1, 2024
2024
-
[6]
Judgment of February 2, 2012 (Pink Lady case), Minshu vol. 66, no. 2, p. 89,
Supreme Court of Japan, “Judgment of February 2, 2012 (Pink Lady case), Minshu vol. 66, no. 2, p. 89,” 2012, case No. 2009 (Ju) 2056; publicity right recognized in case law, scoped to uses exploiting a persona’s customer-drawing power
2012
-
[7]
Speaker verification by human participants with acting voices perceived as different characters: An experimental study on VTuber fans,
D. Hayashi, K. Hakui, R. Yamamoto, and M. Morise, “Speaker verification by human participants with acting voices perceived as different characters: An experimental study on VTuber fans,” inProc. Spring Meeting Acoust. Soc. Jpn., Mar. 2026, pp. 739–742, in Japanese
2026
-
[8]
Probabilistic linear discriminant analysis for inferences about identity,
S. J. D. Prince and J. H. Elder, “Probabilistic linear discriminant analysis for inferences about identity,” inProc. IEEE Int. Conf. Comput. Vis. (ICCV), 2007, pp. 1–8
2007
-
[9]
Front-end factor analysis for speaker verification,
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”IEEE Trans. Audio, Speech, Language Process., vol. 19, no. 4, pp. 788–798, 2011
2011
-
[10]
Bayesian speaker verification with heavy-tailed priors,
P. Kenny, “Bayesian speaker verification with heavy-tailed priors,” in Proc. Speaker Lang. Recognit. Workshop (Odyssey), 2010, p. paper 14
2010
-
[11]
Analysis of i-vector length normalization in speaker recognition systems,
D. Garcia-Romero and C. Y . Espy-Wilson, “Analysis of i-vector length normalization in speaker recognition systems,” inProc. INTER- SPEECH, 2011, pp. 249–252
2011
-
[12]
Comparison of speaker recognition approaches for real appli- cations,
S. Cumani, P. D. Batzu, D. Colibro, C. Vair, P. Laface, and V . Vasi- lakakis, “Comparison of speaker recognition approaches for real appli- cations,” inProc. INTERSPEECH, 2011, pp. 2365–2368
2011
-
[13]
Application-independent evaluation of speaker detection,
N. Brümmer and J. du Preez, “Application-independent evaluation of speaker detection,”Comput. Speech Lang., vol. 20, no. 2–3, pp. 230– 275, 2006
2006
-
[14]
Hubs in space: Popular nearest neighbors in high-dimensional data,
M. Radovanovi ´c, A. Nanopoulos, and M. Ivanovi ´c, “Hubs in space: Popular nearest neighbors in high-dimensional data,”J. Mach. Learn. Res., vol. 11, pp. 2487–2531, 2010
2010
-
[15]
Score normaliza- tion for text-independent speaker verification systems,
R. Auckenthaler, M. Carey, and H. Lloyd-Thomas, “Score normaliza- tion for text-independent speaker verification systems,”Digit. Signal Process., vol. 10, no. 1-3, pp. 42–54, 2000
2000
-
[16]
Rose,Forensic Speaker Identification
P. Rose,Forensic Speaker Identification. London: Taylor & Francis, 2002
2002
-
[17]
Forensic voice comparison and the paradigm shift,
G. S. Morrison, “Forensic voice comparison and the paradigm shift,” Sci. Justice, vol. 49, no. 4, pp. 298–308, 2009
2009
-
[18]
The relevant population in forensic voice comparison: Effects of varying delimitations of social class and age,
V . Hughes and P. Foulkes, “The relevant population in forensic voice comparison: Effects of varying delimitations of social class and age,” Speech Commun., vol. 66, pp. 218–230, 2015
2015
-
[19]
The use of multiple measurements in taxonomic prob- lems,
R. A. Fisher, “The use of multiple measurements in taxonomic prob- lems,”Ann. Eugenics, vol. 7, no. 2, pp. 179–188, 1936
1936
-
[20]
Within-class covariance normalization for SVM-based speaker recognition,
A. O. Hatch, S. Kajarekar, and A. Stolcke, “Within-class covariance normalization for SVM-based speaker recognition,” inProc. INTER- SPEECH, 2006, pp. 1471–1474
2006
-
[21]
The speaker partitioning problem,
N. Brümmer and E. de Villiers, “The speaker partitioning problem,” in Proc. Speaker Lang. Recognit. Workshop (Odyssey), 2010
2010
-
[22]
JVS corpus: Free Japanese multi-speaker voice corpus,
S. Takamichi, K. Mitsui, Y . Saito, T. Koriyama, N. Tanji, and H. Saruwatari, “JVS corpus: Free Japanese multi-speaker voice corpus,” arXiv:1908.06248, 2019
Pith/arXiv arXiv 1908
-
[23]
Common V oice: A massively-multilingual speech corpus,
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common V oice: A massively-multilingual speech corpus,” inProc. Lang. Resour. Eval. Conf. (LREC), 2020
2020
-
[24]
Phoneme recognition using time-delay neural networks,
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, “Phoneme recognition using time-delay neural networks,”IEEE Trans. Acoust., Speech, Signal Process., vol. 37, no. 3, pp. 328–339, 1989
1989
-
[25]
X-vectors: Robust DNN embeddings for speaker recognition,
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” inProc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP), 2018
2018
-
[26]
ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,” inProc. INTERSPEECH, 2020. 23
2020
-
[27]
WavLM: Large-scale self-supervised pre- training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large-scale self-supervised pre- training for full stack speech processing,”IEEE J. Sel. Topics Signal Process., vol. 16, no. 6, pp. 1505–1518, 2022
2022
-
[28]
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”IEEE/ACM Trans. Audio, Speech, Language Process., vol. 29, pp. 3451–3460, 2021
2021
-
[29]
CAM++: A fast and efficient network for speaker verification using context-aware masking,
H. Wang, S. Zheng, Y . Chen, L. Cheng, and Q. Chen, “CAM++: A fast and efficient network for speaker verification using context-aware masking,” inProc. INTERSPEECH, 2023
2023
-
[30]
Reshape dimensions network for speaker recognition,
I. Yakovlev, R. Makarov, A. Balykin, P. Malov, A. Okhotnikov, and N. Torgashov, “Reshape dimensions network for speaker recognition,” inProc. INTERSPEECH, 2024
2024
-
[31]
An enhanced Res2Net with local and global feature fusion for speaker verification,
Y . Chen, S. Zheng, H. Wang, L. Cheng, Q. Chen, and J. Qi, “An enhanced Res2Net with local and global feature fusion for speaker verification,” inProc. INTERSPEECH, 2023, pp. 2228–2232
2023
-
[32]
V oxCeleb: A large-scale speaker identification dataset,
A. Nagrani, J. S. Chung, and A. Zisserman, “V oxCeleb: A large-scale speaker identification dataset,” inProc. INTERSPEECH, 2017
2017
-
[33]
Analysis of score normalization in multilingual speaker recognition,
P. Mat ˇejka, O. Novotný, O. Plchot, L. Burget, M. D. Sánchez, and J. ˇCernocký, “Analysis of score normalization in multilingual speaker recognition,” inProc. INTERSPEECH, 2017, pp. 1567–1571
2017
-
[34]
Local and global scaling reduce hubs in space,
D. Schnitzer, A. Flexer, M. Schedl, and G. Widmer, “Local and global scaling reduce hubs in space,”J. Mach. Learn. Res., vol. 13, pp. 2871– 2902, 2012
2012
-
[35]
Accurate image search using the contextual dissimilarity measure,
H. Jégou, C. Schmid, H. Harzallah, and J. Verbeek, “Accurate image search using the contextual dissimilarity measure,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 32, no. 1, pp. 2–11, 2010
2010
-
[36]
Word translation without parallel data,
G. Lample, A. Conneau, M. Ranzato, L. Denoyer, and H. Jégou, “Word translation without parallel data,” inProc. Int. Conf. Learn. Representations (ICLR), 2018. [Online]. Available: https: //openreview.net/forum?id=H196sainb
2018
-
[37]
A comprehensive empirical comparison of hubness reduction in high-dimensional spaces,
R. Feldbauer and A. Flexer, “A comprehensive empirical comparison of hubness reduction in high-dimensional spaces,”Knowl. Inf. Syst., vol. 59, no. 1, pp. 137–166, 2019
2019
-
[38]
Introducing the V oicePrivacy initiative,
N. Tomashenko, B. M. L. Srivastava, X. Wang, E. Vincent, A. Nautsch, J. Yamagishi, N. Evans, J. Patino, J.-F. Bonastre, P.-G. Noé, and M. Todisco, “Introducing the V oicePrivacy initiative,” inProc. INTER- SPEECH, 2020, pp. 1693–1697
2020
-
[39]
ASVspoof 2015: The first automatic speaker verification spoofing and countermeasures challenge,
Z. Wu, T. Kinnunen, N. Evans, J. Yamagishi, C. Hanilçi, M. Sahidullah, and A. Sizov, “ASVspoof 2015: The first automatic speaker verification spoofing and countermeasures challenge,” inProc. INTERSPEECH, 2015, pp. 2037–2041
2015
-
[40]
ASVspoof 2021: Accelerating progress in spoofed and deepfake speech detection,
J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, X. Liu, K. A. Lee, T. Kinnunen, N. Evans, and H. Delgado, “ASVspoof 2021: Accelerating progress in spoofed and deepfake speech detection,” inProc. Autom. Speaker Verification and Spoofing Countermeasures Challenge (ASVspoof), 2021, pp. 47–54
2021
-
[41]
t-DCF: A detection cost function for the tandem assessment of spoofing countermeasures and automatic speaker verification,
T. Kinnunen, K. A. Lee, H. Delgado, N. Evans, M. Todisco, M. Sahidullah, J. Yamagishi, and D. A. Reynolds, “t-DCF: A detection cost function for the tandem assessment of spoofing countermeasures and automatic speaker verification,” inProc. Speaker Lang. Recognit. Workshop (Odyssey), 2018, pp. 312–319
2018
-
[42]
SASV 2022: The first spoofing- aware speaker verification challenge,
J.-w. Jung, H. Tak, H.-j. Shim, H.-S. Heo, B.-J. Lee, S.-W. Chung, H.-J. Yu, N. Evans, and T. Kinnunen, “SASV 2022: The first spoofing- aware speaker verification challenge,” inProc. INTERSPEECH, 2022, pp. 2893–2897
2022
-
[43]
ASVspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,
X. Wang, H. Delgado, H. Tak, J.-w. Jung, H.-j. Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. Kinnunen, N. Evans, K. A. Lee, and J. Yamagishi, “ASVspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,” inProc. Autom. Speaker Verification and Spoofing Countermeasures Challenge (ASVspoof), 2024, pp. 1–8
2024
-
[44]
SpoofCeleb: Speech deepfake detection and SASV in the wild,
J.-w. Jung, Y . Wu, X. Wang, J.-H. Kim, S. Maiti, Y . Matsunaga, H.-j. Shim, J. Tian, N. Evans, J. S. Chung, W. Zhang, S. Um, S. Takamichi, and S. Watanabe, “SpoofCeleb: Speech deepfake detection and SASV in the wild,”IEEE Open J. Signal Process., vol. 6, pp. 68–77, 2025
2025
-
[45]
Does audio deepfake detection generalize?
N. Müller, P. Czempin, F. Diekmann, A. Froghyar, and K. Böttinger, “Does audio deepfake detection generalize?” inProc. INTERSPEECH, 2022, pp. 2783–2787
2022
-
[46]
V oxBlink2: A 100K+ speaker recognition corpus and the open-set speaker-identification benchmark,
Y . Lin, M. Cheng, F. Zhang, Y . Gao, S. Zhang, and M. Li, “V oxBlink2: A 100K+ speaker recognition corpus and the open-set speaker-identification benchmark,” inProc. INTERSPEECH, 2024, arXiv:2407.11510
Pith/arXiv arXiv 2024
-
[47]
An initial investigation for detecting vocoder fingerprints of fake audio,
X. Yan, J. Yi, J. Tao, C. Wang, H. Ma, T. Wang, S. Wang, and R. Fu, “An initial investigation for detecting vocoder fingerprints of fake audio,” inProc. Int. Workshop on Deepfake Detection for Audio Multimedia (DDAM), 2022, pp. 61–68
2022
-
[48]
Source tracing of audio deepfake systems,
N. Klein, T. Chen, H. Tak, R. Casal, and E. Khoury, “Source tracing of audio deepfake systems,” inProc. INTERSPEECH, 2024, pp. 1100– 1104
2024
-
[49]
Source ver- ification for speech deepfakes,
V . Negroni, D. Salvi, P. Bestagini, and S. Tubaro, “Source ver- ification for speech deepfakes,” inProc. INTERSPEECH, 2025, arXiv:2505.14188
Pith/arXiv arXiv 2025
-
[50]
Identifying source speakers for voice conversion based spoofing attacks on speaker verification systems,
D. Cai, Z. Cai, and M. Li, “Identifying source speakers for voice conversion based spoofing attacks on speaker verification systems,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP), 2023, pp. 1–5
2023
-
[51]
Bias in automated speaker recogni- tion,
W. T. Hutiri and A. Y . Ding, “Bias in automated speaker recogni- tion,” inProc. ACM Conf. Fairness, Accountability, and Transparency (FAccT), 2022, pp. 230–247
2022
-
[52]
S. Takamichi, L. Kürzinger, T. Saeki, S. Shiota, and S. Watanabe, “JTubeSpeech: Corpus of Japanese speech collected from YouTube for speech recognition and speaker verification,” arXiv:2112.09323, 2021
Pith/arXiv arXiv 2021
-
[53]
anime_speaker_embedding: A speaker-embedding model trained to separate Japanese voice actors,
litagin, “anime_speaker_embedding: A speaker-embedding model trained to separate Japanese voice actors,” https://github.com/litagin02/ anime_speaker_embedding, 2025, software artifact (ECAPA-TDNN, BatchNorm→GroupNorm),variant="va"(v0.2.0, 2025-06); ac- cessed 2026-07
2025
-
[54]
Zero-shot voice conversion with diffusion Transformers,
S. Liu, “Zero-shot voice conversion with diffusion Transformers,” arXiv:2411.09943, 2024
Pith/arXiv arXiv 2024
-
[55]
GPT-SoVITS: Few-shot voice cloning and text-to-speech,
RVC-Boss and contributors, “GPT-SoVITS: Few-shot voice cloning and text-to-speech,” https://github.com/RVC-Boss/GPT-SoVITS, 2024, software artifact; no associated academic paper; accessed 2026-07
2024
-
[56]
Irodori-TTS: A flow matching-based text-to-speech model with emoji-driven style control,
C. Arata, “Irodori-TTS: A flow matching-based text-to-speech model with emoji-driven style control,” https://github.com/Aratako/ Irodori-TTS, 2026, software artifact by Aratako (Chihiro Arata); no as- sociated academic paper; Hugging Face models Irodori-TTS-500M-v3 (clone/TTS) and Irodori-TTS-600M-v3-V oiceDesign (designed-voice); accessed 2026-07
2026
-
[57]
Semantic-DACV AE-Japanese-32dim: Lightweight audio V AE for Japanese speech,
——, “Semantic-DACV AE-Japanese-32dim: Lightweight audio V AE for Japanese speech,” https://huggingface.co/Aratako/ Semantic-DACV AE-Japanese-32dim, 2026, software artifact (Irodori- TTS codec) released under the Hugging Face handle “Aratako”; no associated academic paper; accessed 2026-07
2026
-
[58]
Verification of the effectiveness of a speaker verification system for acting voices,
R. Yamamoto, T. Koumura, D. Hayashi, and M. Morise, “Verification of the effectiveness of a speaker verification system for acting voices,” inProc. Spring Meeting Acoust. Soc. Jpn., Mar. 2026, pp. 1123–1124, in Japanese
2026
-
[59]
Acoustical features related to represen- tations of age and gender in acting voices by female voice actors,
D. Hayashi and M. Morise, “Acoustical features related to represen- tations of age and gender in acting voices by female voice actors,”J. Acoust. Soc. Jpn., vol. 82, no. 3, pp. 124–131, 2026, in Japanese
2026
-
[60]
Experimental research on voice perception with character voices by professional voice actors,
D. Hayashi, “Experimental research on voice perception with character voices by professional voice actors,”Bull. Aichi Shukutoku Univ., Fac. Human Informatics, no. 9, pp. 49–62, 2019, in Japanese
2019
-
[61]
Seiy ¯u meikan (voice-actor directory), Seiyu Grand Prix official web site,
Imagica Infos Co., Ltd., “Seiy ¯u meikan (voice-actor directory), Seiyu Grand Prix official web site,” https://seigura.com/directory/, 2026, in Japanese; directory archived 2026-06 (UTC)
2026
-
[62]
librosa: Audio and music signal analysis in Python,
B. McFee, C. Raffel, D. Liang, D. P. W. Ellis, M. McVicar, E. Bat- tenberg, and O. Nieto, “librosa: Audio and music signal analysis in Python,” inProc. Python in Sci. Conf., 2015, pp. 18–24
2015
-
[63]
japanese-hubert-base-k2,
Reazon Human Interaction Lab, “japanese-hubert-base-k2,” https://huggingface.co/reazon-research/japanese-hubert-base-k2, 2023, japanese HuBERT Base model artifact (commit a9f26026), trained on the ReazonSpeech corpus; accessed 2026-07
2023
-
[64]
ReazonSpeech: A free and massive corpus for Japanese ASR,
Y . Yin, D. Mori, and S. Fujimoto, “ReazonSpeech: A free and massive corpus for Japanese ASR,” inProc. 29th Annual Meeting of the Association for Natural Language Processing, 2023, pp. 1134–1139
2023
-
[65]
xvec- tor_jtubespeech: An x-vector speaker-embedding model trained on JTubeSpeech,
Saruwatari Laboratory (SaruLab), The University of Tokyo, “xvec- tor_jtubespeech: An x-vector speaker-embedding model trained on JTubeSpeech,” https://github.com/sarulab-speech/xvector_jtubespeech, software artifact (sarulab-speech), loaded via torch.hub; accessed 2026- 07
2026
-
[66]
VNDB: The Visual Novel Database,
The Visual Novel Database, “VNDB: The Visual Novel Database,” https://vndb.org, community database of visual novels and their voice- actor credits; accessed 2026-07
2026
-
[67]
Copyright act (act no. 48 of 1970, as amended),
Government of Japan, “Copyright act (act no. 48 of 1970, as amended),” Japanese Law Translation Database System, Min- istry of Justice, https://www.japaneselawtranslation.go.jp/en/laws/view/ 3379, 2018, english translation; accessed 2026-07
1970
-
[68]
Controlling the false discovery rate: A practical and powerful approach to multiple testing,
Y . Benjamini and Y . Hochberg, “Controlling the false discovery rate: A practical and powerful approach to multiple testing,”J. Roy. Statist. Soc. B, vol. 57, no. 1, pp. 289–300, 1995. 24
1995
-
[69]
A simple sequentially rejective multiple test procedure,
S. Holm, “A simple sequentially rejective multiple test procedure,” Scand. J. Statist., vol. 6, no. 2, pp. 65–70, 1979
1979
-
[70]
I. T. Jolliffe,Principal Component Analysis, 2nd ed. New York: Springer, 2002
2002
-
[71]
NPLDA: A deep neural PLDA model for speaker verification,
S. Ramoji, P. Krishnan, and S. Ganapathy, “NPLDA: A deep neural PLDA model for speaker verification,” inProc. Speaker Lang. Recog- nit. Workshop (Odyssey), 2020, pp. 202–209
2020
-
[72]
The DET curve in assessment of detection task performance,
A. Martin, G. Doddington, T. Kamm, M. Ordowski, and M. Przybocki, “The DET curve in assessment of detection task performance,” inProc. Eurospeech, 1997, pp. 1895–1898
1997
-
[73]
Robust speech recognition via large-scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” inProc. Int. Conf. Mach. Learn. (ICML), 2023, arXiv:2212.04356
Pith/arXiv arXiv 2023
-
[74]
faster-whisper: Faster Whisper transcription with CTrans- late2,
SYSTRAN, “faster-whisper: Faster Whisper transcription with CTrans- late2,” https://github.com/SYSTRAN/faster-whisper, 2023, cTranslate2 reimplementation of OpenAI Whisper; accessed 2026-07
2023
-
[75]
AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,
J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.- J. Yu, and N. Evans, “AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,” inProc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP), 2022, pp. 6367–6371, arXiv:2110.01200
Pith/arXiv arXiv 2022
-
[76]
Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation,
H. Tak, M. Todisco, X. Wang, J.-w. Jung, J. Yamagishi, and N. Evans, “Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation,” inProc. Speaker Lang. Recognit. Workshop (Odyssey), 2022, pp. 112–119
2022
-
[77]
V oice conversion with just nearest neighbors,
M. Baas, B. van Niekerk, and H. Kamper, “V oice conversion with just nearest neighbors,” inProc. INTERSPEECH, 2023, pp. 2053–2057
2023
-
[78]
Resemblyzer: Analyze and compare voices with deep learning,
Resemble AI, “Resemblyzer: Analyze and compare voices with deep learning,” https://github.com/resemble-ai/Resemblyzer, 2019, imple- ments the GE2E speaker encoder of Wan et al.; accessed 2026-07
2019
-
[79]
Generalized end-to- end loss for speaker verification,
L. Wan, Q. Wang, A. Papir, and I. Lopez Moreno, “Generalized end-to- end loss for speaker verification,” inProc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP), 2018, pp. 4879–4883
2018
-
[80]
The T05 system for the V oiceMOS Challenge 2024: Transfer learning from deep image classifier to naturalness MOS prediction of high-quality synthetic speech,
K. Baba, W. Nakata, Y . Saito, and H. Saruwatari, “The T05 system for the V oiceMOS Challenge 2024: Transfer learning from deep image classifier to naturalness MOS prediction of high-quality synthetic speech,” inProc. IEEE Spoken Lang. Technol. Workshop (SLT), 2024, uTMOSv2
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.