REVIEW 2 major objections 4 minor 28 references
Source tracing works better when a synthetic voice is treated as architecture plus training data plus residual, not as one monolithic class.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 04:38 UTC pith:AEERA6KY
load-bearing objection Solid, modest methodological step for compositional open-set source tracing; the residual-subspace story is under-isolated but the empirical gains are real and cleanly reported. the 2 major comments →
Open-Set Source Tracing as Compositional Factors via Structured Prototypes
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Redefining a synthetic-speech source as the tuple (architecture, training data, residual training factors) and aligning the embedding to structured orthonormal prototypes in factorized subspaces yields higher few-shot open-set attribution accuracy and a smaller generalization gap than monolithic angular-margin learning, because novel combinations can be assembled from the already-learned factor bases.
What carries the argument
Subspace-partitioned structured orthonormal prototypes: the embedding is split into fixed architecture and data subspaces whose bases are orthonormal prototypes built from metadata labels, plus a residual subspace whose energy is softly constrained so that it absorbs unlabeled interactions without collapsing or dominating the factor subspaces.
Load-bearing premise
The architecture and training-data labels supplied by the dataset are complete and independent enough that fixed prototypes can be aligned to them and everything else can safely live in a residual subspace.
What would settle it
Train the same subspace-partitioned model on a corpus whose architecture and data labels are deliberately corrupted or entangled; if few-shot open-set F1-macro then falls to or below the ArcFace baseline, the compositional claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper redefines synthetic-speech source tracing as attribution of a compositional tuple S=(A,D,H) rather than a monolithic architecture label. It proposes Structured Orthonormal Prototypes and a Subspace Partitioning strategy that splits the embedding into architecture (Z_A), training-data (Z_D) and residual (Z_R) subspaces, with fixed orthonormal bases for the labeled factors and an energy constraint on Z_R (Eqs. 2–4). On MLAAD, under a few-shot open-set protocol, the residual-modeling variant (Strategy 2.B) improves All-vs-All F1-macro over ArcFace (84.52 % vs 80.11 % at K=5) and reduces the seen–unseen generalization gap, with a five-way OOD breakdown (Seen / Unseen Combination / Unseen Dataset / Unseen Arch / Both Unseen) reported in Table III and Figure 3.
Significance. If the claimed compositional mechanism holds, the work supplies a concrete geometric recipe for open-set source tracing that scales more favorably than monolithic prototypes (C_A imes C_D sources from only C_A + C_D bases) and offers a path toward factor-level forensic interpretability. The multi-seed means/stds, explicit closed-set vs few-shot protocol, and five-way OOD decomposition are strengths that make the empirical claims falsifiable. The absolute gains remain modest, so the primary contribution is the problem reformulation and the structured-prototype framework rather than a decisive performance leap.
major comments (2)
- The central claim that subspace partitioning enables compositional generalization rests on Z_A and Z_D remaining pure while Z_R absorbs residual H and unlabeled variability (Eqs. 3–4, §II-C). Figure 3 shows the largest relative lift for Strategy 2.B precisely on Unseen Combination and Unseen Architecture. Yet nothing forces the residual energy constraint (µ_ref = 1) to leave the factor subspaces pure: non-linear A×D interactions or unrecorded factors can migrate into Z_R, after which nearest-prototype matching on the full embedding can succeed without true compositional reuse of the marginal prototypes. No ablation freezes or zeros Z_R at test time, nor reports the fraction of cosine similarity for novel (a,d) pairs carried by residual versus factor coordinates. Without that isolation the modest All-vs-All gain is consistent with ordinary prototype regularization plus an extra free subsp
- The weakest modeling assumption is that MLAAD’s architecture and training-data metadata are sufficiently complete and independent labels of the true generative factors (Table II, §II-B, §III-A). Sources with undocumented training data receive independent “unknown” labels, and speaker labels are discarded as inconsistent. If those metadata are incomplete, noisy or entangled with unrecorded factors, the fixed bases U and V cannot be cleanly aligned and the residual subspace cannot be guaranteed to absorb only the intended remainder. A sensitivity analysis (e.g., randomly permuting a fraction of factor labels, or treating “unknown” as a single shared class) would quantify how much of the reported OOD lift depends on metadata fidelity.
minor comments (4)
- Table III reports means ± std over 4 seeds and 100 few-shot trials, but no statistical significance test (paired t-test or bootstrap CI) is given for the 4.4 pp All-vs-All gap between Strategy 2.B and ArcFace; a short significance statement would strengthen the claim.
- Figure 3 caption states “Best-performing configuration selected for each strategy based on validation performance,” yet the main text never lists which λ was chosen for each bar; adding that information would aid reproducibility.
- Notation for the residual energy target is written µ_ref = 1/2 (∥u_a∥^{2} + ∥v_d∥^{2}) = 1; the intermediate equality is redundant once unit-norm prototypes are assumed and could be simplified.
- The abstract claims the approach “significantly outperforms angular-margin baselines”; given the absolute margins, “consistently outperforms” would be more precise.
Circularity Check
No circularity: empirical few-shot F1 gains on held-out MLAAD sources do not reduce by construction to training inputs or self-citations.
full rationale
The paper’s load-bearing claims are empirical comparisons (Table III, Figure 3) of F1-macro under closed-set and few-shot open-set protocols on MLAAD, with architectures/datasets held out of training (Table II). Prototypes for seen factors are fixed orthonormal bases built from training metadata; dynamic prototypes for evaluation sources are averages of K support embeddings computed at test time. The losses (Eqs. 1–4) are ordinary training objectives; none of the reported F1 numbers is algebraically forced by a fitted parameter or by the definition of the subspaces. The self-citation to Almudévar et al. (Interspeech 2024) supplies inspiration for predefined prototypes and is not used as a uniqueness theorem or as the sole justification of the performance claims. Concerns that Z_R may absorb compositional signal rather than isolate it are mechanism/correctness questions, not circular reductions of prediction to input. The derivation chain is therefore self-contained against the external benchmark.
Axiom & Free-Parameter Ledger
free parameters (4)
- MSE regularization strength λ
- Subspace dimensionalities (M_A, M_D, M_R)
- ArcFace scale s and margin m
- Residual reference magnitude μ_ref
axioms (4)
- domain assumption A generative source can be adequately represented as the compositional tuple (A, D, H) whose labeled factors are recoverable from MLAAD metadata.
- ad hoc to paper Frozen orthonormal bases for architecture and data factors minimize class overlap while permitting shared structure for sources that reuse a factor.
- domain assumption Inference-time variables (speaker, prompt, phonetic content) and unlabeled training stochasticity can be absorbed into a residual subspace without destroying factor discriminability.
- standard math Cosine similarity to averaged support embeddings is a valid few-shot open-set decision rule.
invented entities (2)
-
Residual subspace Z_R with energy constraint
no independent evidence
-
Composite factorized prototype p_s = u_a ⊕ v_d
no independent evidence
read the original abstract
Recent research expands beyond binary anti-spoofing with the emergence of Source Tracing, the task of identifying the specific generative origins of synthetic speech. However, current research often equates a "source" with its generative architecture. We propose redefining a source as a compositional tuple of Architecture, Training Data, and other training factors affecting the generated speech. We propose a framework using Structured Orthonormal Prototypes to minimize class overlap and intra-class variance. Our Subspace Partitioning strategy splits the embedding into architecture and data subspaces, while a residual subspace captures stochastic variability, enabling "compositional generalization" for novel factor combinations. This approach improves performance for partially seen sources and maintains robustness in fully open-set scenarios. MLAAD evaluations for Few-Shot open-set Identification show our approach significantly outperforms angular-margin baselines.
Figures
Reference graph
Works this paper leans on
-
[1]
Asvspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge,
Z. Wu, T. Kinnunen, N. Evans, J. Yamagishi, C. Hanilc ¸i, M. Sahidullah, and A. Sizov, “Asvspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge,” inInterspeech 2015, 2015, pp. 2037–2041
2015
-
[2]
The asvspoof 2017 challenge: Assessing the limits of replay spoofing attack detection,
T. Kinnunen, M. Sahidullah, H. Delgado, M. Todisco, N. Evans, J. Ya- magishi, and K. A. Lee, “The asvspoof 2017 challenge: Assessing the limits of replay spoofing attack detection,” inInterspeech 2017, 2017, pp. 2–6
2017
-
[3]
Asvspoof 2019: Future horizons in spoofed and fake audio detection,
M. Todisco, X. Wang, V . Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. Evans, T. H. Kinnunen, and K. A. Lee, “Asvspoof 2019: Future horizons in spoofed and fake audio detection,” inInterspeech 2019, 2019, pp. 1008–1012
2019
-
[4]
Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild,
X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch, and K. A. Lee, “Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 2507–2522, 2023
2021
-
[5]
Asvspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,
X. Wang, H. Delgado, H. Tak, J. weon Jung, H. jin Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. Kinnunen, N. Evans, K. A. Lee, and J. Yamagishi, “Asvspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,” 2024. [Online]. Available: https://arxiv.org/abs/2408.08739
Pith/arXiv arXiv 2024
-
[6]
Add 2022: the first audio deep synthesis detection challenge,
J. Yi, R. Fu, J. Tao, S. Nie, H. Ma, C. Wang, T. Wang, Z. Tian, Y . Bai, C. Fan, S. Liang, S. Wang, S. Zhang, X. Yan, L. Xu, Z. Wen, H. Li, Z. Lian, and B. Liu, “Add 2022: the first audio deep synthesis detection challenge,” inICASSP 2022 – 2022 IEEE International Conference on Acoustics, Speech and Signal Processing, 2022, pp. 9216–9220
2022
-
[7]
Add 2023: the second audio deepfake detection challenge,
J. Yi, J. Tao, R. Fu, X. Yan, C. Wang, T. Wang, C. Y . Zhang, X. Zhang, Y . Zhao, Y . Ren, L. Xu, J. Zhou, H. Gu, Z. Wen, S. Liang, Z. Lian, S. Nie, and H. Li, “Add 2023: the second audio deepfake detection challenge,” inDADA@IJCAI 2023 (Second Audio Deepfake Detection Challenge Workshop), 2023, pp. 125–130
2023
-
[8]
Source Tracing of Audio Deepfake Systems,
N. Klein, T. Chen, H. Tak, R. Casal, and E. Khoury, “Source Tracing of Audio Deepfake Systems,” inInterspeech 2024, 2024, pp. 1100–1104
2024
-
[9]
Mlaad: The multi-language audio anti-spoofing dataset,
N. M. M ¨uller, P. Kawa, W. H. Choong, E. Casanova, E. G¨olge, T. M¨uller, P. Syga, P. Sperl, and K. B ¨ottinger, “Mlaad: The multi-language audio anti-spoofing dataset,” in2024 International Joint Conference on Neural Networks (IJCNN), 2024, pp. 1–7
2024
-
[10]
STOPA: A Dataset of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution,
A. Firc, M. Chhibber, J. Mishra, V . Pratap Singh, T. Kinnunen, and K. Malinka, “STOPA: A Dataset of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution,” inInterspeech 2025, 2025, pp. 1553–1557
2025
-
[11]
Codecfake+: A large-scale neural audio codec-based deepfake speech dataset,
X. Chen, J. Du, H. Wu, L. Zhang, I. Lin, I. Chiu, W. Ren, Y . Tseng, Y . Tsao, J.-S. R. Janget al., “Codecfake+: A large-scale neural audio codec-based deepfake speech dataset,”arXiv preprint arXiv:2501.08238, 2025
Pith/arXiv arXiv 2025
-
[12]
Audio Deepfake Source Tracing using Multi-Attribute Open-Set Identification and Verification,
P. Falez, T. Marteau, D. Lolive, and A. Delhay, “Audio Deepfake Source Tracing using Multi-Attribute Open-Set Identification and Verification,” inInterspeech 2025, 2025, pp. 1528–1532
2025
-
[13]
Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy,
X. Chen, I.-M. Lin, L. Zhang, J. Du, H. Wu, H. yi Lee, and J.-S. R. Jang, “Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy,” inInterspeech 2025, 2025, pp. 1538–1542
2025
-
[14]
Advancing Zero-Shot Open-Set Speech Deepfake Source Tracing,
M. Chhibber, J. Mishra, and T. H. Kinnunen, “Advancing Zero-Shot Open-Set Speech Deepfake Source Tracing,” inOdyssey 2026, 2026, pp. 185–190
2026
-
[15]
Source Verification for Speech Deepfakes ,
V . Negroni, D. Salvi, P. Bestagini, and S. Tubaro, “ Source Verification for Speech Deepfakes ,” inInterspeech 2025, 2025, pp. 1548–1552
2025
-
[16]
Open-Set Source Tracing of Audio Deepfake Systems,
N. Klein, H. Tak, and E. Khoury, “Open-Set Source Tracing of Audio Deepfake Systems,” inInterspeech 2025, 2025, pp. 1578–1582
2025
-
[17]
How people make their own environments: A theory of genotype→environment effects,
S. Scarr and K. McCartney, “How people make their own environments: A theory of genotype→environment effects,”Child Development, vol. 54, no. 2, pp. 424–435, 1983. [Online]. Available: http://www.jstor.org/stable/1129703
arXiv 1983
-
[18]
Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion,
A. Kulkarni, S. Dowerah, T. Alum ¨ae, and M. M. Doss, “Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion,” inInterspeech 2025, 2025, pp. 1533– 1537
2025
-
[19]
Syn- thetic Speech Source Tracing using Metric Learning,
D. Koutsianos, S. Zacharopoulos, Y . Panagakis, and T. Stafylakis, “Syn- thetic Speech Source Tracing using Metric Learning,” inInterspeech 2025, 2025, pp. 1558–1562
2025
-
[20]
Predefined Prototypes for Intra-Class Separation and Disentanglement,
A. Almud ´evar, T. Mariotte, A. Ortega, M. Tahon, L. Vicente, A. Miguel, and E. Lleida, “Predefined Prototypes for Intra-Class Separation and Disentanglement,” inInterspeech 2024, 2024, pp. 3809–3813
2024
-
[21]
An interpretable deep learning model for automatic sound classification,
P. Zinemanas, M. Rocamora, M. Miron, F. Font, and X. Serra, “An interpretable deep learning model for automatic sound classification,”Electronics, vol. 10, no. 7, 2021. [Online]. Available: https://www.mdpi.com/2079-9292/10/7/850
2021
-
[22]
Assessing the impact of speaker identity in speech spoofing detection,
A.-T. Daoet al., “Assessing the impact of speaker identity in speech spoofing detection,” inICASSP 2026 - IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2026
2026
-
[23]
Arcface: Additive angular margin loss for deep face recognition,
J. Deng, J. Guo, J. Yang, N. Xue, I. Kotsia, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 10, p. 5962–5979, Oct. 2022. [Online]. Available: http://dx.doi.org/10.1109/TPAMI.2021.3087709
-
[24]
Extensions of lipschitz mappings into a hilbert space,
W. B. Johnson, J. Lindenstrausset al., “Extensions of lipschitz mappings into a hilbert space,”Contemporary mathematics, vol. 26, no. 189-206, p. 1, 1984
1984
-
[25]
Ring loss: Convex feature nor- malization for face recognition,
Y . Zheng, D. K. Pal, and M. Savvides, “Ring loss: Convex feature nor- malization for face recognition,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018
2018
-
[26]
Vicreg: Variance-invariance- covariance regularization for self-supervised learning,
A. Bardes, J. Ponce, and Y . LeCun, “Vicreg: Variance-invariance- covariance regularization for self-supervised learning,” inInternational Conference on Learning Representations (ICLR), 2022
2022
-
[27]
Mls: A large-scale multilingual dataset for speech research,
V . Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “Mls: A large-scale multilingual dataset for speech research,”ArXiv, vol. abs/2012.03411, 2020
Pith/arXiv arXiv 2012
-
[28]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” 2015. [Online]. Available: https://arxiv.org/abs/1512.03385
Pith/arXiv arXiv 2015
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.