REVIEW 2 major objections 2 minor 50 references
SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization
T0 review · 2 major / 2 minor · reviewed 2026-06-25 · grok-4.3
Pith's one-line read SphereVBx replaces PLDA with T-PSDA to cluster hyperspherical embeddings for simpler and accurate speaker diarization.
desk verdict SphereVBx swaps T-PSDA for PLDA inside the VBx variational setup to cluster hyperspherical EEND-VC embeddings, with a clean parameter-free cosine variant. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA), which enables variational Bayesian inference under a mixture of von Mises-Fisher distributions on the hypersphere.
What would settle it
Running SphereVBx on the same embeddings as standard VBx and finding lower diarization error rates or no improvement on the benchmarks would challenge the central performance claim.
Extended reading notes
Core claim
SphereVBx performs variational inference in a mixture of von Mises-Fisher distributions by replacing the Gaussian PLDA backend with T-PSDA in the VBx formulation, resulting in improved clustering accuracy in cascaded diarization and comparable or better performance in EEND-VC with a significantly simplified clustering stage.
Load-bearing premise
The embeddings produced by the neural front-end lie on a hypersphere and are well described by a mixture of von Mises-Fisher distributions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SphereVBx, a variational Bayesian clustering method for hyperspherical embeddings obtained by replacing the PLDA backend in the existing VBx formulation with Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA), yielding inference under a mixture of von Mises-Fisher distributions. The approach is applied to speaker diarization, with emphasis on the EEND-VC framework; a parameter-free variant (SphereVBx-PF) is introduced that aligns with cosine-like scoring. Experiments on multiple diarization benchmarks are reported to show improved clustering accuracy in cascaded pipelines and comparable or superior performance in EEND-VC while simplifying the clustering stage.
Significance. If the reported benchmark results hold under rigorous evaluation, the work provides a concrete simplification of the clustering component in EEND-VC pipelines without loss of accuracy. The explicit construction of the parameter-free SphereVBx-PF variant, which avoids pretrained backend parameters and ties directly to spherical geometry, is a clear strength that could reduce deployment complexity in practical diarization systems.
major comments (2)
- [§3] §3 (method derivation): the claim that T-PSDA substitution yields a mixture of von Mises-Fisher distributions is presented as a direct replacement, but the manuscript does not explicitly derive or bound the variational update equations under the toroidal model; without these steps it is unclear whether the inference remains tractable at the same computational cost as the original VBx.
- [§4] §4 (experiments): the abstract and results claim improvements or parity on multiple benchmarks, yet no table or section provides the raw DER values, confidence intervals, or statistical significance tests against the EEND-VC baseline with standard PLDA; this omission prevents assessment of whether the observed gains are load-bearing or within noise.
minor comments (2)
- The notation for the toroidal parameters in T-PSDA should be introduced with a short table or explicit mapping to the original PLDA parameters to aid readability.
- Figure captions for the embedding visualizations should state the exact dataset and front-end model used to generate the plotted points.
Simulated Author's Rebuttal
We thank the referee for the constructive review and the recommendation for minor revision. We address each major comment below.
read point-by-point responses
-
Referee: [§3] §3 (method derivation): the claim that T-PSDA substitution yields a mixture of von Mises-Fisher distributions is presented as a direct replacement, but the manuscript does not explicitly derive or bound the variational update equations under the toroidal model; without these steps it is unclear whether the inference remains tractable at the same computational cost as the original VBx.
Authors: We agree that an explicit derivation of the variational updates would improve clarity. In the revised manuscript we will add a dedicated subsection deriving the variational Bayesian updates for the vMF mixture under the T-PSDA parameterization. The derivation follows the standard mean-field variational inference steps for vMF distributions (with the toroidal model inducing the appropriate concentration and mean-direction updates), confirming that the per-iteration cost remains O(NK) and is therefore comparable to the original PLDA-based VBx. revision: yes
-
Referee: [§4] §4 (experiments): the abstract and results claim improvements or parity on multiple benchmarks, yet no table or section provides the raw DER values, confidence intervals, or statistical significance tests against the EEND-VC baseline with standard PLDA; this omission prevents assessment of whether the observed gains are load-bearing or within noise.
Authors: We acknowledge the value of raw numbers and statistical assessment. The revised manuscript will include an expanded results table reporting the full DER values for every system and condition, together with 95% confidence intervals obtained via bootstrap resampling and paired statistical significance tests (McNemar’s test on speaker error decisions) against the EEND-VC + PLDA baseline. revision: yes
Circularity Check
No significant circularity in derivation chain
full rationale
The paper presents SphereVBx as a direct substitution of T-PSDA for PLDA inside the established VBx variational Bayesian clustering formulation, yielding inference over von Mises-Fisher mixtures on hyperspherical embeddings. No equations, predictions, or first-principles results are shown that reduce by construction to fitted inputs, self-definitions, or load-bearing self-citations. The parameter-free SphereVBx-PF variant is explicitly tied to cosine-like scoring without hidden parameters. Empirical results on diarization benchmarks are reported independently of any internal redefinition, making the derivation self-contained against external evaluation.
Assumptions & free parameters
Cite this review
Pith. "Pith review of SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization." pith.science (2026). https://pith.science/paper/VLL7EUC7
@misc{pith2026260624528,
author = {Pith},
title = {Pith review of: SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization},
year = {2026},
howpublished = {\url{https://pith.science/paper/VLL7EUC7}},
note = {Machine review of arXiv:2606.24528}
}
read the original abstract
We propose SphereVBx, a Bayesian clustering framework for hyperspherical embeddings based on Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA). The method follows the variational Bayesian formulation of VBx while replacing the Gaussian Probabilistic Linear Discriminant Analysis (PLDA) backend with T-PSDA, resulting in variational inference in a mixture of von Mises-Fisher distributions. We apply SphereVBx to speaker diarization and in particular to the end-to-end neural diarization with vector clustering (EEND-VC) framework. A parameter-free variant, denoted SphereVBx-PF, corresponds to a spherical similarity model closely related to cosine scoring and does not require pretrained backend parameters. Experiments on multiple diarization benchmarks show that SphereVBx improves clustering accuracy in cascaded diarization pipelines and achieves comparable or better performance in the EEND-VC framework while significantly simplifying its clustering stage.
Figures
Reference graph
Works this paper leans on
-
[1]
This is common in tasks such as speaker recognition [1] and face recog- nition [2]
Introduction Many modern representation learning systems produce embed- dings that are normalized to lie on the unit hypersphere. This is common in tasks such as speaker recognition [1] and face recog- nition [2]. Although cosine similarity is widely used to com- pare such embeddings, probabilistic clustering methods that ex- plicitly model their hypersph...
-
[2]
SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization
Background 2.1. PLDA-based VBx VBx was introduced in [7] as a Bayesian framework for clus- tering a sequence of speaker embeddingsX= [x 1, . . . ,xT ] extracted from short, overlapping speech segments. VBx mod- elsXusing a Hidden Markov Model (HMM) in which each state corresponds to a speaker and transitions represent speaker changes. Conditioned on speak...
work page Pith review arXiv 2026
-
[3]
In particular, the PLDA within- and between-speaker distributions (1) and (2) used in VBx are re- placed by the corresponding T-PSDA distributions (3) and (4)
SphereVBx SphereVBx combines VBx (Section 2.1) with the T-PSDA model (Section 2.2). In particular, the PLDA within- and between-speaker distributions (1) and (2) used in VBx are re- placed by the corresponding T-PSDA distributions (3) and (4). The remaining VB inference framework of VBx is retained. As discussed in Section 2.2, whend=D,κ b = 0, and κw = 1...
-
[4]
cannot-link
Integration to EEND-VC EEND-VC combines (i) a local EEND model that estimates speaker activity within short, overlapping windows with (ii) a global clustering stage that links these local speaker instances across the full recording. For each window, we extract one speaker embedding per active local speaker and pass the set of embeddings to the clustering ...
-
[5]
For AMI and AliMeeting, we consistently use the first channel from the far-field microphone array
Experimental setup We evaluated on the following standard benchmark datasets: AMI [27], AISHELL-4 (AIS) [28], AliMeeting (AliM) [29], NOTSOFAR-1 (NSF) [30], MSDWild (MSD) [31], DIHARD3 full (DH3) [12], RAMC [32], and V oxConverse (V oxC) [33]. For AMI and AliMeeting, we consistently use the first channel from the far-field microphone array. The cascade VB...
2026
-
[6]
The model was trained on16s chunks of the audio using the powerset loss [36], which supports up to four speakers with ar- bitrary overlap
feature extractor followed by a Conformer encoder [35]. The model was trained on16s chunks of the audio using the powerset loss [36], which supports up to four speakers with ar- bitrary overlap. The pretrained model is kept fixed across all experiments and serves as the local EEND. For inference, we segment the audio into16s chunks with a shift of1.6s to ...
-
[7]
Results 6.1. Comparison of VBx vs SphereVBx Table 1 summarizes the performance of the cascade diarization pipeline (V AD + VBx + OSD) under different speaker models for VBx and evaluation protocols. The first row reports the VBx baseline from [24]. In that work, the hyperparametersF A and FB were tuned per dataset on its development partition. In the rema...
-
[8]
Conclusion We introduced SphereVBx, a general clustering method for nor- malized embedding vectors on the unit hypersphere using a T-PSDA speaker model within the VBx inference framework. SphereVBx and its parameter-free version allow a more prin- cipled and greatly simplified clustering pipeline by eliminating several heuristic steps required when applyi...
Show all 50 references
-
[9]
On our own: Oppor- tunities and Risks in the Individualization of Society (PRINS) CZ.02.01.01/00/23 025/0008710
Acknowledgments This work was supported by the project “On our own: Oppor- tunities and Risks in the Individualization of Society (PRINS) CZ.02.01.01/00/23 025/0008710”, which is co-financed by the European Union. Computing on the IT4I supercomputer was supported by the Czech ...
-
[10]
The scientific content, methodology, experiments, and conclusions were developed by the authors
Generative AI Use Disclosure Generative AI tools (e.g., ChatGPT) were used solely for lan- guage editing and paraphrasing to improve the clarity and read- ability of the manuscript. The scientific content, methodology, experiments, and conclusions were developed by the authors...
-
[11]
Margin mat- ters: Towards more discriminative deep neural network embed- dings for speaker recognition,
X. Xiang, S. Wang, H. Huang, Y . Qian, and K. Yu, “Margin mat- ters: Towards more discriminative deep neural network embed- dings for speaker recognition,” in2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Confer- ence (APSIPA ASC). IEEE, 2019...
2019
-
[12]
ArcFace: Additive Angular Margin Loss for Deep Face Recognition,
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “ArcFace: Additive Angular Margin Loss for Deep Face Recognition,” inProceed- ings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), June 2019
2019
-
[13]
Clustering on the unit hypersphere using von mises-fisher dis- tributions
A. Banerjee, I. S. Dhillon, J. Ghosh, S. Sra, and G. Ridgeway, “Clustering on the unit hypersphere using von mises-fisher dis- tributions.”Journal of Machine Learning Research, vol. 6, no. 9, 2005
2005
-
[14]
V on mises-fisher clustering models,
S. Gopal and Y . Yang, “V on mises-fisher clustering models,” in International conference on machine learning. PMLR, 2014, pp. 154–162
2014
-
[15]
Robust speaker clus- tering using mixtures of von mises-fisher distributions for natural- istic audio streams,
H. Dubey, A. Sangwan, and J. H. Hansen, “Robust speaker clus- tering using mixtures of von mises-fisher distributions for natural- istic audio streams,” inProc. Interspeech, 2018, pp. 3603–3607
2018
-
[16]
Proba- bilistic back-ends for online speaker recognition and clustering,
A. Sholokhov, N. Kuzmin, K. A. Lee, and E. S. Chng, “Proba- bilistic back-ends for online speaker recognition and clustering,” inProc. ICASSP, 2023
2023
-
[17]
Bayesian hmm based x-vector clustering for speaker diarization,
M. Diez, L. Burget, S. Wang, J. Rohdin, and J. ˇCernock´y, “Bayesian hmm based x-vector clustering for speaker diarization,” inInterspeech 2019, 2019, pp. 346–350
2019
-
[18]
Bayesian HMM Clustering of x-vector Sequences (VBx) in Speaker Diarization: Theory, Implementation and Analysis on Standard Tasks,
F. Landini, J. Profant, M. Diez, and L. Burget, “Bayesian HMM Clustering of x-vector Sequences (VBx) in Speaker Diarization: Theory, Implementation and Analysis on Standard Tasks,”Com- puter Speech & Language, vol. 71, 2022
2022
-
[19]
The speaker partitioning prob- lem
N. Br ¨ummer and E. De Villiers, “The speaker partitioning prob- lem.” inOdyssey, 2010, p. 34
2010
-
[20]
Probabilistic Spherical Discriminant Analy- sis: An Alternative to PLDA for length-normalized embeddings,
N. Brummer, A. Swart, L. Mosner, A. Silnova, O. Plchot, T. Stafy- lakis, and L. Burget, “Probabilistic Spherical Discriminant Analy- sis: An Alternative to PLDA for length-normalized embeddings,” inProc. Interspeech, 2022, pp. 1446–1450
2022
-
[21]
Toroidal probabilistic spherical discriminant analysis,
A. Silnova, N. Br ¨ummer, A. Swart, and L. Burget, “Toroidal probabilistic spherical discriminant analysis,” inProc. ICASSP. IEEE, 2023
2023
-
[22]
The third DI- HARD diarization challenge,
N. Ryant, P. Singh, V . Krishnamohan, R. Varma, K. Church, C. Cieri, J. Du, S. Ganapathy, and M. Liberman, “The third DI- HARD diarization challenge,” inProc. Interspeech, 2021, pp. 3570–3574
2021
-
[23]
pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,
H. Bredin, “pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,” inProc. Interspeech, 2023, pp. 1983–1987
2023
-
[24]
Integrating end-to- end neural and clustering-based diarization: Getting the best of both worlds,
K. Kinoshita, M. Delcroix, and N. Tawara, “Integrating end-to- end neural and clustering-based diarization: Getting the best of both worlds,” inProc. ICASSP. IEEE, 2021, pp. 7198–7202
2021
-
[25]
Advances in inte- gration of end-to-end neural and clustering-based diarization for real conversational speech,
K. Kinoshita, M. Delcroix, and N. Tawara “Advances in inte- gration of end-to-end neural and clustering-based diarization for real conversational speech,” inInterspeech 2021, 2021, pp. 3565– 3569
2021
-
[26]
VBx for end-to-end neural and clustering-based diarization,
P. P ´alka, J. Han, M. Delcroix, N. Tawara, and L. Burget, “VBx for end-to-end neural and clustering-based diarization,”arXiv preprint arXiv:2510.19572, 2025
2025
-
[27]
Dual-strategy- enhanced ConBiMamba for neural speaker diarization,
Z. Liao, G. Dai, M. Chen, W. Cheng, and W. Xu, “Dual-strategy- enhanced ConBiMamba for neural speaker diarization,”arXiv preprint arXiv:2601.19472, accepted by ICASSP, 2026
2026
-
[28]
Towards fully bayesian speaker recognition: integrating out the between-speaker covariance,
J. Villalba and N. Br ¨ummer, “Towards fully bayesian speaker recognition: integrating out the between-speaker covariance,” in Proc. Interspeech, 2011, pp. 505–508
2011
-
[29]
Discriminative Training of VBx Diariza- tion,
D. Klement, M. Diez, F. Landini, L. Burget, A. Silnova, M. Del- croix, and N. Tawara, “Discriminative Training of VBx Diariza- tion,” inProc. ICASSP. IEEE, 2024, pp. 11 871–11 875
2024
-
[30]
C. M. Bishop,Pattern Recognition and Machine Learning, 1st ed., ser. Information Science and Statistics. New York, NY: Springer, 2006
2006
-
[31]
CSL VBx Deriva- tions,
M. Diez, L. Burget, and L. Federico, “CSL VBx Deriva- tions,” Brno University of Technology, Tech. Rep., 2021. [Online]. Available: https://www.fit.vut.cz/person/mireia/public/ CSL VBHMM tech report.pdf
2021
-
[32]
Mardia and P
K. Mardia and P. Jupp,Directional Statistics, ser. Wiley Series in Probability and Statistics. Wiley, 2000. [Online]. Available: https://books.google.cz/books?id=zjPvAAAAMAAJ
2000
-
[33]
Multi-stream extension of variational bayesian HMM clustering (MS-VBx) for combined end-to-end and vector clustering-based diarization,
M. Delcroix, N. Tawara, M. Diezet al., “Multi-stream extension of variational bayesian HMM clustering (MS-VBx) for combined end-to-end and vector clustering-based diarization,” inProc. In- terspeech, 2023, pp. 3477–3481
2023
-
[34]
DiaPer: End-to-end neural diarization with perceiver-based attractors,
F. Landini, M. Diez, T. Stafylakis, and L. Burget, “DiaPer: End-to-end neural diarization with perceiver-based attractors,” IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing, 2024
2024
-
[35]
Efficient and generalizable speaker diarization via structured pruning of self-supervised models,
J. Han, P. P ´alka, M. Delcroix, F. Landini, J. Rohdin, J. Cernock´y, and L. Burget, “Efficient and generalizable speaker diarization via structured pruning of self-supervised models,”arXiv preprint arXiv:2506.18623, accepted by TASLP, 2026
2026
-
[36]
Pushing the limits of end- to-end diarization,
S. J. Broughton and L. Samarakoon, “Pushing the limits of end- to-end diarization,” inProc. Interspeech, 2025, pp. 5218–5222
2025
-
[37]
The AMI meeting corpus: A pre-announcement,
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V . Karaiskos, W. Kraaij, M. Kronenthalet al., “The AMI meeting corpus: A pre-announcement,” inInterna- tional workshop on machine learning for multimodal interaction. Springer, 2006, pp. 28–39
2006
-
[38]
AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,
Y . Fu, L. Cheng, S. Lv, Y . Jv, Y . Kong, Z. Chen, Y . Hu, L. Xie, J. Wu, H. Bu, X. Xu, J. Du, and J. Chen, “AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,” inProc. Inter- speech, 2021, pp. 3665–3669
2021
-
[39]
M2MeT: The ICASSP 2022 multi-channel multi-party meeting transcription challenge,
F. Yu, S. Zhang, Y . Fu, L. Xie, S. Zheng, Z. Du, W. Huang, P. Guo, Z. Yan, B. Maet al., “M2MeT: The ICASSP 2022 multi-channel multi-party meeting transcription challenge,” inProc. ICASSP. IEEE, 2022, pp. 6167–6171
2022
-
[40]
NOTSOFAR-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,
A. Vinnikov, A. Ivry, A. Hurvitz, I. Abramovski, S. Koubi, I. Gurvich, S. Peer, X. Xiao, B. M. Elizalde, N. Kandaet al., “NOTSOFAR-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,” inProc. Interspeech, 2024, pp. 5003–5007
2024
-
[41]
MSDWild: Multi- modal speaker diarization dataset in the wild,
T. Liu, S. Fan, X. Xiang, H. Song, S. Lin, J. Sun, T. Han, S. Chen, B. Yao, S. Liu, Y . Wu, Y . Qian, and K. Yu, “MSDWild: Multi- modal speaker diarization dataset in the wild,” inProc. Inter- speech, 2022, pp. 1476–1480
2022
-
[42]
Open source MagicData-RAMC: A rich annotated mandarin conversational (RAMC) speech dataset,
Z. Yang, Y . Chen, L. Luo, R. Yang, L. Ye, G. Cheng, J. Xu, Y . Jin, Q. Zhang, P. Zhang, L. Xie, and Y . Yan, “Open source MagicData-RAMC: A rich annotated mandarin conversational (RAMC) speech dataset,” inProc. Interspeech, 2022, pp. 1736– 1740
2022
-
[43]
Spot the conversation: Speaker diarisation in the wild,
J. S. Chung, J. Huh, A. Nagrani, T. Afouras, and A. Zisserman, “Spot the conversation: Speaker diarisation in the wild,” inProc. Interspeech, 2020, pp. 299–303
2020
-
[44]
WavLM: Large-scale self-supervised pre-training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selecte...
2022
-
[45]
Conformer: Convolution-augmented transformer for speech recognition,
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y . Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y . Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. Interspeech, 2020, pp. 5036–5040
2020
-
[46]
Powerset multi-class cross entropy loss for neural speaker diarization,
A. Plaquet and H. Bredin, “Powerset multi-class cross entropy loss for neural speaker diarization,” inProc. Interspeech, 2023, pp. 3222–3226
2023
-
[47]
Leveraging self-supervised learning for speaker diarization,
J. Han, F. Landini, J. Rohdin, A. Silnova, M. Diez, and L. Burget, “Leveraging self-supervised learning for speaker diarization,” in Proc. ICASSP. IEEE, 2025, pp. 1–5
2025
-
[48]
Wespeaker: A research and production oriented speaker embedding learning toolkit,
H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang, Y . Deng, and Y . Qian, “Wespeaker: A research and production oriented speaker embedding learning toolkit,” inProc. ICASSP. IEEE, 2023, pp. 1–5
2023
-
[49]
V oxCeleb2: Deep speaker recognition,
J. S. Chung, A. Nagrani, and A. Zisserman, “V oxCeleb2: Deep speaker recognition,” inProc. Interspeech, 2018, pp. 1086–1090
2018
-
[50]
Multichannel source separation and tracking with ransac and directional statistics,
J. Traa and P. Smaragdis, “Multichannel source separation and tracking with ransac and directional statistics,”IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, vol. 22, no. 12, pp. 2233–2243, 2014
2014
Reviewed June 25, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.