Pith. sign in

REVIEW 2 major objections 2 minor 50 references

SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization

T0 review · 2 major / 2 minor · reviewed 2026-06-25 · grok-4.3

Pith's one-line read SphereVBx replaces PLDA with T-PSDA to cluster hyperspherical embeddings for simpler and accurate speaker diarization.

desk verdict SphereVBx swaps T-PSDA for PLDA inside the VBx variational setup to cluster hyperspherical EEND-VC embeddings, with a clean parameter-free cosine variant. read the letter →

arxiv 2606.24528 v1 pith:VLL7EUC7 submitted 2026-06-23 eess.AS

classification eess.AS
keywords speakerdiarizationvariationalBayesvonMises-FisherdistributionhypersphericalembeddingsEEND-VCclusteringT-PSDABayesian
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes SphereVBx, a variational Bayesian clustering method designed for embeddings that reside on a hypersphere. It adapts the VBx framework by using Toroidal Probabilistic Spherical Discriminant Analysis instead of Gaussian PLDA, which leads to modeling the data as a mixture of von Mises-Fisher distributions. This approach is applied to speaker diarization tasks, including the EEND-VC framework. A parameter-free version avoids the need for pretrained parameters and relates closely to cosine scoring. If the method works as described, it offers a way to simplify the clustering stage while maintaining or improving accuracy across multiple benchmarks.

What carries the argument

Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA), which enables variational Bayesian inference under a mixture of von Mises-Fisher distributions on the hypersphere.

What would settle it

Running SphereVBx on the same embeddings as standard VBx and finding lower diarization error rates or no improvement on the benchmarks would challenge the central performance claim.

Watch

Extended reading notes

Core claim

SphereVBx performs variational inference in a mixture of von Mises-Fisher distributions by replacing the Gaussian PLDA backend with T-PSDA in the VBx formulation, resulting in improved clustering accuracy in cascaded diarization and comparable or better performance in EEND-VC with a significantly simplified clustering stage.

Load-bearing premise

The embeddings produced by the neural front-end lie on a hypersphere and are well described by a mixture of von Mises-Fisher distributions.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper proposes SphereVBx, a variational Bayesian clustering method for hyperspherical embeddings obtained by replacing the PLDA backend in the existing VBx formulation with Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA), yielding inference under a mixture of von Mises-Fisher distributions. The approach is applied to speaker diarization, with emphasis on the EEND-VC framework; a parameter-free variant (SphereVBx-PF) is introduced that aligns with cosine-like scoring. Experiments on multiple diarization benchmarks are reported to show improved clustering accuracy in cascaded pipelines and comparable or superior performance in EEND-VC while simplifying the clustering stage.

Significance. If the reported benchmark results hold under rigorous evaluation, the work provides a concrete simplification of the clustering component in EEND-VC pipelines without loss of accuracy. The explicit construction of the parameter-free SphereVBx-PF variant, which avoids pretrained backend parameters and ties directly to spherical geometry, is a clear strength that could reduce deployment complexity in practical diarization systems.

major comments (2)
  1. [§3] §3 (method derivation): the claim that T-PSDA substitution yields a mixture of von Mises-Fisher distributions is presented as a direct replacement, but the manuscript does not explicitly derive or bound the variational update equations under the toroidal model; without these steps it is unclear whether the inference remains tractable at the same computational cost as the original VBx.
  2. [§4] §4 (experiments): the abstract and results claim improvements or parity on multiple benchmarks, yet no table or section provides the raw DER values, confidence intervals, or statistical significance tests against the EEND-VC baseline with standard PLDA; this omission prevents assessment of whether the observed gains are load-bearing or within noise.
minor comments (2)
  1. The notation for the toroidal parameters in T-PSDA should be introduced with a short table or explicit mapping to the original PLDA parameters to aid readability.
  2. Figure captions for the embedding visualizations should state the exact dataset and front-end model used to generate the plotted points.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive review and the recommendation for minor revision. We address each major comment below.

read point-by-point responses
  1. Referee: [§3] §3 (method derivation): the claim that T-PSDA substitution yields a mixture of von Mises-Fisher distributions is presented as a direct replacement, but the manuscript does not explicitly derive or bound the variational update equations under the toroidal model; without these steps it is unclear whether the inference remains tractable at the same computational cost as the original VBx.

    Authors: We agree that an explicit derivation of the variational updates would improve clarity. In the revised manuscript we will add a dedicated subsection deriving the variational Bayesian updates for the vMF mixture under the T-PSDA parameterization. The derivation follows the standard mean-field variational inference steps for vMF distributions (with the toroidal model inducing the appropriate concentration and mean-direction updates), confirming that the per-iteration cost remains O(NK) and is therefore comparable to the original PLDA-based VBx. revision: yes

  2. Referee: [§4] §4 (experiments): the abstract and results claim improvements or parity on multiple benchmarks, yet no table or section provides the raw DER values, confidence intervals, or statistical significance tests against the EEND-VC baseline with standard PLDA; this omission prevents assessment of whether the observed gains are load-bearing or within noise.

    Authors: We acknowledge the value of raw numbers and statistical assessment. The revised manuscript will include an expanded results table reporting the full DER values for every system and condition, together with 95% confidence intervals obtained via bootstrap resampling and paired statistical significance tests (McNemar’s test on speaker error decisions) against the EEND-VC + PLDA baseline. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity in derivation chain

full rationale

The paper presents SphereVBx as a direct substitution of T-PSDA for PLDA inside the established VBx variational Bayesian clustering formulation, yielding inference over von Mises-Fisher mixtures on hyperspherical embeddings. No equations, predictions, or first-principles results are shown that reduce by construction to fitted inputs, self-definitions, or load-bearing self-citations. The parameter-free SphereVBx-PF variant is explicitly tied to cosine-like scoring without hidden parameters. Empirical results on diarization benchmarks are reported independently of any internal redefinition, making the derivation self-contained against external evaluation.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract-only review supplies no explicit free parameters, axioms, or invented entities; the central modeling choice (von Mises-Fisher mixture via T-PSDA) is treated as an imported modeling assumption whose justification is not visible.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization." pith.science (2026). https://pith.science/paper/VLL7EUC7

@misc{pith2026260624528,
  author       = {Pith},
  title        = {Pith review of: SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VLL7EUC7}},
  note         = {Machine review of arXiv:2606.24528}
}
read the original abstract

We propose SphereVBx, a Bayesian clustering framework for hyperspherical embeddings based on Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA). The method follows the variational Bayesian formulation of VBx while replacing the Gaussian Probabilistic Linear Discriminant Analysis (PLDA) backend with T-PSDA, resulting in variational inference in a mixture of von Mises-Fisher distributions. We apply SphereVBx to speaker diarization and in particular to the end-to-end neural diarization with vector clustering (EEND-VC) framework. A parameter-free variant, denoted SphereVBx-PF, corresponds to a spherical similarity model closely related to cosine scoring and does not require pretrained backend parameters. Experiments on multiple diarization benchmarks show that SphereVBx improves clustering accuracy in cascaded diarization pipelines and achieves comparable or better performance in the EEND-VC framework while significantly simplifying its clustering stage.

Figures

Figures reproduced from arXiv: 2606.24528 by the authors.

Figure 1
Figure 1. Second stage of the EEND-VC pipeline. Dashed ar￾rows represent labels. X denotes speaker embeddings. initial clustering that is converted into hard responsibilities (6) for the first VB iteration. In the DiariZen baseline (Fig. 1a) [16], embeddings extracted from short segments (< 1.6 s) are filtered out because they tend to be unreliable and can degrade cluster￾ing performance. The remaining embeddings are processe… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

50 extracted references · 4 canonical work pages

  1. [1]

    This is common in tasks such as speaker recognition [1] and face recog- nition [2]

    Introduction Many modern representation learning systems produce embed- dings that are normalized to lie on the unit hypersphere. This is common in tasks such as speaker recognition [1] and face recog- nition [2]. Although cosine similarity is widely used to com- pare such embeddings, probabilistic clustering methods that ex- plicitly model their hypersph...

  2. [2]

    SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization

    Background 2.1. PLDA-based VBx VBx was introduced in [7] as a Bayesian framework for clus- tering a sequence of speaker embeddingsX= [x 1, . . . ,xT ] extracted from short, overlapping speech segments. VBx mod- elsXusing a Hidden Markov Model (HMM) in which each state corresponds to a speaker and transitions represent speaker changes. Conditioned on speak...

  3. [3]

    In particular, the PLDA within- and between-speaker distributions (1) and (2) used in VBx are re- placed by the corresponding T-PSDA distributions (3) and (4)

    SphereVBx SphereVBx combines VBx (Section 2.1) with the T-PSDA model (Section 2.2). In particular, the PLDA within- and between-speaker distributions (1) and (2) used in VBx are re- placed by the corresponding T-PSDA distributions (3) and (4). The remaining VB inference framework of VBx is retained. As discussed in Section 2.2, whend=D,κ b = 0, and κw = 1...

  4. [4]

    cannot-link

    Integration to EEND-VC EEND-VC combines (i) a local EEND model that estimates speaker activity within short, overlapping windows with (ii) a global clustering stage that links these local speaker instances across the full recording. For each window, we extract one speaker embedding per active local speaker and pass the set of embeddings to the clustering ...

  5. [5]

    For AMI and AliMeeting, we consistently use the first channel from the far-field microphone array

    Experimental setup We evaluated on the following standard benchmark datasets: AMI [27], AISHELL-4 (AIS) [28], AliMeeting (AliM) [29], NOTSOFAR-1 (NSF) [30], MSDWild (MSD) [31], DIHARD3 full (DH3) [12], RAMC [32], and V oxConverse (V oxC) [33]. For AMI and AliMeeting, we consistently use the first channel from the far-field microphone array. The cascade VB...

  6. [6]

    The model was trained on16s chunks of the audio using the powerset loss [36], which supports up to four speakers with ar- bitrary overlap

    feature extractor followed by a Conformer encoder [35]. The model was trained on16s chunks of the audio using the powerset loss [36], which supports up to four speakers with ar- bitrary overlap. The pretrained model is kept fixed across all experiments and serves as the local EEND. For inference, we segment the audio into16s chunks with a shift of1.6s to ...

  7. [7]

    Results 6.1. Comparison of VBx vs SphereVBx Table 1 summarizes the performance of the cascade diarization pipeline (V AD + VBx + OSD) under different speaker models for VBx and evaluation protocols. The first row reports the VBx baseline from [24]. In that work, the hyperparametersF A and FB were tuned per dataset on its development partition. In the rema...

  8. [8]

    Conclusion We introduced SphereVBx, a general clustering method for nor- malized embedding vectors on the unit hypersphere using a T-PSDA speaker model within the VBx inference framework. SphereVBx and its parameter-free version allow a more prin- cipled and greatly simplified clustering pipeline by eliminating several heuristic steps required when applyi...

Show all 50 references
  1. [9]

    On our own: Oppor- tunities and Risks in the Individualization of Society (PRINS) CZ.02.01.01/00/23 025/0008710

    Acknowledgments This work was supported by the project “On our own: Oppor- tunities and Risks in the Individualization of Society (PRINS) CZ.02.01.01/00/23 025/0008710”, which is co-financed by the European Union. Computing on the IT4I supercomputer was supported by the Czech ...

  2. [10]

    The scientific content, methodology, experiments, and conclusions were developed by the authors

    Generative AI Use Disclosure Generative AI tools (e.g., ChatGPT) were used solely for lan- guage editing and paraphrasing to improve the clarity and read- ability of the manuscript. The scientific content, methodology, experiments, and conclusions were developed by the authors...

  3. [11]

    Margin mat- ters: Towards more discriminative deep neural network embed- dings for speaker recognition,

    X. Xiang, S. Wang, H. Huang, Y . Qian, and K. Yu, “Margin mat- ters: Towards more discriminative deep neural network embed- dings for speaker recognition,” in2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Confer- ence (APSIPA ASC). IEEE, 2019...

  4. [12]

    ArcFace: Additive Angular Margin Loss for Deep Face Recognition,

    J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “ArcFace: Additive Angular Margin Loss for Deep Face Recognition,” inProceed- ings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), June 2019

  5. [13]

    Clustering on the unit hypersphere using von mises-fisher dis- tributions

    A. Banerjee, I. S. Dhillon, J. Ghosh, S. Sra, and G. Ridgeway, “Clustering on the unit hypersphere using von mises-fisher dis- tributions.”Journal of Machine Learning Research, vol. 6, no. 9, 2005

  6. [14]

    V on mises-fisher clustering models,

    S. Gopal and Y . Yang, “V on mises-fisher clustering models,” in International conference on machine learning. PMLR, 2014, pp. 154–162

  7. [15]

    Robust speaker clus- tering using mixtures of von mises-fisher distributions for natural- istic audio streams,

    H. Dubey, A. Sangwan, and J. H. Hansen, “Robust speaker clus- tering using mixtures of von mises-fisher distributions for natural- istic audio streams,” inProc. Interspeech, 2018, pp. 3603–3607

  8. [16]

    Proba- bilistic back-ends for online speaker recognition and clustering,

    A. Sholokhov, N. Kuzmin, K. A. Lee, and E. S. Chng, “Proba- bilistic back-ends for online speaker recognition and clustering,” inProc. ICASSP, 2023

  9. [17]

    Bayesian hmm based x-vector clustering for speaker diarization,

    M. Diez, L. Burget, S. Wang, J. Rohdin, and J. ˇCernock´y, “Bayesian hmm based x-vector clustering for speaker diarization,” inInterspeech 2019, 2019, pp. 346–350

  10. [18]

    Bayesian HMM Clustering of x-vector Sequences (VBx) in Speaker Diarization: Theory, Implementation and Analysis on Standard Tasks,

    F. Landini, J. Profant, M. Diez, and L. Burget, “Bayesian HMM Clustering of x-vector Sequences (VBx) in Speaker Diarization: Theory, Implementation and Analysis on Standard Tasks,”Com- puter Speech & Language, vol. 71, 2022

  11. [19]

    The speaker partitioning prob- lem

    N. Br ¨ummer and E. De Villiers, “The speaker partitioning prob- lem.” inOdyssey, 2010, p. 34

  12. [20]

    Probabilistic Spherical Discriminant Analy- sis: An Alternative to PLDA for length-normalized embeddings,

    N. Brummer, A. Swart, L. Mosner, A. Silnova, O. Plchot, T. Stafy- lakis, and L. Burget, “Probabilistic Spherical Discriminant Analy- sis: An Alternative to PLDA for length-normalized embeddings,” inProc. Interspeech, 2022, pp. 1446–1450

  13. [21]

    Toroidal probabilistic spherical discriminant analysis,

    A. Silnova, N. Br ¨ummer, A. Swart, and L. Burget, “Toroidal probabilistic spherical discriminant analysis,” inProc. ICASSP. IEEE, 2023

  14. [22]

    The third DI- HARD diarization challenge,

    N. Ryant, P. Singh, V . Krishnamohan, R. Varma, K. Church, C. Cieri, J. Du, S. Ganapathy, and M. Liberman, “The third DI- HARD diarization challenge,” inProc. Interspeech, 2021, pp. 3570–3574

  15. [23]

    pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,

    H. Bredin, “pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,” inProc. Interspeech, 2023, pp. 1983–1987

  16. [24]

    Integrating end-to- end neural and clustering-based diarization: Getting the best of both worlds,

    K. Kinoshita, M. Delcroix, and N. Tawara, “Integrating end-to- end neural and clustering-based diarization: Getting the best of both worlds,” inProc. ICASSP. IEEE, 2021, pp. 7198–7202

  17. [25]

    Advances in inte- gration of end-to-end neural and clustering-based diarization for real conversational speech,

    K. Kinoshita, M. Delcroix, and N. Tawara “Advances in inte- gration of end-to-end neural and clustering-based diarization for real conversational speech,” inInterspeech 2021, 2021, pp. 3565– 3569

  18. [26]

    VBx for end-to-end neural and clustering-based diarization,

    P. P ´alka, J. Han, M. Delcroix, N. Tawara, and L. Burget, “VBx for end-to-end neural and clustering-based diarization,”arXiv preprint arXiv:2510.19572, 2025

  19. [27]

    Dual-strategy- enhanced ConBiMamba for neural speaker diarization,

    Z. Liao, G. Dai, M. Chen, W. Cheng, and W. Xu, “Dual-strategy- enhanced ConBiMamba for neural speaker diarization,”arXiv preprint arXiv:2601.19472, accepted by ICASSP, 2026

  20. [28]

    Towards fully bayesian speaker recognition: integrating out the between-speaker covariance,

    J. Villalba and N. Br ¨ummer, “Towards fully bayesian speaker recognition: integrating out the between-speaker covariance,” in Proc. Interspeech, 2011, pp. 505–508

  21. [29]

    Discriminative Training of VBx Diariza- tion,

    D. Klement, M. Diez, F. Landini, L. Burget, A. Silnova, M. Del- croix, and N. Tawara, “Discriminative Training of VBx Diariza- tion,” inProc. ICASSP. IEEE, 2024, pp. 11 871–11 875

  22. [30]

    C. M. Bishop,Pattern Recognition and Machine Learning, 1st ed., ser. Information Science and Statistics. New York, NY: Springer, 2006

  23. [31]

    CSL VBx Deriva- tions,

    M. Diez, L. Burget, and L. Federico, “CSL VBx Deriva- tions,” Brno University of Technology, Tech. Rep., 2021. [Online]. Available: https://www.fit.vut.cz/person/mireia/public/ CSL VBHMM tech report.pdf

  24. [32]

    Mardia and P

    K. Mardia and P. Jupp,Directional Statistics, ser. Wiley Series in Probability and Statistics. Wiley, 2000. [Online]. Available: https://books.google.cz/books?id=zjPvAAAAMAAJ

  25. [33]

    Multi-stream extension of variational bayesian HMM clustering (MS-VBx) for combined end-to-end and vector clustering-based diarization,

    M. Delcroix, N. Tawara, M. Diezet al., “Multi-stream extension of variational bayesian HMM clustering (MS-VBx) for combined end-to-end and vector clustering-based diarization,” inProc. In- terspeech, 2023, pp. 3477–3481

  26. [34]

    DiaPer: End-to-end neural diarization with perceiver-based attractors,

    F. Landini, M. Diez, T. Stafylakis, and L. Burget, “DiaPer: End-to-end neural diarization with perceiver-based attractors,” IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing, 2024

  27. [35]

    Efficient and generalizable speaker diarization via structured pruning of self-supervised models,

    J. Han, P. P ´alka, M. Delcroix, F. Landini, J. Rohdin, J. Cernock´y, and L. Burget, “Efficient and generalizable speaker diarization via structured pruning of self-supervised models,”arXiv preprint arXiv:2506.18623, accepted by TASLP, 2026

  28. [36]

    Pushing the limits of end- to-end diarization,

    S. J. Broughton and L. Samarakoon, “Pushing the limits of end- to-end diarization,” inProc. Interspeech, 2025, pp. 5218–5222

  29. [37]

    The AMI meeting corpus: A pre-announcement,

    J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V . Karaiskos, W. Kraaij, M. Kronenthalet al., “The AMI meeting corpus: A pre-announcement,” inInterna- tional workshop on machine learning for multimodal interaction. Springer, 2006, pp. 28–39

  30. [38]

    AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,

    Y . Fu, L. Cheng, S. Lv, Y . Jv, Y . Kong, Z. Chen, Y . Hu, L. Xie, J. Wu, H. Bu, X. Xu, J. Du, and J. Chen, “AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,” inProc. Inter- speech, 2021, pp. 3665–3669

  31. [39]

    M2MeT: The ICASSP 2022 multi-channel multi-party meeting transcription challenge,

    F. Yu, S. Zhang, Y . Fu, L. Xie, S. Zheng, Z. Du, W. Huang, P. Guo, Z. Yan, B. Maet al., “M2MeT: The ICASSP 2022 multi-channel multi-party meeting transcription challenge,” inProc. ICASSP. IEEE, 2022, pp. 6167–6171

  32. [40]

    NOTSOFAR-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,

    A. Vinnikov, A. Ivry, A. Hurvitz, I. Abramovski, S. Koubi, I. Gurvich, S. Peer, X. Xiao, B. M. Elizalde, N. Kandaet al., “NOTSOFAR-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,” inProc. Interspeech, 2024, pp. 5003–5007

  33. [41]

    MSDWild: Multi- modal speaker diarization dataset in the wild,

    T. Liu, S. Fan, X. Xiang, H. Song, S. Lin, J. Sun, T. Han, S. Chen, B. Yao, S. Liu, Y . Wu, Y . Qian, and K. Yu, “MSDWild: Multi- modal speaker diarization dataset in the wild,” inProc. Inter- speech, 2022, pp. 1476–1480

  34. [42]

    Open source MagicData-RAMC: A rich annotated mandarin conversational (RAMC) speech dataset,

    Z. Yang, Y . Chen, L. Luo, R. Yang, L. Ye, G. Cheng, J. Xu, Y . Jin, Q. Zhang, P. Zhang, L. Xie, and Y . Yan, “Open source MagicData-RAMC: A rich annotated mandarin conversational (RAMC) speech dataset,” inProc. Interspeech, 2022, pp. 1736– 1740

  35. [43]

    Spot the conversation: Speaker diarisation in the wild,

    J. S. Chung, J. Huh, A. Nagrani, T. Afouras, and A. Zisserman, “Spot the conversation: Speaker diarisation in the wild,” inProc. Interspeech, 2020, pp. 299–303

  36. [44]

    WavLM: Large-scale self-supervised pre-training for full stack speech processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selecte...

  37. [45]

    Conformer: Convolution-augmented transformer for speech recognition,

    A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y . Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y . Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. Interspeech, 2020, pp. 5036–5040

  38. [46]

    Powerset multi-class cross entropy loss for neural speaker diarization,

    A. Plaquet and H. Bredin, “Powerset multi-class cross entropy loss for neural speaker diarization,” inProc. Interspeech, 2023, pp. 3222–3226

  39. [47]

    Leveraging self-supervised learning for speaker diarization,

    J. Han, F. Landini, J. Rohdin, A. Silnova, M. Diez, and L. Burget, “Leveraging self-supervised learning for speaker diarization,” in Proc. ICASSP. IEEE, 2025, pp. 1–5

  40. [48]

    Wespeaker: A research and production oriented speaker embedding learning toolkit,

    H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang, Y . Deng, and Y . Qian, “Wespeaker: A research and production oriented speaker embedding learning toolkit,” inProc. ICASSP. IEEE, 2023, pp. 1–5

  41. [49]

    V oxCeleb2: Deep speaker recognition,

    J. S. Chung, A. Nagrani, and A. Zisserman, “V oxCeleb2: Deep speaker recognition,” inProc. Interspeech, 2018, pp. 1086–1090

  42. [50]

    Multichannel source separation and tracking with ransac and directional statistics,

    J. Traa and P. Smaragdis, “Multichannel source separation and tracking with ransac and directional statistics,”IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, vol. 22, no. 12, pp. 2233–2243, 2014

Pith tools

Reviewed June 25, 2026 · model on record in the stance chip above.