Pith. sign in

REVIEW 1 major objections 2 minor 65 references

Bridging Maximum Likelihood and Optimal Transport for Efficient Inference and Model Selection in Stochastic Block Models

T0 review · 1 major / 2 minor · reviewed 2026-06-29 · grok-4.3

Pith's one-line read Unregularized semi-relaxed Gromov-Wasserstein estimators recover stochastic block model parameters and cluster assignments in the asymptotic regime.

desk verdict The paper reinterprets SBM variational inference as semi-relaxed Gromov-Wasserstein, proves asymptotic consistency for the unregularized version, and shows empirically that regularization enables joint parameter recovery plus model selection. read the letter →

arxiv 2605.28488 v2 pith:JRAP3AVZ submitted 2026-05-27 stat.ML cs.LGmath.STstat.TH

classification stat.MLcs.LGmath.STstat.TH
keywords stochasticblockmodeloptimaltransportgromov-wassersteindistanceselectionvariationalinferencenetworkclusteringasymptoticconsistency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reinterprets maximum likelihood variational inference for stochastic block models as a semi-relaxed Gromov-Wasserstein projection under entropic regularization. It proves that removing the regularization yields consistent recovery of both the connectivity matrix and the latent cluster assignments when the number of nodes tends to infinity. Because this consistency does not guarantee good model selection in finite samples, the authors add a regularization term that promotes sparsity in the cluster proportions, allowing the same optimization to recover parameters and select the number of clusters at once.

What carries the argument

semi-relaxed Gromov-Wasserstein projection, which reformulates the variational inference objective and permits direct analysis of consistency for recovering SBM parameters.

What would settle it

A finite-sample simulation in which the regularized srGW estimator selects a number of clusters different from the true value would challenge the claim of simultaneous reliable selection.

Watch

Extended reading notes

Core claim

Maximum likelihood variational inference in stochastic block models is equivalent to a semi-relaxed Gromov-Wasserstein projection with entropic regularization. Unregularized srGW estimators consistently recover the SBM connectivity matrix and latent cluster assignments asymptotically. A regularized formulation performs both parameter estimation and selection of the number of clusters simultaneously in a single optimization problem.

Load-bearing premise

That sparsity-promoting regularization can be introduced to achieve reliable model selection in finite samples while preserving the recovery properties.

Editorial extensions

If this is right

  • Consistent asymptotic recovery of the connectivity matrix and assignments.
  • Joint inference and model selection avoids separate grid search procedures.
  • Regularization addresses the sparsity issue that prevents reliable finite-sample selection.
  • Provides an efficient alternative to traditional heuristic model selection in SBMs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The connection between MLVI and srGW may allow importing other optimal transport techniques to improve SBM inference.
  • Similar reformulations could apply to other network models beyond SBMs.
  • Finite-sample performance might improve further with adaptive regularization parameters.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 2 minor

Summary. The manuscript interprets maximum likelihood variational inference for stochastic block models as an entropically regularized semi-relaxed Gromov-Wasserstein projection. It proves that unregularized srGW estimators are consistent for recovering the connectivity matrix and latent assignments in the asymptotic regime, while acknowledging that this consistency does not yield reliable finite-sample model selection. It then presents an empirical demonstration that an additional regularized formulation performs joint parameter recovery and selection of the number of clusters within a single optimization problem.

Significance. If the consistency result is rigorously established, the work supplies a useful OT-based reinterpretation of MLVI for SBMs and a practical route to model selection that avoids separate grid search. The explicit recognition that asymptotic consistency alone is insufficient for finite-sample selection is a positive feature; the empirical evidence for the regularized joint estimator would need to be robust across multiple regimes to support the efficiency claim.

major comments (1)
  1. [Abstract] Abstract (final paragraph): the central claim that the regularized formulation 'yields estimators that simultaneously recover model parameters and select the number of clusters in a single optimization problem' rests entirely on empirical demonstration; because the manuscript does not supply a corresponding consistency or selection-consistency theorem for the regularized case, the load-bearing step from asymptotic unregularized recovery to finite-sample joint selection requires stronger empirical controls (e.g., explicit comparison against standard criteria such as BIC or ICL on the same simulated and real networks).
minor comments (2)
  1. The notation for the semi-relaxed Gromov-Wasserstein objective and its relation to the SBM likelihood should be introduced with an explicit equation linking the two formulations (currently only described in prose).
  2. Figure captions and experimental details should state the precise form of the sparsity-promoting regularization (e.g., the value of any additional penalty parameter) and the range of network sizes used to illustrate finite-sample behavior.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the careful reading of the manuscript and the constructive feedback. We address the single major comment below.

read point-by-point responses
  1. Referee: [Abstract] Abstract (final paragraph): the central claim that the regularized formulation 'yields estimators that simultaneously recover model parameters and select the number of clusters in a single optimization problem' rests entirely on empirical demonstration; because the manuscript does not supply a corresponding consistency or selection-consistency theorem for the regularized case, the load-bearing step from asymptotic unregularized recovery to finite-sample joint selection requires stronger empirical controls (e.g., explicit comparison against standard criteria such as BIC or ICL on the same simulated and real networks).

    Authors: We agree that the central claim for the regularized formulation is supported by empirical evidence rather than a consistency theorem, as already noted in the manuscript. To strengthen the empirical validation, we will add explicit comparisons of the regularized srGW estimator against BIC and ICL on the same simulated and real networks in the revised version. This will provide more robust controls for the joint recovery and selection performance. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; derivation chain is self-contained

full rationale

The abstract presents an interpretation of MLVI as entropically regularized srGW, followed by a consistency proof for the unregularized case in the asymptotic regime and an empirical demonstration for the regularized formulation's joint parameter recovery and model selection. No quoted equations or steps reduce a claimed prediction or result to a fitted quantity defined in terms of itself, nor do they rely on self-citation chains or imported uniqueness theorems. The paper explicitly flags the gap between asymptotic consistency and finite-sample reliability, treating regularization as an empirical mechanism rather than a derived necessity. This structure keeps the central claims independent of the inputs they are derived from.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract-only review supplies no explicit free parameters, axioms, or invented entities; full manuscript would be required to audit fitted scales, background assumptions, or new constructs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bridging Maximum Likelihood and Optimal Transport for Efficient Inference and Model Selection in Stochastic Block Models." pith.science (2026). https://pith.science/paper/JRAP3AVZ

@misc{pith2026260528488,
  author       = {Pith},
  title        = {Pith review of: Bridging Maximum Likelihood and Optimal Transport for Efficient Inference and Model Selection in Stochastic Block Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JRAP3AVZ}},
  note         = {Machine review of arXiv:2605.28488}
}
read the original abstract

We study inference in stochastic block models (SBMs) through the lens of optimal transport (OT). We first establish that maximum likelihood variational inference (MLVI) can be interpreted as a semi-relaxed Gromov-Wasserstein (srGW) projection with entropic regularization. While this formulation yields accurate clustering, the entropic regularization prevents transport plans to be sparse, hindering intrinsic model selection. Consequently, we investigate unregularized srGW estimators, and prove that they consistently recover both the SBM connectivity matrix and latent cluster assignments in the asymptotic regime. However, this asymptotic property does not translate into reliable model selection in finite samples, and calls for additional mechanisms to promote sparsity in the inferred cluster proportions. We empirically show that such a regularized formulation yields estimators that simultaneously recover model parameters and select the number of clusters in a single optimization problem, thereby avoiding costly grid search or heuristic model selection procedures.

Figures

Figures reproduced from arXiv: 2605.28488 by the authors.

Figure 1
Figure 1. Left : A graph sampled from a Bernoulli SBM with five communities. Right : The [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Evolution of the average ARI w.r.t α, 5 simulated graphs for each α. The higher α, the easier to detect the clusters. 3 2 1 0 log10 ( ) 0 5 10 15 20 Estimated K Assortative =0.03 =0.06 =0.10 =0.13 =0.17 =0.20 True K 3 2 1 0 log10 ( ) 0 5 10 15 20 Hub 3 2 1 0 log10 ( ) 0 5 10 15 20 Disassortative [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Average selected K∗ versus (log of) the sparsity hyper-parameter λ. widely used community detection methods respectively based on modularity maximization and flow compression. We also consider two model-based competitors: Greed [20], which performs greedy maximization of the Integrated Classification Likelihood (ICL), and the blockmodels package [57], which implements variational EM with ICL-based model selection. E… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Tree structure of our results up to Theorem 1. [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]
Figure 5
Figure 5. Figure 5: Sampled Adjacency matrices for the 3 SBM scenarions described in Section 5, i.e. [PITH_FULL_IMAGE:figures/full_fig_p024_5.png]
Figure 6
Figure 6. Figure 6: Evolution of ARI wrt α, the greater α, the easier to detect the clusters. We take a graph with N = 103 nodes. β = 0.03 and α ∈ [β, 0.2]. Each algorithm, excepts Louvain is searching for at most K = 20 clusters, the real number of cluster is K∗ = 5. Here the classes are…
Figure 7
Figure 7. Figure 7: Running times for different graph clustering methods. [PITH_FULL_IMAGE:figures/full_fig_p025_7.png]
Figure 8
Figure 8. Figure 8: Running Times of srGW NLL on cpu vs gpu. [PITH_FULL_IMAGE:figures/full_fig_p026_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

65 extracted references · 8 canonical work pages

  1. [1]

    A note on the relations between mixture models, maximum- likelihood and entropic optimal transport, 2025

    Titouan Vayer and Etienne Lasalle. A note on the relations between mixture models, maximum- likelihood and entropic optimal transport, 2025

  2. [2]

    Rohde, and Heiko Hoffmann

    Soheil Kolouri, Gustavo K. Rohde, and Heiko Hoffmann. Sliced wasserstein distance for learning gaussian mixture models. In2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3427–3436, 2018. doi: 10.1109/CVPR.2018.00361

  3. [3]

    Entropic optimal transport is maximum-likelihood deconvolution, 2018

    Philippe Rigollet and Jonathan Weed. Entropic optimal transport is maximum-likelihood deconvolution, 2018

  4. [4]

    Stochastic blockmodels: First steps.Social networks, 5(2):109–137, 1983

    Paul W Holland, Kathryn Blackmond Laskey, and Samuel Leinhardt. Stochastic blockmodels: First steps.Social networks, 5(2):109–137, 1983

  5. [5]

    Estimation and prediction for stochastic blockstruc- tures.Journal of the American statistical association, 96(455):1077–1087, 2001

    Krzysztof Nowicki and Tom A B Snijders. Estimation and prediction for stochastic blockstruc- tures.Journal of the American statistical association, 96(455):1077–1087, 2001

  6. [6]

    A mixture model for random graphs

    Jean-Jacques Daudin, Franck Picard, and Stephane Robin. A mixture model for random graphs. Statistics and Computing, 18(2):173–183, 2008. doi: 10.1007/s11222-007-9046-7

  7. [7]

    Network analysis in the social sciences.science, 323(5916):892–895, 2009

    Stephen P Borgatti, Ajay Mehra, Daniel J Brass, and Giuseppe Labianca. Network analysis in the social sciences.science, 323(5916):892–895, 2009

  8. [8]

    Journal of Statistical Mechanics: Theory and Experiment2008(10), 10008 (2008) https://doi.org/10.1088/1742-5468/2008/ 10/p10008

    Vincent D Blondel, Jean-Loup Guillaume, Renaud Lambiotte, and Etienne Lefebvre. Fast unfolding of communities in large networks.Journal of Statistical Mechanics: Theory and Experiment, 2008(10):P10008, October 2008. ISSN 1742-5468. doi: 10.1088/1742-5468/2008/ 10/p10008

Show all 65 references
  1. [9]

    Community detection in graphs.Physics reports, 486(3-5):75–174, 2010

    Santo Fortunato. Community detection in graphs.Physics reports, 486(3-5):75–174, 2010

  2. [10]

    A tutorial on spectral clustering.Statistics and computing, 17(4):395–416, 2007

    Ulrike V on Luxburg. A tutorial on spectral clustering.Statistics and computing, 17(4):395–416, 2007

  3. [11]

    Consistency of spectral clustering in stochastic block models

    Jing Lei and Alessandro Rinaldo. Consistency of spectral clustering in stochastic block models. The Annals of Statistics, 43(1), February 2015. ISSN 0090-5364. doi: 10.1214/14-aos1274

  4. [12]

    Mean-field theory of graph neural networks in graph partitioning.Advances in neural information processing systems, 31, 2018

    Tatsuro Kawamoto, Masashi Tsubaki, and Tomoyuki Obuchi. Mean-field theory of graph neural networks in graph partitioning.Advances in neural information processing systems, 31, 2018

  5. [13]

    Neurocut: A neural approach for robust graph partitioning

    Rishi Shah, Krishnanshu Jain, Sahil Manchanda, Sourav Medya, and Sayan Ranu. Neurocut: A neural approach for robust graph partitioning. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 2584–2595, 2024

  6. [14]

    Scalable gromov-wasserstein learning for graph partitioning and matching.Advances in neural information processing systems, 32, 2019

    Hongteng Xu, Dixin Luo, and Lawrence Carin. Scalable gromov-wasserstein learning for graph partitioning and matching.Advances in neural information processing systems, 32, 2019

  7. [15]

    Semi- relaxed gromov wasserstein divergence with applications on graphs.CoRR, abs/2110.02753, 2021

    C´edric Vincent-Cuaz, R´emi Flamary, Marco Corneli, Titouan Vayer, and Nicolas Courty. Semi- relaxed gromov wasserstein divergence with applications on graphs.CoRR, abs/2110.02753, 2021

  8. [16]

    Heat kernel based community detection

    Kyle Kloster and David F Gleich. Heat kernel based community detection. InProceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 1386–1395, 2014

  9. [17]

    Stochastic blockmodels and community structure in networks.Physical Review E—Statistical, Nonlinear, and Soft Matter Physics, 83(1):016107, 2011

    Brian Karrer and Mark EJ Newman. Stochastic blockmodels and community structure in networks.Physical Review E—Statistical, Nonlinear, and Soft Matter Physics, 83(1):016107, 2011

  10. [18]

    Assessing a mixture model for clustering with the integrated completed likelihood.IEEE transactions on pattern analysis and machine intelligence, 22(7):719–725, 2000

    Christophe Biernacki, Gilles Celeux, and G ´erard Govaert. Assessing a mixture model for clustering with the integrated completed likelihood.IEEE transactions on pattern analysis and machine intelligence, 22(7):719–725, 2000

  11. [19]

    Improved bayesian inference for the stochastic block model with application to large networks.Computational Statistics & Data Analysis, 60:12–31, 2013

    Aaron F McDaid, Thomas Brendan Murphy, Nial Friel, and Neil J Hurley. Improved bayesian inference for the stochastic block model with application to large networks.Computational Statistics & Data Analysis, 60:12–31, 2013. 11

  12. [20]

    Hierarchical clustering with discrete latent variable models and the integrated classification likelihood.Advances in Data Analysis and Classification, 15(4):957–986, 2021

    Etienne Cˆome, Nicolas Jouvin, Pierre Latouche, and Charles Bouveyron. Hierarchical clustering with discrete latent variable models and the integrated classification likelihood.Advances in Data Analysis and Classification, 15(4):957–986, 2021. doi: 10.1007/s11634-021-00440-z

  13. [21]

    Chao Gao, Yu Lu, and Harrison H. Zhou. Rate-optimal graphon estimation.The Annals of Statistics, 43(6), December 2015. ISSN 0090-5364. doi: 10.1214/15-aos1354

  14. [22]

    Maximum likelihood estimation of sparse networks with missing observations.Journal of Statistical Planning and Inference, 215:299–329, 2021

    Solenne Gaucher and Olga Klopp. Maximum likelihood estimation of sparse networks with missing observations.Journal of Statistical Planning and Inference, 215:299–329, 2021. ISSN 0378-3758. doi: https://doi.org/10.1016/j.jspi.2021.04.003

  15. [23]

    Optimality of variational inference for stochasticblock model with missing links

    Solenne Gaucher and Olga Klopp. Optimality of variational inference for stochasticblock model with missing links. In M. Ranzato, A. Beygelzimer, Y . Dauphin, P.S. Liang, and J. Wortman Vaughan, editors,Advances in Neural Information Processing Systems, volume 34, pages 19947–1...

  16. [24]

    Spectral clustering of graphs with the bethe hessian, 2014

    Alaa Saade, Florent Krzakala, and Lenka Zdeborov ´a. Spectral clustering of graphs with the bethe hessian, 2014

  17. [25]

    Recovering communities in the general stochastic block model without knowing the parameters, 2015

    Emmanuel Abbe and Colin Sandon. Recovering communities in the general stochastic block model without knowing the parameters, 2015

  18. [26]

    Grundlehren der mathematis- chen Wissenschaften

    C´edric Villani.Optimal transport : old and new / C ´edric Villani. Grundlehren der mathematis- chen Wissenschaften. Springer, Berlin, 2009. ISBN 978-3-540-71049-3

  19. [27]

    Computational optimal transport, 2020

    Gabriel Peyr ´e and Marco Cuturi. Computational optimal transport, 2020

  20. [28]

    A survey on optimal transport for machine learning: Theory and applications.IEEE Access, 2025

    Luiz Manella Pereira and M Hadi Amini. A survey on optimal transport for machine learning: Theory and applications.IEEE Access, 2025

  21. [29]

    Gromov—wasserstein distances and the metric approach to object matching

    Facundo M´emoli. Gromov—wasserstein distances and the metric approach to object matching. Found. Comput. Math., 11(4):417–487, August 2011. ISSN 1615-3375

  22. [30]

    American Mathematical Society, 2023

    Karl-Theodor Sturm.The space of spaces: curvature bounds and gradient flows on the space of metric measure spaces, volume 290. American Mathematical Society, 2023

  23. [31]

    The gromov–wasserstein distance between networks and stable network invariants.Information and Inference: A Journal of the IMA, 8(4):757–787, 2019

    Samir Chowdhury and Facundo M´emoli. The gromov–wasserstein distance between networks and stable network invariants.Information and Inference: A Journal of the IMA, 8(4):757–787, 2019

  24. [32]

    Gromov-wasserstein learning for graph matching and node embedding

    Hongteng Xu, Dixin Luo, Hongyuan Zha, and Lawrence Carin Duke. Gromov-wasserstein learning for graph matching and node embedding. InInternational conference on machine learning, pages 6932–6941. PMLR, 2019

  25. [33]

    Quantized gromov-wasserstein

    Samir Chowdhury, David Miller, and Tom Needham. Quantized gromov-wasserstein. InJoint European Conference on Machine Learning and Knowledge Discovery in Databases, pages 811–827. Springer, 2021

  26. [34]

    Gromov-Wasserstein Averaging of Kernel and Distance Matrices

    Gabriel Peyr´e, Marco Cuturi, and Justin Solomon. Gromov-Wasserstein Averaging of Kernel and Distance Matrices. InProc. 33rd International Conference on Machine Learning, Proc. 33rd International Conference on Machine Learning, New-York, United States, June 2016

  27. [35]

    Optimal transport for structured data with application on graphs

    Vayer Titouan, Nicolas Courty, Romain Tavenard, Chapel Laetitia, and R´emi Flamary. Optimal transport for structured data with application on graphs. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors,Proceedings of the 36th International Conference on Machine Learning, v...

  28. [36]

    Generalized spectral clustering via gromov-wasserstein learning

    Samir Chowdhury and Tom Needham. Generalized spectral clustering via gromov-wasserstein learning. InInternational Conference on Artificial Intelligence and Statistics, pages 712–720. PMLR, 2021

  29. [37]

    Optimal transport-based cluster- ing of attributed graphs with an application to road traffic data, 2025

    Ioana Gavra, Ketsia Guichard-Sustowski, and Lo¨ıc Le Marrec. Optimal transport-based cluster- ing of attributed graphs with an application to road traffic data, 2025. 12

  30. [38]

    Learning graphons via struc- tured gromov-wasserstein barycenters

    Hongteng Xu, Dixin Luo, Lawrence Carin, and Hongyuan Zha. Learning graphons via struc- tured gromov-wasserstein barycenters. InProceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 10505–10513, 2021

  31. [39]

    A gromov-wasserstein geometric view of spectrum-preserving graph coarsening

    Yifan Chen, Rentian Yao, Yun Yang, and Jie Chen. A gromov-wasserstein geometric view of spectrum-preserving graph coarsening. InInternational Conference on Machine Learning, pages 5257–5281. PMLR, 2023

  32. [40]

    Gromov-wasserstein factorization models for graph clustering

    Hongtengl Xu. Gromov-wasserstein factorization models for graph clustering. InProceedings of the AAAI conference on artificial intelligence, volume 34, pages 6478–6485, 2020

  33. [41]

    Online graph dictionary learning

    C´edric Vincent-Cuaz, Titouan Vayer, R´emi Flamary, Marco Corneli, and Nicolas Courty. Online graph dictionary learning. InInternational conference on machine learning, pages 10564–10574. PMLR, 2021

  34. [42]

    Robust graph dictionary learning

    Weijie Liu, Jiahao Xie, Chao Zhang, Makoto Yamada, Nenggan Zheng, and Hui Qian. Robust graph dictionary learning. InThe Eleventh International Conference on Learning Representa- tions, 2022

  35. [43]

    Generative graph dictionary learning

    Zhichen Zeng, Ruike Zhu, Yinglong Xia, Hanqing Zeng, and Hanghang Tong. Generative graph dictionary learning. InInternational Conference on Machine Learning, pages 40749–40769. PMLR, 2023

  36. [44]

    Template based graph neural network with optimal transport distances.Advances in Neural Information Processing Systems, 35:11800–11814, 2022

    C´edric Vincent-Cuaz, R ´emi Flamary, Marco Corneli, Titouan Vayer, and Nicolas Courty. Template based graph neural network with optimal transport distances.Advances in Neural Information Processing Systems, 35:11800–11814, 2022

  37. [45]

    Wasserstein barycenter matching for graph size generalization of message passing neural networks

    Xu Chu, Yujie Jin, Xin Wang, Shanghang Zhang, Yasha Wang, Wenwu Zhu, and Hong Mei. Wasserstein barycenter matching for graph size generalization of message passing neural networks. InInternational Conference on Machine Learning, pages 6158–6184. PMLR, 2023

  38. [46]

    Reimagining graph classification from a prototype view with optimal transport: Algorithm and theorem

    Chen Qian, Huayi Tang, Hong Liang, and Yong Liu. Reimagining graph classification from a prototype view with optimal transport: Algorithm and theorem. InProceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining, pages 2444–2454, 2024

  39. [47]

    The quest for the GRAph level autoencoder (GRALE)

    Paul Krzakala, Gabriel Melo, Charlotte Laclau, Florence d’Alch´e Buc, and R´emi Flamary. The quest for the GRAph level autoencoder (GRALE). InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026

  40. [48]

    Distributional reduction: Unifying dimensionality reduction and clustering with gromov-wasserstein.Transactions on Machine Learning Research Journal, 2025

    Hugues Van Assel, C´edric Vincent-Cuaz, Nicolas Courty, R´emi Flamary, Pascal Frossard, and Titouan Vayer. Distributional reduction: Unifying dimensionality reduction and clustering with gromov-wasserstein.Transactions on Machine Learning Research Journal, 2025

  41. [49]

    Generalized dimension reduction using semi-relaxed gromov-wasserstein distance

    Ranthony A Clark, Tom Needham, and Thomas Weighill. Generalized dimension reduction using semi-relaxed gromov-wasserstein distance. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 16082–16090, 2025

  42. [50]

    Kernel k-means: spectral clustering and normalized cuts

    Inderjit S Dhillon, Yuqiang Guan, and Brian Kulis. Kernel k-means: spectral clustering and normalized cuts. InProceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining, pages 551–556, 2004

  43. [51]

    A review of stochastic block models and extensions for graph clustering.Applied Network Science, 4(1):122, 2019

    Clement Lee and Darren J Wilkinson. A review of stochastic block models and extensions for graph clustering.Applied Network Science, 4(1):122, 2019

  44. [52]

    Daudin, and Laurent Pierre

    Alain Celisse, J.-J. Daudin, and Laurent Pierre. Consistency of maximum-likelihood and variational estimators in the Stochastic Block Model. working paper or preprint, May 2011

  45. [53]

    Convergence of the groups posterior distribution in latent or stochastic block models.Bernoulli, 21(1):537–573, 2015

    Mahendra Mariadassou and Catherine Matias. Convergence of the groups posterior distribution in latent or stochastic block models.Bernoulli, 21(1):537–573, 2015

  46. [54]

    A graph matching approach to balanced data sub- sampling for self-supervised learning

    Hugues Van Assel and Randall Balestriero. A graph matching approach to balanced data sub- sampling for self-supervised learning. InNeurIPS 2024 Workshop: Self-Supervised Learning- Theory and Practice, 2024. 13

  47. [55]

    Itera- tive bregman projections for regularized transportation problems.SIAM Journal on Scientific Computing, 37(2):A1111–A1138, 2015

    Jean-David Benamou, Guillaume Carlier, Marco Cuturi, Luca Nenna, and Gabriel Peyr´e. Itera- tive bregman projections for regularized transportation problems.SIAM Journal on Scientific Computing, 37(2):A1111–A1138, 2015

  48. [56]

    The map equation.The European Physical Journal Special Topics, 178(1):13–23, 2009

    Martin Rosvall, Daniel Axelsson, and Carl T Bergstrom. The map equation.The European Physical Journal Special Topics, 178(1):13–23, 2009

  49. [57]

    Blockmodels: A r-package for estimating in latent block model and stochastic block model, with various probability functions, with or without covariates, 2016

    Jean-Benoist Leger. Blockmodels: A r-package for estimating in latent block model and stochastic block model, with various probability functions, with or without covariates, 2016

  50. [58]

    Adjusting for chance clustering comparison measures

    Simone Romano, Nguyen Xuan Vinh, James Bailey, and Karin Verspoor. Adjusting for chance clustering comparison measures. 17(1), 2016. ISSN 1532-4435

  51. [59]

    Comparing partitions.J

    Lawrence Hubert and Phipps Arabie. Comparing partitions.J. Classif., 2(1):193–218, December 1985

  52. [60]

    Cambridge Series in Statistical and Probabilistic Mathematics

    Roman Vershynin.High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, Cambridge, 2 edition, 2026

  53. [61]

    Asymptotic statistics (1998)

    AW van der Vaart. Asymptotic statistics (1998). 1998

  54. [62]

    Pot python optimal transport (version 0.9.5), 2024

    R´emi Flamary, C´edric Vincent-Cuaz, Nicolas Courty, Alexandre Gramfort, Oleksii Kachaiev, Huy Quang Tran, Laur`ene David, Cl´ement Bonet, Nathan Cassereau, Th´eo Gnassounou, Eloi Tanguy, Julie Delon, Antoine Collas, Sonia Mazelet, Laetitia Chapel, Tanguy Kerdoncuff, Xizheng Y...

  55. [63]

    the alternative version is true for all SBMs whereas the original only holds for models which satisfyE[A ij|Z] = Θ ZiZj (the Bernoulli model is a specific case here)

  56. [64]

    projection

    the proof of proposition of the alternative version exploits the KL divergence, whereas the one of the original version exploits the definition Bregman divergence. Moreover, a summary of what proved as preliminar to the statement of Theorem 1 cans be seen in Figure 4. 2Recall ...

  57. [65]

    Justification: The research does not involve human subjects, therefore IRB approval is not applicable

    Institutional review board (IRB) approvals or equivalent for research with human subjects Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board (IRB) approvals...

Pith tools

Reviewed June 29, 2026 · model on record in the stance chip above.