Pith. sign in

REVIEW 1 cited by

An Empirical Study on Clustering Pretrained Embeddings: Is Deep Strictly Better?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.05183 v1 pith:6CTI6LC2 submitted 2022-11-09 cs.CV cs.LG

classification cs.CVcs.LG
keywords methodsclusteringdeepembeddingsfacedatasetsdiscriminativeempirical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Recent research in clustering face embeddings has found that unsupervised, shallow, heuristic-based methods -- including $k$-means and hierarchical agglomerative clustering -- underperform supervised, deep, inductive methods. While the reported improvements are indeed impressive, experiments are mostly limited to face datasets, where the clustered embeddings are highly discriminative or well-separated by class (Recall@1 above 90% and often nearing ceiling), and the experimental methodology seemingly favors the deep methods. We conduct a large-scale empirical study of 17 clustering methods across three datasets and obtain several robust findings. Notably, deep methods are surprisingly fragile for embeddings with more uncertainty, where they match or even perform worse than shallow, heuristic-based methods. When embeddings are highly discriminative, deep methods do outperform the baselines, consistent with past results, but the margin between methods is much smaller than previously reported. We believe our benchmarks broaden the scope of supervised clustering methods beyond the face domain and can serve as a foundation on which these methods could be improved. To enable reproducibility, we include all necessary details in the appendices, and plan to release the code.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hyperspherical Variational Autoencoders Using Efficient Spherical Cauchy Distribution

    stat.ML 2025-06 accept novelty 6.0 of 10

    Spherical Cauchy latent variables give hyperspherical VAEs an exact Möbius reparameterization and stable, Bessel-free KL evaluation, matching vMF locally while running faster and remaining stable in high dimensions.

Pith tools