Pith. sign in

REVIEW 2 cited by

A Comprehensive Survey on Deep Clustering: Taxonomy, Challenges, and Future Directions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.07579 v1 pith:YZEVJITT submitted 2022-06-15 cs.LG cs.AI

classification cs.LGcs.AI
keywords clusteringdeeplearningrepresentationbeendatamethodssurvey
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Clustering is a fundamental machine learning task which has been widely studied in the literature. Classic clustering methods follow the assumption that data are represented as features in a vectorized form through various representation learning techniques. As the data become increasingly complicated and complex, the shallow (traditional) clustering methods can no longer handle the high-dimensional data type. With the huge success of deep learning, especially the deep unsupervised learning, many representation learning techniques with deep architectures have been proposed in the past decade. Recently, the concept of Deep Clustering, i.e., jointly optimizing the representation learning and clustering, has been proposed and hence attracted growing attention in the community. Motivated by the tremendous success of deep learning in clustering, one of the most fundamental machine learning tasks, and the large number of recent advances in this direction, in this paper we conduct a comprehensive survey on deep clustering by proposing a new taxonomy of different state-of-the-art approaches. We summarize the essential components of deep clustering and categorize existing methods by the ways they design interactions between deep representation learning and clustering. Moreover, this survey also provides the popular benchmark datasets, evaluation metrics and open-source implementations to clearly illustrate various experimental settings. Last but not least, we discuss the practical applications of deep clustering and suggest challenging topics deserving further investigations as future directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Information-Theoretic Generative Clustering of Documents

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A clustering method that replaces document embeddings with language-model probabilities over generated texts achieves state-of-the-art results on four document datasets.

  2. An Adaptive Framework for Multi-View Clustering Leveraging Conditional Entropy Optimization

    cs.AI 2024-12 reject novelty 4.0 of 10

    CE-MVC combines NMI-based and conditional-entropy-based view weighting with per-view autoencoders, and reports top clustering accuracy on DIGIT, COIL, RGB-D, and Caltech.

Pith tools