Pith. sign in

REVIEW 2 cited by

Integrating Document Clustering and Topic Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1309.6874 v1 pith:A3EA45VY submitted 2013-09-26 cs.LG cs.CLcs.IRstat.ML

classification cs.LGcs.CLcs.IRstat.ML
keywords topicclusteringdocumentmodeltopicsmodelingclusterclusters
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Document clustering and topic modeling are two closely related tasks which can mutually benefit each other. Topic modeling can project documents into a topic space which facilitates effective document clustering. Cluster labels discovered by document clustering can be incorporated into topic models to extract local topics specific to each cluster and global topics shared by all clusters. In this paper, we propose a multi-grain clustering topic model (MGCTM) which integrates document clustering and topic modeling into a unified framework and jointly performs the two tasks to achieve the overall best performance. Our model tightly couples two components: a mixture component used for discovering latent groups in document collection and a topic model component used for mining multi-grain topics including local topics specific to each cluster and global topics shared across clusters.We employ variational inference to approximate the posterior of hidden variables and learn model parameters. Experiments on two datasets demonstrate the effectiveness of our model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding Developer Pain Points in Federated Learning: Insights from Stack Overflow and GitHub

    cs.SE 2026-07 conditional novelty 6.0 of 10

    Federated-learning developers' biggest public pain points are environment setup, API/version breakage, non-IID training instability, and evaluation/privacy integration, with Stack Overflow skewing How and GitHub skewing Why.

  2. Estimating the Effective Topics of Articles and journals Abstract Using LDA And K-Means Clustering Algorithm

    cs.IR 2025-08 unverdicted novelty 2.0 of 10

    An application of LDA and K-Means clustering with WordNet to extract topics and keyphrases from abstracts, with no released code, data, or quantitative evaluation.

Pith tools