REVIEW 2 cited by
Topic Modeling with Contextualized Word Representation Clusters
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Topic Modeling with Contextualized Word Representation Clusters
read the original abstract
Clustering token-level contextualized word representations produces output that shares many similarities with topic models for English text collections. Unlike clusterings of vocabulary-level word embeddings, the resulting models more naturally capture polysemy and can be used as a way of organizing documents. We evaluate token clusterings trained from several different output layers of popular contextualized language models. We find that BERT and GPT-2 produce high quality clusterings, but RoBERTa does not. These cluster models are simple, reliable, and can perform as well as, if not better than, LDA topic models, maintaining high topic quality even when the number of topics is large relative to the size of the local collection.
Forward citations
Cited by 2 Pith papers
-
Disentangling Similarity and Relatedness in Topic Models
Topic models lie on a spectrum from thematic-relatedness-rich to similarity-rich, and that position predicts which downstream tasks they handle well.
-
Disentangling Similarity and Relatedness in Topic Models
Topic-model families occupy distinct positions on a similarity–relatedness plane, and those positions predict which downstream tasks they help or hurt.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.