Pith. sign in

REVIEW 1 cited by

An Analysis of Lemmatization on Topic Models of Morphologically Rich Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1608.03995 v2 pith:VKYO5JJY submitted 2016-08-13 cs.CL

classification cs.CL
keywords modelstopiclemmatizationeffectmorphologicallyrichstemmingword
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Topic models are typically represented by top-$m$ word lists for human interpretation. The corpus is often pre-processed with lemmatization (or stemming) so that those representations are not undermined by a proliferation of words with similar meanings, but there is little public work on the effects of that pre-processing. Recent work studied the effect of stemming on topic models of English texts and found no supporting evidence for the practice. We study the effect of lemmatization on topic models of Russian Wikipedia articles, finding in one configuration that it significantly improves interpretability according to a word intrusion metric. We conclude that lemmatization may benefit topic models on morphologically rich languages, but that further investigation is needed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Topic Modeling and Sentiment Analysis on Japanese Online Media's Coverage of Nuclear Energy

    cs.CL 2024-11 conditional novelty 5.0 of 10

    Japanese YouTube news coverage and comments on nuclear energy cluster into 16 topics, with an overall slightly negative sentiment that is partly a known artifact of the sentiment model's bias.

Pith tools