Pith. sign in

REVIEW 2 cited by

Sparse topic modeling via spectral decomposition and thresholding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.06730 v1 pith:MHDMWEVH submitted 2023-10-10 stat.ME

classification stat.ME
keywords matrixassumptionproceduretopic-wordcorpusparameterregimesdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The probabilistic Latent Semantic Indexing model assumes that the expectation of the corpus matrix is low-rank and can be written as the product of a topic-word matrix and a word-document matrix. In this paper, we study the estimation of the topic-word matrix under the additional assumption that the ordered entries of its columns rapidly decay to zero. This sparsity assumption is motivated by the empirical observation that the word frequencies in a text often adhere to Zipf's law. We introduce a new spectral procedure for estimating the topic-word matrix that thresholds words based on their corpus frequencies, and show that its $\ell_1$-error rate under our sparsity assumption depends on the vocabulary size $p$ only via a logarithmic term. Our error bound is valid for all parameter regimes and in particular for the setting where $p$ is extremely large; this high-dimensional setting is commonly encountered but has not been adequately addressed in prior literature. Furthermore, our procedure also accommodates datasets that violate the separability assumption, which is necessary for most prior approaches in topic modeling. Experiments with synthetic data confirm that our procedure is computationally fast and allows for consistent estimation of the topic-word matrix in a wide variety of parameter regimes. Our procedure also performs well relative to well-established methods when applied to a large corpus of research paper abstracts, as well as the analysis of single-cell and microbiome data where the same statistical model is relevant but the parameter regimes are vastly different.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tensor Topic Modeling Via HOSVD

    math.ST 2024-12 conditional novelty 6.0 of 10

    A HOSVD-based estimator for Tucker-decomposed tensor topic models recovers factor matrices and core tensor with entry-wise l1 error rates.

  2. Graph Topic Modeling for Documents with Spatial or Covariate Dependencies

    cs.LG 2024-12 conditional novelty 6.0 of 10

    GpLSI extends frequentist pLSI with graph total-variation denoising of left singular vectors, yielding improved topic mixture estimation on short documents and high-probability error bounds under low-p and anchor-docu...

Pith tools