Pith. sign in

REVIEW 1 cited by

Enhancing Short-Text Topic Modeling with LLM-Driven Context Expansion and Prefix-Tuned VAEs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03071 v2 pith:RO5MWJRV submitted 2024-10-04 cs.CL cs.IR

classification cs.CLcs.IR
keywords topicmodelingmodelsshort-texttextsdatalanguagepropose
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Topic modeling is a powerful technique for uncovering hidden themes within a collection of documents. However, the effectiveness of traditional topic models often relies on sufficient word co-occurrence, which is lacking in short texts. Therefore, existing approaches, whether probabilistic or neural, frequently struggle to extract meaningful patterns from such data, resulting in incoherent topics. To address this challenge, we propose a novel approach that leverages large language models (LLMs) to extend short texts into more detailed sequences before applying topic modeling. To further improve the efficiency and solve the problem of semantic inconsistency from LLM-generated texts, we propose to use prefix tuning to train a smaller language model coupled with a variational autoencoder for short-text topic modeling. Our method significantly improves short-text topic modeling performance, as demonstrated by extensive experiments on real-world datasets with extreme data sparsity, outperforming current state-of-the-art topic models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Advanced Topic Modeling Techniques for Categorizing Software Vulnerabilities

    cs.CR 2026-07 reject novelty 2.5 of 10

    Existing embedding-based topic models produce interpretable clusters on Cisco vulnerability Threat text, but without quantitative coherence scores, baselines, or downstream prioritization metrics.

Pith tools