Pith. sign in

REVIEW 2 cited by

Mitigating Data Sparsity for Short Text Topic Modeling by Topic-Semantic Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.12878 v1 pith:4JIXRPUI submitted 2022-11-23 cs.CL

classification cs.CL
keywords datatopiccontrastivelearningshortsparsitytextmodeling
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

To overcome the data sparsity issue in short text topic modeling, existing methods commonly rely on data augmentation or the data characteristic of short texts to introduce more word co-occurrence information. However, most of them do not make full use of the augmented data or the data characteristic: they insufficiently learn the relations among samples in data, leading to dissimilar topic distributions of semantically similar text pairs. To better address data sparsity, in this paper we propose a novel short text topic modeling framework, Topic-Semantic Contrastive Topic Model (TSCTM). To sufficiently model the relations among samples, we employ a new contrastive learning method with efficient positive and negative sampling strategies based on topic semantics. This contrastive learning method refines the representations, enriches the learning signals, and thus mitigates the sparsity issue. Extensive experimental results show that our TSCTM outperforms state-of-the-art baselines regardless of the data augmentation availability, producing high-quality topics and topic distributions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Topic Interpretability for Neural Topic Modeling through Topic-wise Contrastive Learning

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A topic-wise contrastive regularizer using precomputed NPMI similarities improves coherence and diversity of neural topic model topics on 20NG, Yahoo, and NYTimes.

  2. Aspect-Based Summarization with Self-Aspect Retrieval Enhanced Generation

    cs.CL 2025-04 conditional novelty 4.0 of 10

    SARESG prunes documents to aspect-relevant sentences via embedding similarity before LLM summarization, reporting gains over selective-context baselines on three datasets.

Pith tools