Pith. sign in

REVIEW 1 cited by

Topic Modeling with Wasserstein Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.12374 v2 pith:TYAUKOAY submitted 2019-07-24 cs.IR cs.AIcs.LG

classification cs.IRcs.AIcs.LG
keywords topictopicsautoencodersbetterdirichletdiscoverdistributionexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a novel neural topic model in the Wasserstein autoencoders (WAE) framework. Unlike existing variational autoencoder based models, we directly enforce Dirichlet prior on the latent document-topic vectors. We exploit the structure of the latent space and apply a suitable kernel in minimizing the Maximum Mean Discrepancy (MMD) to perform distribution matching. We discover that MMD performs much better than the Generative Adversarial Network (GAN) in matching high dimensional Dirichlet distribution. We further discover that incorporating randomness in the encoder output during training leads to significantly more coherent topics. To measure the diversity of the produced topics, we propose a simple topic uniqueness metric. Together with the widely used coherence measure NPMI, we offer a more wholistic evaluation of topic quality. Experiments on several real datasets show that our model produces significantly better topics than existing topic models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding Cross-Domain Adaptation in Low-Resource Topic Modeling

    cs.CL 2025-06 conditional novelty 5.0 of 10

    DALTA adapts a variational topic model from a high-resource source domain to a low-resource target domain via adversarial latent alignment, separate decoders, and a consistency loss, with a claimed generalization bound.

Pith tools