Pith. sign in

REVIEW 1 cited by

Extractive Summary as Discrete Latent Variables

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1811.05542 v2 pith:XZE5YF4S submitted 2018-11-14 cs.CL cs.LGstat.ML

classification cs.CLcs.LGstat.ML
keywords textlanguagelatentmethodstokensvariableschoosecompare
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we compare various methods to compress a text using a neural model. We find that extracting tokens as latent variables significantly outperforms the state-of-the-art discrete latent variable models such as VQ-VAE. Furthermore, we compare various extractive compression schemes. There are two best-performing methods that perform equally. One method is to simply choose the tokens with the highest tf-idf scores. Another is to train a bidirectional language model similar to ELMo and choose the tokens with the highest loss. If we consider any subsequence of a text to be a text in a broader sense, we conclude that language is a strong compression code of itself. Our finding justifies the high quality of generation achieved with hierarchical method, as their latent variables are nothing but natural language summary. We also conclude that there is a hierarchy in language such that an entire text can be predicted much more easily based on a sequence of a small number of keywords, which can be easily found by classical methods as tf-idf. We speculate that this extraction process may be useful for unsupervised hierarchical text generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MeloBottleneck: Self-Supervised Melody Skeleton Extraction with a Latent Subsequence Bottleneck

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Self-supervised latent-subsequence bottlenecks extract coherent melody skeletons that transfer more robustly than pseudo-label note classifiers and improve ornament-robust retrieval.

Pith tools