Pith. sign in

REVIEW 1 cited by

Efficient Vector Representation for Documents through Corruption

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1707.02377 v1 pith:VT5NOMLO submitted 2017-07-08 cs.CL cs.LG

classification cs.CLcs.LG
keywords documentdoc2veccmodelrepresentationcorruptionefficientembeddingslearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present an efficient document representation learning framework, Document Vector through Corruption (Doc2VecC). Doc2VecC represents each document as a simple average of word embeddings. It ensures a representation generated as such captures the semantic meanings of the document during learning. A corruption model is included, which introduces a data-dependent regularization that favors informative or rare words while forcing the embeddings of common and non-discriminative ones to be close to zero. Doc2VecC produces significantly better word embeddings than Word2Vec. We compare Doc2VecC with several state-of-the-art document representation learning algorithms. The simple model architecture introduced by Doc2VecC matches or out-performs the state-of-the-art in generating high-quality document representations for sentiment analysis, document classification as well as semantic relatedness tasks. The simplicity of the model enables training on billions of words per hour on a single machine. At the same time, the model is very efficient in generating representations of unseen documents at test time.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Using Images to Find Context-Independent Word Representations in Vector Space

    cs.CL 2024-11 reject novelty 6.0 of 10

    A method that represents each word by the concatenated autoencoder latent codes of images of its dictionary definition terms, evaluated on word similarity, categorization, and outlier detection.

Pith tools