Pith. sign in

REVIEW 4 cited by

Unsupervised Learning of Sentence Embeddings using Compositional n-Gram Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1703.02507 v3 pith:ZTM7WKUY submitted 2017-03-07 cs.CL cs.AIcs.IR

classification cs.CLcs.AIcs.IR
keywords embeddingsunsupervisedrepresentationssentencewordapplicationsbenchmarkcompositional
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The recent tremendous success of unsupervised word embeddings in a multitude of applications raises the obvious question if similar methods could be derived to improve embeddings (i.e. semantic representations) of word sequences as well. We present a simple but efficient unsupervised objective to train distributed representations of sentences. Our method outperforms the state-of-the-art unsupervised models on most benchmark tasks, highlighting the robustness of the produced general-purpose sentence embeddings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Sequence-to-Sequence Models Perceive Language Styles?

    cs.CL 2019-08 conditional novelty 6.0 of 10

    Style in text is represented by the covariance matrix of seq2seq semantic vectors, enabling a whitening-coloring style transfer algorithm.

  2. Hamming Sentence Embeddings for Information Retrieval

    cs.IR 2019-08 conditional novelty 6.0 of 10

    A neural compressor turns sentence embeddings into binary codes that retain semantic similarity performance on STS benchmarks while cutting memory by up to 256:1.

  3. When Deep Learning Meets Information Retrieval-based Bug Localization: A Survey

    cs.SE 2025-04 conditional novelty 4.0 of 10

    A systematic literature survey of 61 deep-learning-based IRBL studies that proposes taxonomies and a performance overview, concluding that DL mitigates lexical gap, code structure, and cold-start issues while LLM-base...

  4. BioBridge: Unified Bio-Embedding with Bridging Modality in Code-Switched EMR

    cs.CL 2024-12 conditional novelty 4.0 of 10

    BioBridge improves emergency triage classification on Korean-English code-switched EMRs by adding language segment tokens and BioSent2Vec medical features to transformer encoders, with modest gains over baselines.

Pith tools