Pith. sign in

REVIEW 1 cited by

Paraphrastic Representations at Scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.15114 v2 pith:B5CNFE26 submitted 2021-04-30 cs.CL

classification cs.CL
keywords modelsdataparaphrasticsemanticsignificantlysimilaritytrainingcode
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a system that allows users to train their own state-of-the-art paraphrastic sentence representations in a variety of languages. We also release trained models for English, Arabic, German, French, Spanish, Russian, Turkish, and Chinese. We train these models on large amounts of data, achieving significantly improved performance from the original papers proposing the methods on a suite of monolingual semantic similarity, cross-lingual semantic similarity, and bitext mining tasks. Moreover, the resulting models surpass all prior work on unsupervised semantic textual similarity, significantly outperforming even BERT-based models like Sentence-BERT (Reimers and Gurevych, 2019). Additionally, our models are orders of magnitude faster than prior work and can be used on CPU with little difference in inference speed (even improved speed over GPU when using more CPU cores), making these models an attractive choice for users without access to GPUs or for use on embedded devices. Finally, we add significantly increased functionality to the code bases for training paraphrastic sentence models, easing their use for both inference and for training them for any desired language with parallel data. We also include code to automatically download and preprocess training data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SEFD: Semantic-Enhanced Framework for Detecting LLM-Generated Text

    cs.CL 2024-11 conditional novelty 5.0 of 10

    SEFD combines retrieval-based semantic similarity with existing detectors and an adaptive pool to improve detection of paraphrased LLM-generated text in sequential streams.

Pith tools