Pith. sign in

REVIEW 2 cited by

Comparative Analysis of Word Embeddings for Capturing Word Similarities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.03812 v1 pith:YCBNEQKQ submitted 2020-05-08 cs.CL cs.LG

classification cs.CLcs.LG
keywords wordsimilaritiesembeddingslanguagedistributedembeddinganalysiscapturing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Distributed language representation has become the most widely used technique for language representation in various natural language processing tasks. Most of the natural language processing models that are based on deep learning techniques use already pre-trained distributed word representations, commonly called word embeddings. Determining the most qualitative word embeddings is of crucial importance for such models. However, selecting the appropriate word embeddings is a perplexing task since the projected embedding space is not intuitive to humans. In this paper, we explore different approaches for creating distributed word representations. We perform an intrinsic evaluation of several state-of-the-art word embedding methods. Their performance on capturing word similarities is analysed with existing benchmark datasets for word pairs similarities. The research in this paper conducts a correlation analysis between ground truth word similarities and similarities obtained by different word embedding methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Measuring Contextual Informativeness in Child-Directed Text

    cs.CL 2024-12 conditional novelty 6.0 of 10

    An LLM-based scorer predicts human-judged contextual informativeness in children's stories with a Spearman correlation of 0.4983, outperforming baselines and generalizing to adult text.

  2. A Comparative Analysis of Transformer and LSTM Models for Detecting Suicidal Ideation on Reddit

    cs.LG 2024-11 conditional novelty 4.0 of 10

    RoBERTa outperforms other transformers and BERT-embedded LSTMs on a new 37,821-post Reddit dataset for suicidal-ideation detection, but the evaluation has annotation and significance gaps.

Pith tools