Pith. sign in

REVIEW 2 cited by

A Survey of Word Embeddings Evaluation Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1801.09536 v1 pith:62LVSXEQ submitted 2018-01-21 cs.CL

classification cs.CL
keywords evaluationmethodswordembeddingsproposingrepresentationsableadequate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Word embeddings are real-valued word representations able to capture lexical semantics and trained on natural language corpora. Models proposing these representations have gained popularity in the recent years, but the issue of the most adequate evaluation method still remains open. This paper presents an extensive overview of the field of word embeddings evaluation, highlighting main problems and proposing a typology of approaches to evaluation, summarizing 16 intrinsic methods and 12 extrinsic methods. I describe both widely-used and experimental methods, systematize information about evaluation datasets and discuss some key challenges.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    For English-to-Czech MT evaluation, embeddings that win intrinsic semantic similarity benchmarks (SimCSE) perform worst after fine-tuning, while intrinsically poor models (XLM-R, FERNET) perform best.

  2. BlueGlass: A Framework for Composite AI Safety

    cs.AI 2025-07 conditional novelty 5.0 of 10

    BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...

Pith tools