Pith. sign in

REVIEW 1 cited by

On the Dimensionality of Word Embedding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1812.04224 v1 pith:MIRXN3EB submitted 2018-12-11 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords worddimensionalityembeddingbias-varianceembeddingstrade-offlossselection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we provide a theoretical understanding of word embedding and its dimensionality. Motivated by the unitary-invariance of word embedding, we propose the Pairwise Inner Product (PIP) loss, a novel metric on the dissimilarity between word embeddings. Using techniques from matrix perturbation theory, we reveal a fundamental bias-variance trade-off in dimensionality selection for word embeddings. This bias-variance trade-off sheds light on many empirical observations which were previously unexplained, for example the existence of an optimal dimensionality. Moreover, new insights and discoveries, like when and how word embeddings are robust to over-fitting, are revealed. By optimizing over the bias-variance trade-off of the PIP loss, we can explicitly answer the open question of dimensionality selection for word embedding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Topological Improvement of the Overall Performance of Sparse Evolutionary Training: Motif-Based Structural Optimization of Sparse MLPs Project

    cs.NE 2025-06 reject novelty 2.0 of 10

    Grouping neurons into blocks of size 2 with shared weights trims training time by 30 to 43 percent with accuracy losses of 1 to 4 percent on two datasets, by the paper's own measurements.

Pith tools