Grouping neurons into blocks of size 2 with shared weights trims training time by 30 to 43 percent with accuracy losses of 1 to 4 percent on two datasets, by the paper's own measurements.
On the Dimensionality of Word Embedding
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this paper, we provide a theoretical understanding of word embedding and its dimensionality. Motivated by the unitary-invariance of word embedding, we propose the Pairwise Inner Product (PIP) loss, a novel metric on the dissimilarity between word embeddings. Using techniques from matrix perturbation theory, we reveal a fundamental bias-variance trade-off in dimensionality selection for word embeddings. This bias-variance trade-off sheds light on many empirical observations which were previously unexplained, for example the existence of an optimal dimensionality. Moreover, new insights and discoveries, like when and how word embeddings are robust to over-fitting, are revealed. By optimizing over the bias-variance trade-off of the PIP loss, we can explicitly answer the open question of dimensionality selection for word embedding.
citation-role summary
citation-polarity summary
fields
cs.NE 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
A Topological Improvement of the Overall Performance of Sparse Evolutionary Training: Motif-Based Structural Optimization of Sparse MLPs Project
Grouping neurons into blocks of size 2 with shared weights trims training time by 30 to 43 percent with accuracy losses of 1 to 4 percent on two datasets, by the paper's own measurements.