Controlled Experiments for Word Embeddings

Adriaan M. J. Schakel; Benjamin J. Wilson

arxiv: 1510.02675 · v2 · pith:UR7ACT4Knew · submitted 2015-10-09 · 💻 cs.CL

Controlled Experiments for Word Embeddings

Benjamin J. Wilson , Adriaan M. J. Schakel This is my paper

classification 💻 cs.CL

keywords wordexperimentsco-occurrencenoiseapproachcontrolleddistributionembeddings

0 comments

read the original abstract

An experimental approach to studying the properties of word embeddings is proposed. Controlled experiments, achieved through modifications of the training corpus, permit the demonstration of direct relations between word properties and word vector direction and length. The approach is demonstrated using the word2vec CBOW model with experiments that independently vary word frequency and word co-occurrence noise. The experiments reveal that word vector length depends more or less linearly on both word frequency and the level of noise in the co-occurrence distribution of the word. The coefficients of linearity depend upon the word. The special point in feature space, defined by the (artificial) word with pure noise in its co-occurrence distribution, is found to be small but non-zero.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Semantic Gradients Interactions in SSD: A Case Study in Racial Identity and Hate Speech
cs.CL 2026-05 unverdicted novelty 6.0

Interaction SSD extends semantic differential modeling with main, interaction, and conditional gradients to test moderation by group identity, applied to racial differences in hate speech ratings on the UC Berkeley corpus.