A method that represents each word by the concatenated autoencoder latent codes of images of its dictionary definition terms, evaluated on word similarity, categorization, and outlier detection.
Automated Generation of Multilingual Clusters for the Evaluation of Distributed Representations
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We propose a language-agnostic way of automatically generating sets of semantically similar clusters of entities along with sets of "outlier" elements, which may then be used to perform an intrinsic evaluation of word embeddings in the outlier detection task. We used our methodology to create a gold-standard dataset, which we call WikiSem500, and evaluated multiple state-of-the-art embeddings. The results show a correlation between performance on this dataset and performance on sentiment analysis.
fields
cs.CL 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Using Images to Find Context-Independent Word Representations in Vector Space
A method that represents each word by the concatenated autoencoder latent codes of images of its dictionary definition terms, evaluated on word similarity, categorization, and outlier detection.