Pith. sign in

Interpretable Word Sense Representations via Definition Generation: The Case of Semantic Change Analysis

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We propose using automatically generated natural language definitions of contextualised word usages as interpretable word and word sense representations. Given a collection of usage examples for a target word, and the corresponding data-driven usage clusters (i.e., word senses), a definition is generated for each usage with a specialised Flan-T5 language model, and the most prototypical definition in a usage cluster is chosen as the sense label. We demonstrate how the resulting sense labels can make existing approaches to semantic change analysis more interpretable, and how they can allow users -- historical linguists, lexicographers, or social scientists -- to explore and intuitively explain diachronic trajectories of word meaning. Semantic change analysis is only one of many possible applications of the `definitions as representations' paradigm. Beyond being human-readable, contextualised definitions also outperform token or usage sentence embeddings in word-in-context semantic similarity judgements, making them a new promising type of lexical representation for NLP.

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Towards Universal Semantics With Large Language Models

cs.CL · 2025-05-17 · conditional · novelty 7.0

Fine-tuned 1B and 8B LLMs generate NSM explications that score higher than GPT-4o on the paper's automatic legality, substitutability, and cross-translatability metrics.

citing papers explorer

Showing 1 of 1 citing paper.

  • Towards Universal Semantics With Large Language Models cs.CL · 2025-05-17 · conditional · none · ref 16 · internal anchor

    Fine-tuned 1B and 8B LLMs generate NSM explications that score higher than GPT-4o on the paper's automatic legality, substitutability, and cross-translatability metrics.