A new dataset and surprisal-based metric estimate that roughly 20% of naturally occurring generic sentences are weak generalisations and that generics are more context-sensitive than explicit quantifiers.
The Language of Generalization
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Language provides simple ways of communicating generalizable knowledge to each other (e.g., "Birds fly", "John hikes", "Fire makes smoke"). Though found in every language and emerging early in development, the language of generalization is philosophically puzzling and has resisted precise formalization. Here, we propose the first formal account of generalizations conveyed with language that makes quantitative predictions about human understanding. We test our model in three diverse domains: generalizations about categories (generic language), events (habitual language), and causes (causal language). The model explains the gradience in human endorsement through the interplay between a simple truth-conditional semantic theory and diverse beliefs about properties, formalized in a probabilistic model of language understanding. This work opens the door to understanding precisely how abstract knowledge is learned from language.
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Generics are puzzling. Can language models find the missing piece?
A new dataset and surprisal-based metric estimate that roughly 20% of naturally occurring generic sentences are weak generalisations and that generics are more context-sensitive than explicit quantifiers.