Using a designed linguistic schema and BERT probabilities, the paper reports 77,118 sheaf-contextual and 36.9 million CbD-contextual instances from Simple English Wikipedia, with Euclidean distance as the best statistical predictor.
Developments in Sheaf-Theoretic Models of Natural Language Ambiguities
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Sheaves are mathematical objects consisting of a base which constitutes a topological space and the data associated with each open set thereof, e.g. continuous functions defined on the open sets. Sheaves have originally been used in algebraic topology and logic. Recently, they have also modelled events such as physical experiments and natural language disambiguation processes. We extend the latter models from lexical ambiguities to discourse ambiguities arising from anaphora. To begin, we calculated a new measure of contextuality for a dataset of basic anaphoric discourses, resulting in a higher proportion of contextual models-82.9%-compared to previous work which only yielded 3.17% contextual models. Then, we show how an extension of the natural language processing challenge, known as the Winograd Schema, which involves anaphoric ambiguities can be modelled on the Bell-CHSH scenario with a contextual fraction of 0.096.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1roles
extension 1polarities
extend 1representative citing papers
citing papers explorer
-
Quantum-Like Contextuality in Large Language Models
Using a designed linguistic schema and BERT probabilities, the paper reports 77,118 sheaf-contextual and 36.9 million CbD-contextual instances from Simple English Wikipedia, with Euclidean distance as the best statistical predictor.