Humans and large language models show similar comprehension failures on garden-path sentences, with stronger models correlating more closely with human performance across three tasks.
Computational Sentence-level Metrics Predicting Human Sentence Comprehension
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The majority of research in computational psycholinguistics has concentrated on the processing of words. This study introduces innovative methods for computing sentence-level metrics using multilingual large language models. The metrics developed sentence surprisal and sentence relevance and then are tested and compared to validate whether they can predict how humans comprehend sentences as a whole across languages. These metrics offer significant interpretability and achieve high accuracy in predicting human sentence reading speeds. Our results indicate that these computational sentence-level metrics are exceptionally effective at predicting and elucidating the processing difficulties encountered by readers in comprehending sentences as a whole across a variety of languages. Their impressive performance and generalization capabilities provide a promising avenue for future research in integrating LLMs and cognitive science.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models
Humans and large language models show similar comprehension failures on garden-path sentences, with stronger models correlating more closely with human performance across three tasks.