A review of 58 papers finds that large language models pass some Theory of Mind tests but remain brittle and fall short of human performance.
Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
A review of 58 papers finds that large language models pass some Theory of Mind tests but remain brittle and fall short of human performance.