A review of 58 papers finds that large language models pass some Theory of Mind tests but remain brittle and fall short of human performance.
In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
A review of 58 papers finds that large language models pass some Theory of Mind tests but remain brittle and fall short of human performance.