Pith. sign in

REVIEW 3 cited by

Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.19619 v2 pith:OV57GA6A submitted 2023-10-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords situatedevaluationholisticlandscapellmsmachinemodelsposition
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large Language Models (LLMs) have generated considerable interest and debate regarding their potential emergence of Theory of Mind (ToM). Several recent inquiries reveal a lack of robust ToM in these models and pose a pressing demand to develop new benchmarks, as current ones primarily focus on different aspects of ToM and are prone to shortcuts and data leakage. In this position paper, we seek to answer two road-blocking questions: (1) How can we taxonomize a holistic landscape of machine ToM? (2) What is a more effective evaluation protocol for machine ToM? Following psychological studies, we taxonomize machine ToM into 7 mental state categories and delineate existing benchmarks to identify under-explored aspects of ToM. We argue for a holistic and situated evaluation of ToM to break ToM into individual components and treat LLMs as an agent who is physically situated in environments and socially situated in interactions with humans. Such situated evaluation provides a more comprehensive assessment of mental states and potentially mitigates the risk of shortcuts and data leakage. We further present a pilot study in a grid world setup as a proof of concept. We hope this position paper can facilitate future research to integrate ToM with LLMs and offer an intuitive means for researchers to better position their work in the landscape of ToM. Project page: https://github.com/Mars-tin/awesome-theory-of-mind

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-Based Social Simulations Require a Boundary

    cs.CY 2025-06 conditional novelty 6.0 of 10

    LLM-based social simulations are scientifically useful only within boundaries set by behavioral variance, and current validation practice under-checks variance.

  2. Computer Vision Models Show Human-Like Sensitivity to Geometric and Topological Concepts

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Transformers outperformed young children and matched their error profile on a geometry odd-one-out task, while vision-language models underperformed vision-only models.

  3. Position: Theory of Mind Benchmarks are Broken for Large Language Models

    cs.AI 2024-12 conditional novelty 6.0 of 10

    The paper proposes that LLM theory-of-mind evaluation should measure functional adaptation to partners, not just literal prediction of their behavior, and shows the two can diverge sharply in simple games.

Pith tools