Pith. sign in

REVIEW 11 cited by

How FaR Are Large Language Models From Agents with Theory-of-Mind?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.03051 v1 pith:T7RMH2T3 submitted 2023-10-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords inferencesmodelsllmsactionactionsmentalotherstates
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

"Thinking is for Doing." Humans can infer other people's mental states from observations--an ability called Theory-of-Mind (ToM)--and subsequently act pragmatically on those inferences. Existing question answering benchmarks such as ToMi ask models questions to make inferences about beliefs of characters in a story, but do not test whether models can then use these inferences to guide their actions. We propose a new evaluation paradigm for large language models (LLMs): Thinking for Doing (T4D), which requires models to connect inferences about others' mental states to actions in social scenarios. Experiments on T4D demonstrate that LLMs such as GPT-4 and PaLM 2 seemingly excel at tracking characters' beliefs in stories, but they struggle to translate this capability into strategic action. Our analysis reveals the core challenge for LLMs lies in identifying the implicit inferences about mental states without being explicitly asked about as in ToMi, that lead to choosing the correct action in T4D. To bridge this gap, we introduce a zero-shot prompting framework, Foresee and Reflect (FaR), which provides a reasoning structure that encourages LLMs to anticipate future challenges and reason about potential actions. FaR boosts GPT-4's performance from 50% to 71% on T4D, outperforming other prompting methods such as Chain-of-Thought and Self-Ask. Moreover, FaR generalizes to diverse out-of-distribution story structures and scenarios that also require ToM inferences to choose an action, consistently outperforming other methods including few-shot in-context learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. The Essence of Contextual Understanding in Theory of Mind: A Study on Question Answering with Story Characters

    cs.CL 2025-01 conditional novelty 7.0 of 10

    A new benchmark tests LLMs on theory-of-mind questions about novel characters, showing that humans with book knowledge outperform the best LLMs.

  3. Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding

    cs.CL 2024-12 conditional novelty 7.0 of 10

    Friends-MMC is a multi-modal multi-party conversation dataset from the TV show Friends with face and speaker annotations, and baselines show that speaker identification benefits from combining visual and textual cues.

  4. Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

    cs.LG 2024-12 conditional novelty 7.0 of 10

    ExploreToM uses A* search over a domain-specific language to generate adversarial theory-of-mind stories that make LLMs, including GPT-4o, score as low as 0% and 9%, and fine-tuning on the data lifts ToMi accuracy by ...

  5. CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new 100-image benchmark shows that state-of-the-art multimodal models detect subtle, socially meaningful visual clues at near-chance levels and hallucinate accusations on innocent images.

  6. From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Attention heads in multimodal LLMs linearly encode agents' beliefs, and steering those heads along probe-derived directions improves first- and second-order belief accuracy on the new GridToM benchmark.

  7. Where You Go is Who You Are: Behavioral Theory-Guided LLMs for Inverse Reinforcement Learning

    cs.AI 2025-05 conditional novelty 6.0 of 10

    SILIC uses LLM-guided inverse reinforcement learning and Theory of Planned Behavior chain reasoning to infer age, gender, income, and employment from travel trajectories, reportedly beating SVM, XGBoost, CatBoost, and...

  8. Effects of structure on reasoning in instance-level Self-Discover

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Unstructured natural-language reasoning plans outperform dynamically generated JSON reasoning plans in an instance-level Self-Discover framework, with relative gains up to 18.90% on MATH.

  9. Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning

    cs.AI 2025-07 conditional novelty 4.0 of 10

    An LLM-augmented Bayesian inverse planning model, LAIP, generates hypotheses and action likelihoods, then uses Bayes' rule to infer agent preferences, outperforming LLM-only baselines.

  10. A Layered Architecture for Developing and Enhancing Capabilities in Large Language Model-based Software Systems

    cs.SE 2024-11 conditional novelty 4.0 of 10

    A layered architecture with model, inference, and application layers, plus a capability-mapping process, guides where to implement features like structured output and domain knowledge in LLM systems.

  11. Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A thesis proposal repurposing two prior papers on LM agents for text games, framed as a path to theory-of-mind AI, with no new theory-of-mind evidence.

Pith tools