Pith. sign in

REVIEW 3 cited by

HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16755 v1 pith:C7PPJU5K submitted 2023-10-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords higher-orderlanguagellmsmindtheorybenchmarkhi-tomlarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Theory of Mind (ToM) is the ability to reason about one's own and others' mental states. ToM plays a critical role in the development of intelligence, language understanding, and cognitive processes. While previous work has primarily focused on first and second-order ToM, we explore higher-order ToM, which involves recursive reasoning on others' beliefs. We introduce HI-TOM, a Higher Order Theory of Mind benchmark. Our experimental evaluation using various Large Language Models (LLMs) indicates a decline in performance on higher-order ToM tasks, demonstrating the limitations of current LLMs. We conduct a thorough analysis of different failure cases of LLMs, and share our thoughts on the implications of our findings on the future of NLP.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Social World Model for Lifelong Social Intelligence

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    The Social World Model supplies a five-dimension decomposition and closed-loop training loop that lets a 7B open model match Gemini 3 Flash on social metrics while showing zero forgetting on ASCENT-Bench.

  2. Reinforcing Human Behavior Simulation via Verbal Feedback

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    DITTO uses RL with verbal feedback to train LLMs for human behavior simulation, reporting 36% average gains over base models and outperforming GPT-5.4 on 6 of 10 SOUL benchmark tasks.

  3. Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization

    cs.CL 2025-11 unverdicted novelty 5.0 of 10

    MAGIC-HMO is a multi-agent framework that treats Chinese short-form creative NLG as heterogeneous multi-objective optimization over personalized constraints plus explanation reliability and outperforms baselines on a ...

Pith tools