Introduces NCP-ExploreToM framework to evaluate LLMs on inducing belief states via planning and action, with GPT-5 succeeding on ~80% of tasks and outperforming humans.
arXiv preprint arXiv:2412.19726 , year=
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Dialogue among partially-observing LLM household robots reduces action conflicts 41–93 points yet lowers task success because hallucinated entity mentions cancel belief alignment.
Hybrid Bayesian-graph LLM agent reaches competitive performance against large models and achieves 67% win rate against humans in controlled Avalon play, outperforming baselines and human teammates.
LLMs excel at retrospective mental-state labeling on naturalistic dialogues but mostly fail a context-free prospective probe that maps isolated mental-state profiles to dialogue trajectories, despite expert 100% accuracy.
citing papers explorer
-
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
Introduces NCP-ExploreToM framework to evaluate LLMs on inducing belief states via planning and action, with GPT-5 succeeding on ~80% of tasks and outperforming humans.
-
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
Dialogue among partially-observing LLM household robots reduces action conflicts 41–93 points yet lowers task success because hallucinated entity mentions cancel belief alignment.
-
Bayesian Social Deduction with Graph-Informed Language Models
Hybrid Bayesian-graph LLM agent reaches competitive performance against large models and achieves 67% win rate against humans in controlled Avalon play, outperforming baselines and human teammates.
-
DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories
LLMs excel at retrospective mental-state labeling on naturalistic dialogues but mostly fail a context-free prospective probe that maps isolated mental-state profiles to dialogue trajectories, despite expert 100% accuracy.