Pith. sign in

hub

Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

31 Pith papers cite this work, alongside 79 external citations. Polarity classification is still indexing.

31 Pith papers citing it
79 external citations · external index

hub tools

citation-role summary

background 3 contradiction 1

citation-polarity summary

polarities

background 3 contest 1

representative citing papers

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs

cs.CL · 2026-06-26 · unverdicted · novelty 7.0

Extending Werewolf with a Jester faction whose win condition inverts suspicion reveals that LLMs frequently fail at triadic incentive reasoning, with Jesters winning 60-70% of games while wolves make self-defeating early votes.

ProactBench: Beyond What The User Asked For

cs.LG · 2026-05-09 · unverdicted · novelty 7.0

ProactBench measures LLM conversational proactivity in three phases using 198 multi-agent dialogues and finds recovery behavior hard to predict from existing benchmarks.

Bayesian Social Deduction with Graph-Informed Language Models

cs.AI · 2025-06-21 · unverdicted · novelty 7.0

Hybrid Bayesian-graph LLM agent reaches competitive performance against large models and achieves 67% win rate against humans in controlled Avalon play, outperforming baselines and human teammates.

GAIA: a benchmark for General AI Assistants

cs.CL · 2023-11-21 · unverdicted · novelty 7.0

GAIA benchmark shows humans at 92% accuracy on simple real-world questions far outperform current AI systems at 15%, proposing this gap as a key milestone for general AI.

MindZero: Learning Online Mental Reasoning With Zero Annotations

cs.AI · 2026-05-29 · unverdicted · novelty 5.0

MindZero is a self-supervised RL framework that trains MLLMs for online Theory of Mind reasoning by rewarding mental-state hypotheses that best explain observed actions via a planner, then distills this into fast inference.

citing papers explorer

Showing 31 of 31 citing papers.