Pith. sign in

Paper Citation Record · LEDGER

Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 54 inbound Pith citation observations for arXiv:2302.08399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.08399 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 54 of 54 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:22:45.056255Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

79
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 74814fe1-819c-4f67-a8b1-7db06893d7b9 · inbound

GAIA: a benchmark for General AI Assistants cites this paper.

GAIA: a benchmark for General AI Assistants Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:46:03.367546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T15:46:03.247029Z digest=sha256:beda38e10062f6d6bfd26ab13b99fe1a48353ed416a4c55706c8ed9b5361629b

Observation 1649c9ec-82ab-4eb1-a22c-e240d09e33e8 · inbound

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks cites this paper.

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T15:22:45.056255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:22:45.056255Z digest=sha256:68c24766e375d2f5a2d2637c2f215c4c69b8a59b9e12901ed7e1bb55b9463ee4

Observation 2338b580-a693-4d3d-a534-5ebe96e9046b · inbound

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning cites this paper.

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:15:09.382305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T21:14:13.351140Z digest=sha256:dfbdf7118238b2ca1a8fdead2df72315b8d34a3b4b9f208390f10885a4a4520f

Observation 7db379fd-99a6-4e02-8f69-2e81ad8da768 · inbound

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States cites this paper.

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:29.153547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:46:29.153547Z digest=sha256:4293b3099776d715fe5b518954694ff51402d6309a5b072e4af1137e6f30ece2

Observation 8277298c-d1ee-4270-87dc-b7bc6b8f79f9 · inbound

Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes cites this paper.

Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:14.922995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:44:14.922995Z digest=sha256:a774677601ce4bd8d342710bc3add72e57510c94724de4a4a1ecde0cbcb359fa

Observation 1c908b8d-0609-4f01-afd6-24c7428bea7e · inbound

Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation cites this paper.

Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:25.885019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:57:25.885019Z digest=sha256:c3bf54db0e56a8e243c41489ef98fb242d7d01c63ed7226f713e7f1be37eb8fc

Observation 54814bd8-0158-4116-890a-620cbc0a365d · inbound

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs cites this paper.

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:51:30.735912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:51:30.735912Z digest=sha256:18ff96e324ee600e69ae7e42f2e996a09da91a4975edc37f16fa61818f49502b

Observation 4f45d0ac-3793-4ce4-b70e-57931e169a76 · inbound

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models cites this paper.

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:28:26.225894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:28:26.225894Z digest=sha256:389a8276972ad7b167b19553587e5bfeb5f1123ce45c52354be62d7302d41385

Observation f6f8376c-1c0d-4adf-a669-f81a9d92728e · inbound

Bayesian Social Deduction with Graph-Informed Language Models cites this paper.

Bayesian Social Deduction with Graph-Informed Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:32:09.406955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T07:27:20.027179Z digest=sha256:051bc613c0ac134f1e6338f5bd2042964119233f99ca9b55dc18a2842013ed02

Observation 689c85c9-638d-403a-a4bb-8fd9186afa97 · inbound

Mechanistic Interpretability Needs Philosophy cites this paper.

Mechanistic Interpretability Needs Philosophy Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.392344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T23:49:19.683025Z digest=sha256:fd2da76d3226b6ccee5e9964078c260bdca6c528a6230619783fde1d4d6462b2

Observation a122cc52-9969-4734-9b7d-e596807e2cf7 · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.394970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.394970Z digest=sha256:ef59af66f0dce9ed807f24fff8047ded7fb001508d3a5888250ca131c9bf8129

Observation afb47a2e-fa82-44b8-89da-0efbba159c7d · inbound

Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration cites this paper.

Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.124597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T07:23:56.627581Z digest=sha256:4210f452113a5ac83479d842431be222b14aa5007c31bb0c7cffbb8dafc057ed

Observation 3f912863-fd85-44ee-b125-cddc8c6d7130 · inbound

Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning cites this paper.

Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:57.810009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:57.810009Z digest=sha256:3de5a33b886d2418cc4e119464fbf99433968862fd641ba157be0618abe28d99

Observation 53995b7b-2978-407b-86c7-8db638d4f317 · inbound

Strategy Adaptation in Large Language Model Werewolf Agents cites this paper.

Strategy Adaptation in Large Language Model Werewolf Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:43.283347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:42:43.283347Z digest=sha256:20f5b50c38a05c291fdf6aff4514cc0c394101b15a67cecb7bd883ca0fd59de7

Observation 922a93a1-2cda-4220-8252-7824287ecfff · inbound

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning cites this paper.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.639155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.639155Z digest=sha256:7527302ccb00db3f881f923e0ac8f3cb91cc5ba12b8c93435130da9e679911ce

Observation d998f9ac-10d9-484d-af9f-ab6379ed77fd · inbound

Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition? cites this paper.

Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition? Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T05:53:21.310753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:53:21.310753Z digest=sha256:ebb0576fa66d0fc8e96b90dd052a1bcf9d9b6b91752c44ca57f3b639301b699c

Observation c3e10985-b2d6-494f-b893-85e0af453ac7 · inbound

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies cites this paper.

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:31:29.424187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:31:11.974395Z digest=sha256:66aaed794f135433bb15a642f16d8def227cea8eaece66e79b862c6d536e12f2

Observation 989d248c-1087-4cf8-9957-3f2370b93312 · inbound

Gradual Cognitive Externalization: From Modeling Cognition to Constituting It cites this paper.

Gradual Cognitive Externalization: From Modeling Cognition to Constituting It Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:48.086520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:33:50.078604Z digest=sha256:35acaf05f514922f17b45fb909fa2bb23a8c0ebf3b728c982ac9d5faee7452df

Observation 8a0be8ec-6c49-4634-9eaa-74e7ef352908 · inbound

Network Effects and Agreement Drift in LLM Debates cites this paper.

Network Effects and Agreement Drift in LLM Debates Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:46:08.057855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:50:20.833398Z digest=sha256:73a223a2b2dd8abb42fee46cf296aa7036b8476868d3f5331e1380261df857e2

Observation aa691441-60a7-42d2-a723-fa4ff1cfaab6 · inbound

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding cites this paper.

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:01:49.323113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T06:58:19.094492Z digest=sha256:9a7e305d68b968a5ea8d5f11e36863a4ee594ff6e2dc73fc6694fc6aea5f24f0

Observation e6195f40-4131-4d09-bae1-327a0408f14f · inbound

Impact of Task Phrasing on Presumptions in Large Language Models cites this paper.

Impact of Task Phrasing on Presumptions in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:47:11.361474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T19:14:04.840446Z digest=sha256:8a14b72c116dd7dff539ae7815e1669ba7a24f33476bd6750f19296664c77b8d

Observation a47a5663-6d9e-4e43-934a-5a7cc0bae300 · inbound

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior cites this paper.

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T02:24:40.718851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T15:36:36.257870Z digest=sha256:6209f8d76a43e46a912e7d1727a96375f74ee5a8f4e9a11198365c77e8afffc5

Observation b1b96d59-26d0-4dcb-8ee1-76d52ab7bcd9 · inbound

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior cites this paper.

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T18:28:58.421710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T18:26:14.265380Z digest=sha256:c512102885be11bf11200c1da7022710d212af681792caa295015951e60ecc09

Observation c698d33d-198d-4019-b869-0e9a6f28604c · inbound

ProactBench: Beyond What The User Asked For cites this paper.

ProactBench: Beyond What The User Asked For Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:16:16.052376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T02:14:01.145443Z digest=sha256:35d380ad361baf16501573fc118321f8df9795a0a699b5de689bad1dc5a4be37

Observation 910cca44-b35e-4301-8c36-71eeaef84d61 · inbound

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents cites this paper.

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:27.203788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:59:52.659739Z digest=sha256:69affd2622237417cb2502c5be2f9e85694d9802c968e504bb9e7523cc57e0f8

Observation f2bcdd2b-bc34-427e-bc2c-433fee88eb68 · inbound

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents cites this paper.

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:29:12.834145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T23:27:45.394730Z digest=sha256:c10ed63296667d8d55034464a194c0299da28acb8983f141163f586ec87c31d0

Observation 705ce2c7-cab5-4b35-9270-5e97b2bb7a40 · inbound

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue cites this paper.

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T21:39:03.445907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T21:36:51.291779Z digest=sha256:04f2c6e9348e3ae9487131d8cc4731cfe5db236560e569456bf1babc18007c96

Observation 64bedbf5-ef58-4daf-9b23-796ad86f1c62 · inbound

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance cites this paper.

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:32:49.729980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T22:29:17.058647Z digest=sha256:52fe752d148897350b7dd9c63fcef9917ffa15969a6f3653b8f73e90f07b8b4d

Observation ec8f5611-4a4a-4e04-b445-5b83bcdfc64a · inbound

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks cites this paper.

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:13:11.895524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:10:31.059095Z digest=sha256:3804c546b159286f2bff47256ea80bc5a1d9891964ce9850c1c7d1fd980847f3

Observation 07044555-5fc0-426b-a02a-1543268c205e · inbound

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind cites this paper.

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:09:46.264689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:09:37.399954Z digest=sha256:a4428c8bc9f04e4cb85c4df41e96ea96d627912c05991123a0645aea6ef292ef

Observation c199515d-e77c-4adc-b2b8-d2da8e10ee66 · inbound

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models cites this paper.

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:45:20.656455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:41:35.532363Z digest=sha256:df9b80546ea4381a14da72dcc9b982b7fd21de1fbf80b0b4e62c8f298cd7d84d

Observation 33c1074d-40a5-4d69-b625-f1871ef006f2 · inbound

Voluntary Collusion with Secret Tools in Competing LLM Agents cites this paper.

Voluntary Collusion with Secret Tools in Competing LLM Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:03:40.780449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:01:22.732390Z digest=sha256:f5307fb0a5aba7832f5e265a9595c3d86190b7767dced63a5b2f30e03165bc5e

Observation 2c7cdacb-ec2a-40eb-a818-b7ef9bc8b574 · inbound

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs cites this paper.

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.213072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:15:27.939886Z digest=sha256:d07227d2471d48721ec6395c7274fd9c16efb95fd09179a670e2869d627be8d4

Observation 578d6264-66cf-4fb5-87eb-2151f969a46d · inbound

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models cites this paper.

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.961583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T07:17:07.720589Z digest=sha256:027343c6839398defba6ef901243bbb6eee7fe67912d1c92480db91ba1e26e9c

Observation b8572d77-fe79-49a8-b0f7-b8f84b5d9bab · inbound

MindZero: Learning Online Mental Reasoning With Zero Annotations cites this paper.

MindZero: Learning Online Mental Reasoning With Zero Annotations Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:56:10.637583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T21:57:17.194711Z digest=sha256:b7072f82a3315903a07a89c4ad24ca20379fb654a0d846cb0d7ba1e61bb08e6f

Observation c73e78d3-a925-4449-aed6-0fa229f085fe · inbound

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents cites this paper.

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:56.983506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:07:07.135522Z digest=sha256:faac739132d88e93e4162655648b48d2f8a54362734050f6eb3f936a9094f072

Observation 794d8033-5c5d-4ab4-b63f-2df312c469d7 · inbound

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning cites this paper.

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:28.207742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:42:38.122144Z digest=sha256:1c11615d68dabbd954b6e4c65592727dd16add40b0fc79f5bab4929b39ac6849

Observation ad45c732-6d1f-46d4-b57d-af9c6be58ae5 · inbound

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism cites this paper.

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.444533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:39:02.190984Z digest=sha256:f618e8042e3261445b219dfddc7e0ce14a60420ff464db8a945ce062db84e2b6

Observation c6e2bd2b-00cc-485f-a59f-eebca50ca35b · inbound

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism cites this paper.

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T11:47:18.453636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:47:18.453636Z digest=sha256:31b9ae4b2cc4fcfc4556452fb43cbeb3ca8c195c6af8b93f664540530eba7ad0

Observation d0716f81-f13a-455e-8c32-d9352e95038c · inbound

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning cites this paper.

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:32.702392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:44:31.919126Z digest=sha256:008980c548201f6c01026e36a5e9e0efbcc33ebee776325482d2f2790b1563f8

Observation 1d834461-315d-4c78-96e1-7941103a4f1c · inbound

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence cites this paper.

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T11:11:46.562648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:11:46.562648Z digest=sha256:11c9cd37388ade288b66cf2dd38c5566149b7590ea6f2117538713ac28664f46

Observation 3e6295fa-ac46-410a-bd95-8f5575fd0c36 · inbound

A Survey of Large Language Models for Perception and Measurement of Human Psychology cites this paper.

A Survey of Large Language Models for Perception and Measurement of Human Psychology Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:04:57.157095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T16:59:25.825681Z digest=sha256:44eb350e716cefad2c188f776a06416d5d77b8ff84e839b3a15fa2524537e1c1

Observation ec190ba6-5f1c-4c6d-a5a6-a1d2387351bf · inbound

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure cites this paper.

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:45.975167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:10:46.067374Z digest=sha256:d5c4ab26190be20f7569d621b3094f6db70f39f2e8e9602d4bd14e797647bb81

Observation c145cecc-70a6-4ffd-a569-88c561c2dc89 · inbound

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs cites this paper.

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:13:53.392659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T04:48:58.181882Z digest=sha256:70cdd1409707d0e9a7d8a82377b526c0c30fb68f5655f4a7677590f5dfe56359

Observation 9d1084a6-4b77-4b2a-8fa4-6cff5ffc887a · inbound

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models cites this paper.

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-30T01:34:08.819629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T01:32:40.504759Z digest=sha256:0a418e64118ad47ff0e78536d25d00b7bd44ef18862fde6b12ddc75e8dc915aa

Observation 83cdd38f-ff5c-4b51-a867-55f095bd3ad4 · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.709551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:a43423935d4c84805ee707bb23a513db45d6a7a81a3a0a4912fdbe9b162ff399

Observation 22d2a7af-efa9-4cad-abcc-2ac87a1b027b · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T10:15:59.479435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:15:59.479435Z digest=sha256:a5ce17e54cbd1e30193cf7a04b9776aeff29c68deab66c112370382cc727ed4c

Observation 86c5abea-47e1-4765-9ee2-fadab25216b8 · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T07:15:01.579186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:15:01.579186Z digest=sha256:be2965fb4f0bb0fdc4a7fdd8ca735eb98708383c8e3dd0da4f2975424fb90eac

Observation 8276b524-31b1-4e7c-baa1-268296de9bca · inbound

Belief-reality separation lives in routing over a shared value slot in language models cites this paper.

Belief-reality separation lives in routing over a shared value slot in language models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T07:25:57.402299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:25:57.402299Z digest=sha256:1a20ebe096fc8a5f61851f7b1ca0dacd32b27cbadeed14e5f4c2f06d36396895

Observation 60ace4c0-58ee-40fa-a436-d4e569ff3a88 · inbound

The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt cites this paper.

The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T02:43:30.604719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:43:30.604719Z digest=sha256:c1c4ec78d9897d664936ed56e6bf49f6273cdbe81706376ca8d6ca764bd3e577

Observation 4429bffe-0e6e-47c8-afce-fa78f5848f81 · inbound

Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments cites this paper.

Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:46:40.996070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:46:40.996070Z digest=sha256:285c40b8a9745f4683bb519c41430e74c59ceab0cd1f2edf91ec2cfc61880452

Observation b915499d-f861-40a8-a962-dcbf650d71d8 · inbound

Perceived AGI: Believability as Dimensional Completeness, Not Capability cites this paper.

Perceived AGI: Believability as Dimensional Completeness, Not Capability Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T22:09:15.497907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:09:15.497907Z digest=sha256:0f35e0e424c3dfe41f887a5cd02b3c591e291b13f1dac1437cfaa9992e3e8634

Observation 86e59290-b4d6-4b56-8747-c00352253314 · inbound

Mental World Modeling cites this paper.

Mental World Modeling Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-30T11:07:38.427392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:07:38.427392Z digest=sha256:e72a13a74bd86dc9810b24d698ad165224be78abddc32828004fb47dec1e2399

Observation efd59943-ef7c-46cf-a890-1f6a849ee043 · inbound

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning cites this paper.

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:58.012382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:57:58.012382Z digest=sha256:98c17c9a0accb073a360fca9a98f53134d1e8ee94acdcb43811e4f15390fde33