Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models in Theory of Mind Tasks

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2302.02083.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.02083 v7

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:11:57.606122Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

176
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7ebc297f-623f-474e-83f2-d993cace0173 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Evaluating Large Language Models in Theory of Mind Tasks

Reference 184

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:12:31.697625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:eadd526d329f3f552244c9d25e7386124e32867e15f8b59fcafd5265617966c9

Observation 83bf78cd-26e0-44dc-b9d9-1f57f27beeb2 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Evaluating Large Language Models in Theory of Mind Tasks

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:36:18.177714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:ec10b57ac5bf6a78d163f6217e8d66e53bf0e6e98fe517850fa8c646903f4891

Observation 60fe1ccc-95a3-41cf-b7f5-c3cfd7f4ada9 · inbound

SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering cites this paper.

SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering Evaluating Large Language Models in Theory of Mind Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:11:57.606122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:11:57.606122Z digest=sha256:3662f5ae81cd8f069c728ff9a73f7dd7e3a614c9d7614508fa8c1802270abba4

Observation 439ce8ad-d210-49df-999e-122040f2c452 · inbound

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments cites this paper.

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments Evaluating Large Language Models in Theory of Mind Tasks

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T11:57:16.229278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:57:08.314088Z digest=sha256:aad9fc85699ded01b99876119bd679c1efdbfa754d12cfac6686d1f1130fcbb8

Observation b8e7fb44-2314-44bf-b0f9-3eba03df5ff4 · inbound

XToM: Exploring the Multilingual Theory of Mind for Large Language Models cites this paper.

XToM: Exploring the Multilingual Theory of Mind for Large Language Models Evaluating Large Language Models in Theory of Mind Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:24.257851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:24.257851Z digest=sha256:b0dca5e3609347a11a97671712a6cec40d3e20903c634add80749c58ea71cb8f

Observation 818afb4d-91ae-48c4-a0b6-0061a4073722 · inbound

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models cites this paper.

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models Evaluating Large Language Models in Theory of Mind Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:28:24.444328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:28:24.444328Z digest=sha256:834a11411d3318fcb63d65d95f1e5ca0435af2eeef3aa82541a8385867a63da8

Observation b5176ce3-a507-4e6e-93cc-7aa339c8e6fd · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Evaluating Large Language Models in Theory of Mind Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.859398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.859398Z digest=sha256:4b64ca30d2f68a763e637d7c9d07feababa8872fea1bb77f1c42a52fceaa0a53

Observation 3db051cc-f060-4333-ac85-9e32efa28e73 · inbound

A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools cites this paper.

A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools Evaluating Large Language Models in Theory of Mind Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:52.871383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:45:52.871383Z digest=sha256:d9fe56a520816edfa95104161505459199a374ecc3af6914af6deb26d80d6daf

Observation 9a69d7a8-8997-4c9d-955a-dc06a9ba2835 · inbound

Synergizing Logical Reasoning, Knowledge Management and Collaboration in Multi-Agent LLM System cites this paper.

Synergizing Logical Reasoning, Knowledge Management and Collaboration in Multi-Agent LLM System Evaluating Large Language Models in Theory of Mind Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:04.214279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:39:04.214279Z digest=sha256:a97a5498e1ff8cede138b94dd933bec1b00f68a6c9e2b5a708faa0a5872ba581

Observation 87536116-4f1f-4bf0-bdc7-99cd2b88fb1f · inbound

Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration cites this paper.

Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration Evaluating Large Language Models in Theory of Mind Tasks

Reference 7

Resolution
metadata mismatch
doi, observed 2026-05-19T07:27:08.601964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T07:23:56.627581Z digest=sha256:7b21ad3b7eb1e165a0445ea26fd1f79272666cc7fe693f55df86eb7a60da6fed

Observation a7d67963-12f6-4b35-8730-ccb355fcadda · inbound

Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning cites this paper.

Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning Evaluating Large Language Models in Theory of Mind Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:56.921317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:56.921317Z digest=sha256:a6808e2cf190fa4de74d2e52e3188eb3f83c7b53b8e53717802ae15fee317cb3

Observation aa49467c-29a1-4cfd-aefe-e89034e35f02 · inbound

Referential ambiguity and clarification requests: comparing human and LLM behaviour cites this paper.

Referential ambiguity and clarification requests: comparing human and LLM behaviour Evaluating Large Language Models in Theory of Mind Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:18.429275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:18.429275Z digest=sha256:3c751aed1af4ad859e2ab820f1fad386287de20b9b0f7a091cfe514cc48f1128

Observation 1e1dbb0e-a190-4d4b-b6bb-4a1d5eb046ee · inbound

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning cites this paper.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Evaluating Large Language Models in Theory of Mind Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.278651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.278651Z digest=sha256:454bac44fd3979a695bcf0dbb4712291d68b63ac527f189315e7c8c700c20e37

Observation 51ae5e3d-4528-42bc-aa53-5d6a2632d481 · inbound

Do Large Language Models Have a Planning Theory of Mind? Evidence from MindGames: a Multi-Step Persuasion Task cites this paper.

Do Large Language Models Have a Planning Theory of Mind? Evidence from MindGames: a Multi-Step Persuasion Task Evaluating Large Language Models in Theory of Mind Tasks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:21:18.736835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:21:18.736835Z digest=sha256:069fa186a6146cb010d8840f98e8eab7a87048703d37f1b375e09cd70f49ad61

Observation c4da330a-b34e-4de7-b732-7b850c1a77d2 · inbound

SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models cites this paper.

SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models Evaluating Large Language Models in Theory of Mind Tasks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:42.724363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:42.724363Z digest=sha256:ce72269bc37289e70dc6cecb946f778158698a36a6e49124bf8669350ad65c0f

Observation 8efacc87-1684-4144-ad46-b7fe1dae4705 · inbound

The Incomplete Bridge: How AI Research (Mis)Engages with Psychology cites this paper.

The Incomplete Bridge: How AI Research (Mis)Engages with Psychology Evaluating Large Language Models in Theory of Mind Tasks

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-06T11:19:15.799794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:19:15.799794Z digest=sha256:c5c9c4c83ed74461fe666669df425f986f51007a1fc55c3df26335c0fb898cfc

Observation 65b3c992-78f1-4148-a060-f5a327d8142f · inbound

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue cites this paper.

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue Evaluating Large Language Models in Theory of Mind Tasks

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-05T11:46:53.206198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:46:53.206198Z digest=sha256:17cd0e0aa96474ed90a5a561a4042f5a7b6d82a38ea3396f9e14bc2e0f27670a

Observation 2ac3479e-559f-4eb7-959a-cad129263656 · inbound

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents cites this paper.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents Evaluating Large Language Models in Theory of Mind Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.423054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.423054Z digest=sha256:fc09c1c325ca2972b0994ce95e1440c87928b8049447d2df432ed432eedf941f

Observation 3b59a820-a089-44ad-9167-90df738e3c22 · inbound

One Model, Two Minds: A Context-Gated Graph Learner that Recreates Human Biases cites this paper.

One Model, Two Minds: A Context-Gated Graph Learner that Recreates Human Biases Evaluating Large Language Models in Theory of Mind Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:21:50.462570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:21:50.462570Z digest=sha256:86fd56d54953086cb0b4943230b4a827eba33c5a5513c436f8bddb727d230c41

Observation e61906d8-b39e-4735-a074-1cc1184d5f06 · inbound

Designing Psychometric Bias Measures for ChatBots: An Application to Racial Bias Measurement cites this paper.

Designing Psychometric Bias Measures for ChatBots: An Application to Racial Bias Measurement Evaluating Large Language Models in Theory of Mind Tasks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:12:51.799106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T22:12:15.876447Z digest=sha256:9b62b648c95814811f2fcccf8da058e6ae491fc300278ba73ef7df5b07b41869

Observation 6534f454-410e-42ce-b984-86263344e4fa · inbound

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About? cites this paper.

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About? Evaluating Large Language Models in Theory of Mind Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:42:03.440313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:42:03.440313Z digest=sha256:9b1d0e7a214d2e311b32273cb44235ede820163e1ba19cb3e39e4c7ff02c42ab

Observation 9fa7ea78-7e23-4709-9fdc-7cc218331595 · inbound

Tacit Coordination of Large Language Models cites this paper.

Tacit Coordination of Large Language Models Evaluating Large Language Models in Theory of Mind Tasks

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T07:22:45.534266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:22:45.534266Z digest=sha256:55714584ac586e0c6ba16adcf5b186b2bca48fade272ef85224098ab657eaaa8

Observation da056b69-54ac-4304-8a9c-54bad4205d8f · inbound

Holos: A Web-Scale LLM-Based Multi-Agent System for the Agentic Web cites this paper.

Holos: A Web-Scale LLM-Based Multi-Agent System for the Agentic Web Evaluating Large Language Models in Theory of Mind Tasks

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:20:57.917204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T13:19:07.715653Z digest=sha256:188430f1ed12cce8e03459991510ab35e96f58b81b4ee03aeb4091fafeb5d041

Observation ae4278fe-fc73-4da0-bd18-f87ac078a6c0 · inbound

Gradual Cognitive Externalization: From Modeling Cognition to Constituting It cites this paper.

Gradual Cognitive Externalization: From Modeling Cognition to Constituting It Evaluating Large Language Models in Theory of Mind Tasks

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:48.069309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:33:50.078604Z digest=sha256:bd0d9351f052a96418b62b23bd90aeff8b050c9346244fbb4db6add1841563fa

Observation 69956e50-598b-462a-8ec1-4cf13063cfcc · inbound

Dynamics of Cognitive Heterogeneity: Investigating Behavioral Biases in Multi-Stage Supply Chains with LLM-Based Simulation cites this paper.

Dynamics of Cognitive Heterogeneity: Investigating Behavioral Biases in Multi-Stage Supply Chains with LLM-Based Simulation Evaluating Large Language Models in Theory of Mind Tasks

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:11:20.559013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T06:06:48.422940Z digest=sha256:bcb923832192266489cd5735b574fd79fb79ede7104a49cd8c0d7622d3ecb49f

Observation 27781969-6e61-4d58-98c2-00c8f521e710 · inbound

Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents cites this paper.

Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents Evaluating Large Language Models in Theory of Mind Tasks

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:11:05.456589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T23:17:21.835439Z digest=sha256:266964b4c40dfcdeefc38aef935b90f204c404bd69f7d78789d39f0e0dc8c9a5

Observation 1cae9632-b97b-423a-ad37-541a76ead2be · inbound

Don't Make the LLM Read the Graph: Make the Graph Think cites this paper.

Don't Make the LLM Read the Graph: Make the Graph Think Evaluating Large Language Models in Theory of Mind Tasks

Reference 3

Resolution
metadata mismatch
doi, observed 2026-05-08T21:44:17.760428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T11:41:37.667121Z digest=sha256:fd48d3919857f9472a7ef2f60b1a9c792f62cc50a27de871622338e536224806

Observation 9840fad6-02c6-48e7-8c51-7f9b00268ce4 · inbound

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning cites this paper.

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning Evaluating Large Language Models in Theory of Mind Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:13.583239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T08:04:15.238840Z digest=sha256:71c3133757c1ce3fdbfeb520a5a47a21eb9b0727554a72e49bc8669ea10de933

Observation c1629171-dab9-4425-8ed1-1df6d9c01445 · inbound

Truth or Tribe: How In-group Favoritism Prioritize Facts in Persona Agents cites this paper.

Truth or Tribe: How In-group Favoritism Prioritize Facts in Persona Agents Evaluating Large Language Models in Theory of Mind Tasks

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:08.654166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:41:05.398486Z digest=sha256:6bf017d52a02562c18fbb3da1e51a4046b268aa579533e38bf7ca0b81c471280

Observation 22ca6735-abd6-44bf-8b26-e1bdb4dff4d2 · inbound

The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment cites this paper.

The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment Evaluating Large Language Models in Theory of Mind Tasks

Reference 199

Resolution
metadata mismatch
doi, observed 2026-05-11T01:55:50.908671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:54:49.131461Z digest=sha256:bf270f25547981160c9d4ee9d655d58be6c2deae47e0497e7ad70fcc8fc3232f

Observation 35043146-8204-41b9-a520-3ac5905ff0ac · inbound

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents cites this paper.

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Evaluating Large Language Models in Theory of Mind Tasks

Reference 30

Resolution
metadata mismatch
doi, observed 2026-05-12T05:01:20.886600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:59:52.659739Z digest=sha256:761ab6e6abb66360159c2ff002834f6ebc0d5bb24cef31df2f51a4d5ed44fe50

Observation 50593306-9772-43de-bc61-32be8d8b2cae · inbound

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents cites this paper.

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Evaluating Large Language Models in Theory of Mind Tasks

Reference 30

Resolution
metadata mismatch
doi, observed 2026-05-20T23:29:12.583601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T23:27:45.394730Z digest=sha256:d33888b2d14692e6358c4612e618cea6de37f7edac312d708dd72fed472c5b78

Observation 4c123530-e486-449c-967e-0a5232c58e65 · inbound

What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code cites this paper.

What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code Evaluating Large Language Models in Theory of Mind Tasks

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:08:05.084361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T05:06:46.360174Z digest=sha256:a6f99d29c17d0621c4c1aa3d902e3088089dab33a91aff7f7c32120d4f991d47

Observation 986fe9de-1b21-47b2-93e9-f682ad720dd5 · inbound

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind cites this paper.

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind Evaluating Large Language Models in Theory of Mind Tasks

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:09:46.284684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:09:37.399954Z digest=sha256:ea3c2226229885925fd92ca6915e689e7986bf6f05e36f78fcfa5e2dc7b325dd

Observation 966f203d-06b8-4cae-ad0b-5488cb8a1429 · inbound

AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence cites this paper.

AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence Evaluating Large Language Models in Theory of Mind Tasks

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:54:45.612290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T08:53:21.922481Z digest=sha256:05c0ec7e25a0114c9075c0213ff4843aa37f0fff58d2b20128e6f723fdccd861

Observation f135b0a1-233c-4de2-863b-acb5a2b5b6aa · inbound

AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence cites this paper.

AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence Evaluating Large Language Models in Theory of Mind Tasks

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:04:56.568411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:02:35.794126Z digest=sha256:41ccd8f05b52b8ef65d2ae2bf0972b80593050e5d7ba6cc561472e94b9d897e4

Observation e791bbf4-78ef-4352-a83d-62345b5cff33 · inbound

Evaluating Large Language Models in a Complex Hidden Role Game cites this paper.

Evaluating Large Language Models in a Complex Hidden Role Game Evaluating Large Language Models in Theory of Mind Tasks

Reference 9

Resolution
metadata mismatch
doi, observed 2026-05-25T00:40:08.218573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:36:49.353675Z digest=sha256:7fe26635ecf57bca675100e3dc2abe150a21a4c4f877886fe636d62a9c861968

Observation ecab5c85-5f2a-4334-a661-3994de44ea3f · inbound

Voluntary Collusion with Secret Tools in Competing LLM Agents cites this paper.

Voluntary Collusion with Secret Tools in Competing LLM Agents Evaluating Large Language Models in Theory of Mind Tasks

Reference 10

Resolution
metadata mismatch
doi, observed 2026-06-29T17:03:40.248399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:01:22.732390Z digest=sha256:4781876ca66a24968664292203b3b2e68f2c592f9ef150886f0732db3fc3dbfd

Observation b657bf33-1ed2-4e48-8b6b-d119c3727766 · inbound

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents cites this paper.

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents Evaluating Large Language Models in Theory of Mind Tasks

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:56.986554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:07:07.135522Z digest=sha256:691706483c1876530d81fc7818bdbb379e6411a60d0215a7cb485c9035c627a3

Observation 2d1685c7-b646-4a5d-881c-45cf02d2f7c9 · inbound

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning cites this paper.

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Evaluating Large Language Models in Theory of Mind Tasks

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:28.251048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:42:38.122144Z digest=sha256:6ea471070dafcb41b7fc5e55cbccdd246fd9f66ae972ab443f718e374516327d

Observation d9c23742-28e3-48dd-a66c-20ddd3e4b0ac · inbound

A Survey of Large Language Models for Perception and Measurement of Human Psychology cites this paper.

A Survey of Large Language Models for Perception and Measurement of Human Psychology Evaluating Large Language Models in Theory of Mind Tasks

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:04:57.144776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T16:59:25.825681Z digest=sha256:498146889845544dde6ecc2bb1461acb60b37ca426bca9b932021020bb17803f

Observation eb0b92dc-7025-4496-a19e-293b60882bf3 · inbound

Embodied Explainability and Ontological Obstacles: Why We Struggle to Explain the Answers of Large Language Models (LLMs) cites this paper.

Embodied Explainability and Ontological Obstacles: Why We Struggle to Explain the Answers of Large Language Models (LLMs) Evaluating Large Language Models in Theory of Mind Tasks

Reference 72

Resolution
verified exact
doi, observed 2026-06-26T21:30:03.725211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T06:54:48.941504Z digest=sha256:d274505bbd1ec7e2b1d14512d4e4a53af597d9dd506618e8db75bf3fe357e287

Observation 19ab65e3-ada9-46b4-8062-35ecadf5645e · inbound

ToxiREX: A Dataset on Toxic REasoning in ConteXt cites this paper.

ToxiREX: A Dataset on Toxic REasoning in ConteXt Evaluating Large Language Models in Theory of Mind Tasks

Reference 187

Resolution
metadata mismatch
doi, observed 2026-06-29T04:43:07.126975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T04:33:18.794505Z digest=sha256:dbb3885801f3850f284d471a8b55b2d6d9529812df60cd294dd7c39f653b546c

Observation a06b16e8-868d-4adb-af30-780ec62111cc · inbound

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models cites this paper.

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models Evaluating Large Language Models in Theory of Mind Tasks

Reference 62

Resolution
metadata mismatch
doi, observed 2026-06-30T01:34:08.777412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T01:32:40.504759Z digest=sha256:57d4c64e3e3fb606813afb0096dd88f234fd0862affd2f95fd1070a1f45b83f8

Observation c8683c41-d09b-4a34-ba3a-335b740fdc83 · inbound

Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents cites this paper.

Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents Evaluating Large Language Models in Theory of Mind Tasks

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:20.142185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T06:53:39.502788Z digest=sha256:d781928eef07560d2f3bfde9101447d4f54fd94954f6c2b445a380ecdf2d0154

Observation da795e3f-c986-46e1-89fa-c232d5ec6918 · inbound

Beyond Skepticism: Evaluating LLMs Pedagogical Intent Reasoning with the Adaptive Pedagogical Vigilance Framework cites this paper.

Beyond Skepticism: Evaluating LLMs Pedagogical Intent Reasoning with the Adaptive Pedagogical Vigilance Framework Evaluating Large Language Models in Theory of Mind Tasks

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:28:33.613030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T15:26:40.974564Z digest=sha256:23ca55336fa1a2522ef66fdb4a4b700d7bd1beaac08cdf5e833b1d44adb77c48

Observation 84814518-9829-4235-b250-ee316385be28 · inbound

AgentSociety 2: An Integrated Research Environment for Executable Social Science cites this paper.

AgentSociety 2: An Integrated Research Environment for Executable Social Science Evaluating Large Language Models in Theory of Mind Tasks

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-02T11:43:50.574560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:43:50.574560Z digest=sha256:9a0fb8522b2379beeb2e7f72866262f1854f9a44683189a7a05041901ce22857

Observation 0e8c9ee2-1ef8-43fc-ab00-8086c5a629c2 · inbound

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation cites this paper.

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation Evaluating Large Language Models in Theory of Mind Tasks

Reference 164

Resolution
unresolved
no resolver link, observed 2026-07-31T21:55:17.619128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T21:55:17.619128Z digest=sha256:93b15c9f9c1f94068552fd231e50b6b9a2dcb175412dc99fbb5bf7621e135df2

Observation 74b814d5-e1b2-4821-8e66-d3fecded039f · inbound

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning cites this paper.

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning Evaluating Large Language Models in Theory of Mind Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:56.756766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:57:56.756766Z digest=sha256:7644256573b61fdd5c7e4e5c63400180010268ae70c0764908f088c90a9ff68b