Pith. sign in

Paper Citation Record · LEDGER

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks

As of 10 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 5 inbound Pith citation observations for arXiv:2502.08796.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08796 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:41:25.883309Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T07:22:25.349398Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:14.268938Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca079f7b-05f2-4fdc-8ea2-f058cecdcc6a · outbound

This paper cites Who is Mistaken?.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks Who is Mistaken?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:41:25.839197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:41:25.839197Z digest=sha256:8d78aba9b9629ff3658e91e8ebb0e00635f01f50a952c920d761559534f17962

Observation d69feb74-5de3-4ce2-b52f-40cca0141306 · outbound

This paper cites Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:41:26.065083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.844911Z digest=sha256:cb71f24a7031b8e39378a8d17797e0fec318771a69005ae48f55590f7203ed43

Observation 5766bd49-031c-4d77-8ea7-165b610321ad · outbound

This paper cites an unresolved cited work.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:41:25.999510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.863301Z digest=sha256:66f393b02672a6f4ee8cad49a8f5fd4c294b2d6512491d3763083274586233ca

Observation 88b625b0-be49-42b9-ae4d-e2ac0208b656 · outbound

This paper cites In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:41:25.983246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.868031Z digest=sha256:1d315cd7484db0a26adbaeed0d81ef415c93137f578735e473148eeaf1c6cbe8

Observation a654277e-92ad-4a7d-91fe-b5aea22fbc2b · outbound

This paper cites john thinks that mary thinks that.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks john thinks that mary thinks that

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:41:25.967583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.872585Z digest=sha256:608fbdfc0c6af32f36cc6e329793985ed05c86d28d3b2f247f1f600ab9495953

Observation 28d1d3ea-edd4-448e-b1a2-29d55363add4 · outbound

This paper cites Verbosity Bias in Preference Labeling by Large Language Models.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks Verbosity Bias in Preference Labeling by Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:41:25.877634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:41:25.877634Z digest=sha256:14b25cf9a78644b1aa56c29e0b6e9721a96072d0eccad4e2111969cdb177d764

Observation b5da36a1-8bcf-4b9b-8e64-88d99e5bbb27 · outbound

This paper cites an unresolved cited work.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:41:25.950863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.883309Z digest=sha256:1b491d17c9417e5f8c0363a81e46d66755dc12c5bd301622dfd7e73b66204117

Observation 635658df-82c3-4eb8-8c6f-aabba11bd492 · outbound

This paper cites theory of mind.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks theory of mind

Reference 1985

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:41:26.112518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.824224Z digest=sha256:f5a2e900cdf7edf14a73fea623e6e60b55534143edfb105ef38c10ecd8dbc7ff

Observation 1de39887-0f2e-45f7-a99a-a8408a21541d · outbound

This paper cites Annals of the New York Academy of Sciences, 1167:103–114.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks Annals of the New York Academy of Sciences, 1167:103–114

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:41:26.032633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.854400Z digest=sha256:4a4228edffe7b166e8e6c4da6f19941721e081ca7d72f51195a7a53d7243f6f2

Observation 8c4842f1-6ab3-4250-a4bd-5a747f616e94 · outbound

This paper cites an unresolved cited work.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:41:26.015459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.858851Z digest=sha256:b241ecb0bf31fef30164342df74aebb06f73a16a9bb309ad33a0097cc9462af5

Observation 5042e66b-99d6-4b14-9f8e-259563d7306c · outbound

This paper cites Beaudoin C., Leblanc É., Gagner C., and Beauchamp M.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks Beaudoin C., Leblanc É., Gagner C., and Beauchamp M

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:41:26.081555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.833946Z digest=sha256:8a79dc40e2aa2eeb2385f87bd3689cc7b1e72d6c9f9dbfd6272f21767f0312cc

Observation 86702ea1-7808-4808-a233-abcec26de164 · outbound

This paper cites In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1112–1125.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1112–1125

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:41:26.127387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.818762Z digest=sha256:d0597fc4ccbaa1580846e69706257c0374851826b06a6009e04bed10f4ea16e9

Observation 32b8cf54-dc9d-4f69-a9f4-6afaa8b9fc0b · outbound

This paper cites Jan-Christoph Klie, Bonnie Webber, and Iryna Gurevych.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks Jan-Christoph Klie, Bonnie Webber, and Iryna Gurevych

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:41:26.049508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.850013Z digest=sha256:a1a854df23b3a5e7027bcef87288bf616d2a75bfe244956286e0618ad2d0e00c

Observation 4cf67cfb-9aaa-4181-be55-e6a80067ecc9 · outbound

This paper cites Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai.

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:41:26.097403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:41:25.828970Z digest=sha256:105f9d284c3e3408f33cf851dd8632b5fa74c87f6b98c767008eb160f7149fc5

Pith citing papers

Observation 8582d296-503b-4c2f-9de2-1188a5a317e3 · inbound

On Emergent Social World Models -- Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language Models cites this paper.

On Emergent Social World Models -- Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language Models A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:17:22.486999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:13:40.377973Z digest=sha256:f6e31553901c3dcd57b55a80de1c20a24dd2493fca416f8a9cf18a6452f86c84

Observation 456439ea-fb2f-4fcf-842a-9d31ce88f359 · inbound

Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations cites this paper.

Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:57:42.630853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T17:55:04.458343Z digest=sha256:48077c8680efaaafaef6372be27df4151c3d9c1e286aa759de7268eca05010d4

Observation aa045270-4f7a-4551-8351-c8cd9393df86 · inbound

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches cites this paper.

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks

Reference 244

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:14.270428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:35:01.285534Z digest=sha256:55e980f8680c99d6342b196e8bd00db889812d58ccc5c479b1a85ffaa83f8887

Observation 99b0aac6-1f44-4960-9505-0bd796011684 · inbound

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches cites this paper.

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks

Reference 260

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:35:34.033269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:22:25.349398Z digest=sha256:26970311b811d22bfd01a77ae093bfeaf1e2af242622d08ed4d76d74cc0ad3bd

Observation bbbaa650-2d14-487b-89f2-25e2a1c38aea · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:15:44.636103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:0f511a9e551a88d4d674473a39d118832feba9fa519e0c801233629dd44648d5