Pith. sign in

REVIEW 3 major objections 4 minor 45 references

This paper claims that theory of mind should be a switch, not an always-on ability, and specifies the causal conditions that flip it.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 11:06 UTC pith:FWCWSH2U

load-bearing objection A useful theoretical scaffold for the 'when' of mentalizing, but the IA enabling cause is not enforced by Eq. 6 and the decision procedure is not yet resource-rational. the 3 major comments →

arxiv 2606.16944 v2 pith:FWCWSH2U submitted 2026-06-15 cs.AI cs.HC

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence

classification cs.AI cs.HC MSC 68T0191A2691B06
keywords theory of mindmentalizingstructural causal modelconflictepistemic accuracyinformation asymmetryresource-rational reasoninggame frame recognition
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks when AI systems should bother mentalizing—reading other agents' beliefs, intentions, and knowledge gaps—rather than how. It answers with a structural causal model, a DAG, in which four inputs (conflict complexity, information asymmetry, objective tractability, and agent sophistication) feed five mediators to produce a three-state theory-of-mind engagement trigger. The central claim is that mentalizing is causally warranted only when there is information asymmetry (the enabling cause), the agent cannot access an analytical solution, and/or the agent believes it is miscalibrated relative to its opponent. The paper's primary outcome is epistemic accuracy, not behavior, which it argues decouples social reasoning from behavioral policy and generalizes to coordination and cooperation. A sympathetic reader would see this as the first principled 'when' counterpart to the many 'how' models of machine mentalizing.

Core claim

The core discovery is a parameter-free (at the structural level) causal specification of ToM engagement as a two-stage threshold process. Stage one triggers engagement when a weighted sum of information asymmetry, inaccessible tractability (1−AT), and relative-sophistication miscalibration |RS−1| exceeds θE. Stage two accepts or rejects the mentalizing output based on observable signals, sophistication, and conflict complexity. The combined state ToM ∈ {0,1,2} then sets the mixture weights of analytical, intuitive, and mentalizing reasoning modes in the epistemic accuracy equation. The paper argues that epistemic accuracy is a cleaner, decoupled optimization target than behavior because agen

What carries the argument

The central object is the structural causal DAG with the mechanistic ToM node. The load-bearing identity is the engagement trigger equation, E = 1[λ1·IA + λ2(1−AT) + λ3·|RS−1| > θE], which gives two non-enabling pathways (tractability via low AT or low POT, reasoning-depth via miscalibration) and one enabling-cause pathway (IA, entered additively but conceptually a gate). The acceptance stage adds a second threshold on OS, S, and C. The outcome equation EA = w1*f_analytical(AT) + w2*f_ToM(RS) + w3*f_intuitive decouples epistemic accuracy from behavior; behavior (CB) is downstream as the joint product of EA and RS.

Load-bearing premise

The whole trigger depends on compressing an agent's reasoning depth, game-frame recognition, and opponent modeling into a single fixed number S on [0,1], and the paper's own limitations admit that S should likely be endogenous and updated.

What would settle it

Give an AI full information about an opponent (IA=0) in a simple analytically solvable game (e.g., a one-shot Prisoner's Dilemma with known payoffs), with accurate self-other calibration (RS=1). The model predicts ToM=0 and no epistemic gain from mentalizing. If a high-sophistication agent demonstrably improves its prediction of the opponent's action under these conditions by mentalizing (e.g., because the opponent uses a heuristic), the engagement condition is misspecified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • AI agents can use C, IA, OT, S, POT, AT, and RS to decide before acting whether mentalizing is warranted, saving resources and avoiding detrimental over-mentalizing.
  • ToM engagement becomes a falsifiable decision procedure: the model predicts three causal pathways and two thresholds, so simulations can test whether contextual engagement matches full-engagement epistemic accuracy at lower reasoning cost.
  • Epistemic accuracy as the outcome provides a clean loss function for learning-based social reasoning, separable from behavioral policy.
  • The model explains submentalizing as the behavioral limiting case: when ToM is not engaged, behavior is driven by RS alone.
  • The modular DAG structure extends to coordination and cooperation with re-weighted edges, and the 'expensive cognition' framing applies to causal reasoning and planning.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the causal model is right, a practical test emerges: deploy a ToM-capable system in a cooperative task with full information symmetry (IA≈0) and high tractability; the model predicts no added benefit from mentalizing. The search-and-rescue helper failures described in the paper are thus not implementation failures but predicted consequences of mis-specified task design.
  • The additive treatment of IA is a potential weakening; a multiplicative gate would make the enabling-cause claim more precise and testable. Choosing between them via simulation would sharpen predictions.
  • The static scalar S is the fragile component; treating S as an observable distribution over reasoning depth, frame recognition, and opponent modeling, updated by observable signals, would let the model handle agents who are deep but frame-confused—a case the paper itself notes defeats the RS pathway.
  • The acceptance threshold θA suggests an intervention strategy for AI transparency: by generating stronger observable signals, a teammate could push an AI's mentalizing output from rejected to accepted, effectively steering when AI trusts its social reasoning.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a structural causal model, formalized as a DAG, to answer the question of when theory of mind (ToM) should be engaged in conflict scenarios rather than how it should be implemented. The model has four exogenous variables (C, IA, OT, S), five endogenous mediators (OS, POT, AT, PS, RS), a mechanistic ToM node with three states (0, 1, 2), and a primary outcome of epistemic accuracy (EA). ToM engagement is governed by a two-stage threshold process in Eqs. (6)–(8): an engagement trigger based on a weighted combination of information asymmetry, low accessible tractability, and relative-sophistication miscalibration, followed by an acceptance decision based on observable signals, sophistication, and conflict complexity. Epistemic accuracy is modeled as a weighted mixture of analytical, intuitive, and mentalizing modes in Eq. (9). The paper claims this framework provides AI systems with a principled, resource-rational decision procedure for mentalizing, decouples social reasoning from behavioral policy, and generalizes beyond conflict. No simulation or empirical results are reported; validation is explicitly deferred to future work.

Significance. If fully realized, the framework would address a genuine gap in the AI-ToM literature, which has focused on mechanisms for mentalizing rather than the conditions under which mentalizing is causally warranted. The choice of epistemic accuracy as the outcome variable, rather than observable behavior, is well-motivated and potentially fruitful, as is the explicit separation of objective and perceived variables (e.g., OT vs. POT). The paper is transparent about its unresolved components and limitations, which is commendable. However, the formal content as presently stated is too underspecified to support the advertised contribution: the central equations contain unspecified functions and numerous free parameters, and the formal engagement rule contradicts the stated enabling-cause role of IA. These issues are load-bearing because the abstract and conclusion promise a 'principled, resource-rational decision procedure' that the model does not yet actually supply.

major comments (3)
  1. [§3.2, §3.4, §3.7 (Eq. 6)] The conceptual model states that IA is an enabling cause: 'without IA, there is nothing to mentalize about and ToM engagement has no useful causal work to do regardless of other conditions.' However, Eq. (6) defines E = 1[λ1·IA + λ2(1−AT) + λ3|RS−1| > θE], which is additive. With IA = 0, the inequality can still hold whenever λ2(1−AT) + λ3|RS−1| exceeds θE. Thus the formal model permits ToM engagement under perfect information symmetry, directly contradicting the stated necessity. The paper acknowledges this in the paragraph after Eq. (6) and in §5, but the choice of the additive form is not a harmless parameterization: it changes the causal claim. The central claim that IA is an enabling cause is therefore not represented by the formal equations. The model should either adopt a multiplicative gate (e.g., E = 1[IA·(λ2(1−AT) + λ3|RS−1|) > θE]) or restate IA as a contributing factor rather
  2. [§3.7, §4.1] The paper claims to provide a 'principled, resource-rational decision procedure' for mentalizing. As written, however, the model is a qualitative causal skeleton. Eqs. (1), (3), and (9) contain entirely unspecified functional forms (f1, g, f_analytical, f_ToM, f_intuitive). Eq. (6) contains six free parameters (λ1, λ2, λ3, θE; plus λ4, λ5, λ6, θA in Eq. (7)), and the mixture weights w1, w2, w3 in Eq. (9) are only described verbally as functions of the ToM state. No parameter estimation, calibration, or identification strategy is provided. Consequently, the model as given cannot yield quantitative or unambiguous qualitative predictions, and the claimed falsifiability is not established: with enough free parameters and unspecified functions, any observed engagement pattern can be accommodated. Moreover, 'resource-rational' is never formalized: there is no computational-cost term or expecte
  3. [§2.1, §3.2, §5] Sophistication S is compressed into a fixed scalar on [0,1] that is assumed to subsume recursive reasoning depth, game-frame recognition, and opponent modeling, with higher S implying better frame recognition. Yet the paper itself describes in §2.1 an agent with high reasoning depth but poor game-frame recognition who can be confidently wrong, and in §5 it admits that S 'almost certainly should be endogenous plus update during repeated interactions' and that its internal structure is undefined. Because S enters the key equations (2)–(7) and (9), the model's predictions are contingent on a construct whose measurement and aggregation are unspecified. This is a reasonable simplification for a conceptual framework, but it prevents the claimed decision procedure from being instantiated for actual AI systems. The distinction between depth and frame recognition matters for the reasoning-depth p
minor comments (4)
  1. [§3.7 (Eq. 7)] The sentence following Eq. (7) states that 'θA is conceptually distinct from θE: the former governs situational triggering, hwile the latter governs confidence in the mentalizing output.' This appears to be reversed: θE is the engagement threshold governing situational triggering, while θA governs acceptance. Please correct the wording and the typo 'hwile'.
  2. [Throughout] There are several typographical and formatting errors: 'casual predictions' should be 'causal predictions' (§3.6); 'haracterize' should be 'characterize' (§3.1); 'strucutres' should be 'structures' (§4.1); 'as no useful causal work' should be 'has no useful causal work' (§3.2); 'laid our in detail' should be 'laid out in detail' (§3.6). The LaTeX artifact 'IA99KT oM' in the edge list should be rendered as a dashed arrow (IA ⊸ ToM).
  3. [Abstract] The abstract states that 'Simulation validation, empirical human-machine teaming studies, and ethical considerations arising from conflict-optimized mentalizing are discussed.' The first two items are only discussed as future work, not presented as results. This could mislead readers; consider rewording to 'are identified as necessary next steps' or similar.
  4. [§3.7] Equation (9) defines EA = w1·f_analytical(AT) + w2·f_ToM(RS) + w3·f_intuitive + ε6. The dependency of the weights on the ToM state is described only in a bulleted list. For clarity, define w1, w2, w3 as explicit functions of ToM (e.g., w2 = 0 for ToM ∈ {0,1}, w2 = c>0 for ToM=2) so that the mixture is formally specified.

Circularity Check

0 steps flagged

No significant circularity: Eq. 6 and the DAG are stipulated, not fitted or derived from the quantities they 'predict'; the IA/enabling-cause mismatch is an internal consistency issue, not a circular reduction.

full rationale

The paper does not fit a parameter and then rename the fit as a prediction. The engagement rule (Eq. 6) is openly posited as a functional form—'Functional form specification, including the choice of nonlinear versus linear relationships, parameter estimation, and node measurement scales, is explicitly deferred to the simulation phase' (Sec. 3.7)—and the three 'causal pathways' are just the variable groups in that defining equation. A model's consequences following from its defining equations is ordinary model semantics, not circular derivation. The only self-citations (e.g., [18] for ToM-U) are not load-bearing reductions: Eq. 9 defines epistemic accuracy on the page, and the engagement model does not invoke ToM-U to compute Eq. 6. The prose claim that IA is an enabling cause and the additive form of Eq. 6 are indeed in tension; the paper itself acknowledges that a multiplicative gate 'would more precisely formalize the enabling cause relationship' but chooses the additive form for empirical flexibility. That is an internal-consistency/correctness problem, not a circular step in which an output equals an input. Similarly, leaving functional forms, weights, thresholds, and the internal structure of S unspecified limits empirical content but does not make the model's claims circular. External references (e.g., de Weerd et al. [9], Miller [32], Heyes [24], Santiesteban et al. [40]) provide independent, non-self-citational evidence for the motivating regularities. Under the required standard—quote the paper and exhibit the reduction—no circularity can be shown.

Axiom & Free-Parameter Ledger

17 free parameters · 7 axioms · 4 invented entities

The model is an uninstantiated specification: every causal equation carries a free parameter with no fitted values, and the central constructs (ToM node, RS, AT, EA) exist only as definitions. The causal edges themselves are assumptions about how variables interact, drawn from literature but not estimated.

free parameters (17)
  • α
    Eq. 2 weight on OS in POT update; value unfitted.
  • β
    Eq. 2 weight on C degrading POT; value unfitted.
  • γ
    Eq. 2 weight on S improving calibration; value unfitted.
  • δ
    Eq. 3 self-projection anchoring weight in [0,1]; value unfitted.
  • β1
    Eq. 5 weight on POT in AT; unfitted.
  • β2
    Eq. 5 weight on C (decreasing AT); unfitted.
  • β3
    Eq. 5 weight on S in AT; unfitted.
  • λ1
    Eq. 6 weight on IA for engagement; unfitted.
  • λ2
    Eq. 6 weight on (1−AT); unfitted.
  • λ3
    Eq. 6 weight on |RS−1|; unfitted.
  • θE
    Eq. 6 engagement threshold; treated as fixed for initial analyses.
  • λ4
    Eq. 7 weight on OS for acceptance; unfitted.
  • λ5
    Eq. 7 weight on S for acceptance; unfitted.
  • λ6
    Eq. 7 weight on C (decreasing acceptance); unfitted.
  • θA
    Eq. 7 acceptance threshold; unfitted.
  • w1, w2, w3
    Eq. 9 mixture weights summing to 1; assigned heuristically by ToM state, values unfitted.
  • f1, g, f_analytical, f_ToM, f_intuitive
    Unspecified functional forms; the paper states these are 'intentionally left unspecified' (Sec. 3.7).
axioms (7)
  • standard math Acyclic DAG semantics of Pearl's structural causal model
    The paper invokes Pearl [35] and uses DAG mechanics (Sec. 3.1, 3.6).
  • standard math Logistic sigmoid for bounded variables
    Used in Eqs. 2 and 5 to map to [0,1].
  • domain assumption ToM is contextually engaged, not always-on
    Adopted from cited evidence (de Weerd et al. [9], Heyes [24], Miller [32]) and stated as a foundational design principle in Sec. 2.3; the entire model depends on it.
  • domain assumption Information asymmetry is an enabling cause for ToM: without IA, mentalizing has no useful causal work
    Stated in Sec. 3.2 and used to justify the IA→ToM edge; however Eq. (6) implements IA additively, not as a strict gate, creating a tension the paper acknowledges.
  • ad hoc to paper Agent sophistication is a fixed scalar composite S
    Sec. 3.2 collapses recursive depth, game-frame recognition, and opponent modeling into one number; the paper grants in Sec. 5 it should be endogenous and internally structured.
  • ad hoc to paper Engagement threshold is an additive linear combination of IA, (1−AT), |RS−1|
    Eq. (6); the paper notes a multiplicative alternative would enforce the enabling-cause gate but chooses additive for empirical flexibility.
  • ad hoc to paper Acceptance threshold is an additive linear combination of OS, S, C
    Eq. (7); no derivation given, with thresholds left to future simulation.
invented entities (4)
  • ToM mechanism node with states {0,1,2} no independent evidence
    purpose: Represents whether mentalizing is off, engaged-and-rejected, or engaged-and-accepted
    The paper introduces a latent mechanistic node without a measurement procedure or external handle.
  • Epistemic Accuracy (EA) no independent evidence
    purpose: Primary outcome and proposed optimization target, decoupled from behavior
    Defined as a weighted mixture in Eq. (9) but no operationalization or benchmark is supplied; grounding is deferred to unpublished ToM-U work [18].
  • Relative Sophistication (RS) no independent evidence
    purpose: Ratio S/PS used as the reasoning-depth trigger for ToM engagement
    Introduced as a latent miscalibration index; no data support.
  • Accessible Tractability (AT) no independent evidence
    purpose: Captures whether the agent can derive the analytical solution, mediating engagement
    A theoretical construct with no empirical anchor in the paper.

pith-pipeline@v1.3.0-alltime-deepseek · 13184 in / 16996 out tokens · 169576 ms · 2026-08-02T11:06:19.343737+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of A Causal Model of Theory of Mind in Conflict for Artificial Intelligence." pith.science (2026). https://pith.science/paper/FWCWSH2U

@misc{pith2026260616944,
  author       = {Pith},
  title        = {Pith review of: A Causal Model of Theory of Mind in Conflict for Artificial Intelligence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FWCWSH2U}},
  note         = {Machine review of arXiv:2606.16944}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Theory of mind (ToM), the capacity to ascribe mental states to others and use those ascriptions for prediction and inference, is widely assumed to be essential for effective human-machine integration. Existing AI-ToM models address \emph{how} to mentalize, but leave the question of when largely unaddressed. The central question is: under what situational and agent-level conditions is ToM engagement causally warranted in conflict? This paper presents a structural causal model formalized as a directed acyclic graph (DAG), treating ToM as a mechanism activated by situational and agent-level conditions rather than as an always-on capacity. The model specifies four exogenous variables capturing situational and agent-level conditions, five endogenous mediators, and a mechanistic ToM node producing engagement states through three distinct causal pathways: a tractability pathway, a reasoning-depth pathway, and an enabling-cause pathway. The primary outcome is epistemic accuracy, which decouples social reasoning from behavioral policy and generalizes across social phenomena beyond conflict. The framework gives AI systems a principled, resource-rational decision procedure for mentalizing, with implications for efficiency, trust, and the development of robust artificial social intelligence. Simulation validation, empirical human-machine teaming studies, and ethical considerations arising from conflict-optimized mentalizing are discussed.

Figures

Figures reproduced from arXiv: 2606.16944 by Nikolos Gurney.

Figure 1
Figure 1. Figure 1: Structural causal model of Theory of Mind engagement in conflict. Solid grey ar [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

45 extracted references · 2 linked inside Pith

  1. [1]

    Investigation of automated vehicle effects on driver’s behavior and traffic performance.Transportation research procedia, 15:761–770, 2016

    Erfan Aria, Johan Olstam, and Christoph Schwietering. Investigation of automated vehicle effects on driver’s behavior and traffic performance.Transportation research procedia, 15:761–770, 2016

  2. [2]

    Baker, Rebecca Saxe, and Joshua B

    Chris L. Baker, Rebecca Saxe, and Joshua B. Tenenbaum. Action understanding as inverse planning.Cognition, 113(3):329–349, 2009

  3. [3]

    The curse of knowledge in reasoning about false beliefs.Psychological science, 18(5):382–386, 2007

    Susan AJ Birch and Paul Bloom. The curse of knowledge in reasoning about false beliefs.Psychological science, 18(5):382–386, 2007

  4. [4]

    The curse of knowledge in economic settings: An experimental analysis.Journal of political Economy, 97(5):1232– 1254, 1989

    Colin Camerer, George Loewenstein, and Martin Weber. The curse of knowledge in economic settings: An experimental analysis.Journal of political Economy, 97(5):1232– 1254, 1989

  5. [5]

    A cognitive hierarchy model of games.The Quarterly Journal of Economics, 119(3):861–898, 2004

    Colin F Camerer, Teck-Hua Ho, and Juin-Kuan Chong. A cognitive hierarchy model of games.The Quarterly Journal of Economics, 119(3):861–898, 2004

  6. [6]

    Human–agent teaming for multirobot control: A review of human factors issues.IEEE Transactions on Human-Machine Systems, 44(1):13–29, 2014

    Jessie YC Chen and Michael J Barnes. Human–agent teaming for multirobot control: A review of human factors issues.IEEE Transactions on Human-Machine Systems, 44(1):13–29, 2014

  7. [7]

    Theory of mind in large language models: Assessment and enhancement

    Ruirui Chen, Weifeng Jiang, Chengwei Qin, and Cheston Tan. Theory of mind in large language models: Assessment and enhancement. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 31539–31558, 2025

  8. [8]

    Grounding in communication.Perspectives on socially shared cognition, pages 127—-149, 1991

    Herbert H Clark and Susan E Brennan. Grounding in communication.Perspectives on socially shared cognition, pages 127—-149, 1991

  9. [9]

    Higher-order theory of mind is especially useful in unpredictable negotiations.Autonomous Agents and Multi-Agent Systems, 36(2):30, 2022

    Harmen de Weerd, Rineke Verbrugge, and Bart Verheij. Higher-order theory of mind is especially useful in unpredictable negotiations.Autonomous Agents and Multi-Agent Systems, 36(2):30, 2022

  10. [10]

    An implemented theory of mind to improve human- robot shared plans execution

    Sandra Devin and Rachid Alami. An implemented theory of mind to improve human- robot shared plans execution. InProceedings of the 11th ACM/IEEE HRI, pages 319– 326, 2016

  11. [11]

    Perspective taking as egocentric anchoring and adjustment.Journal of personality and social psychology, 87(3):327, 2004

    Nicholas Epley, Boaz Keysar, Leaf Van Boven, and Thomas Gilovich. Perspective taking as egocentric anchoring and adjustment.Journal of personality and social psychology, 87(3):327, 2004. 18

  12. [12]

    Jonathan Freeman, Lixiao Huang, Mackenzie Wood, and S. J. Cauffman. Evaluating artificial social intelligence in an urban search and rescue task environment. InCom- putational Theory of Mind for Human-Machine Teams, volume 13775 ofLNCS, pages 72–84. Springer, 2022

  13. [13]

    The neural basis of mentalizing.Neuron, 50(4):531–534, 2006

    Chris D Frith and Uta Frith. The neural basis of mentalizing.Neuron, 50(4):531–534, 2006

  14. [14]

    Siba Ghrear, Adam Baimel, Taeh Haddock, and Susan A. J. Birch. Are the classic false belief tasks cursed? Young children are just as likely as older children to pass a false belief task when they are not required to overcome the curse of knowledge.PLOS ONE, 16(2):e0244141, 2021

  15. [15]

    Gmytrasiewicz and Prashant Doshi

    Piotr J. Gmytrasiewicz and Prashant Doshi. A framework for sequential planning in multi-agent settings.Journal of Artificial Intelligence Research, 24:49–79, 2005

  16. [16]

    Automation bias: a systematic review of frequency, effect mediators, and mitigators.Journal of the American Medical Informatics Association, 19(1):121–127, 2012

    Kate Goddard, Abdul Roudsari, and Jeremy C Wyatt. Automation bias: a systematic review of frequency, effect mediators, and mitigators.Journal of the American Medical Informatics Association, 19(1):121–127, 2012

  17. [17]

    Alison Gopnik and Henry M. Wellman. Why the child’s theory of mind really is a theory.Mind & Language, 7(1–2):145–171, 1992

  18. [18]

    The theory of mind utility: Formal specification of a mentalizing mechanism, 2026

    Nikolos Gurney and Stacy Marsella. The theory of mind utility: Formal specification of a mentalizing mechanism, 2026

  19. [19]

    Pynadath

    Nikolos Gurney, Stacy Marsella, Volkan Ustun, and David V. Pynadath. Operational- izing theories of theory of mind: A survey. InComputational Theory of Mind for Human-Machine Teams, volume 13775 ofLNCS, pages 3–20. Springer, 2022

  20. [20]

    Pynadath

    Nikolos Gurney and David V. Pynadath. Robots with theory of mind for humans: A survey. InProceedings of the 31st IEEE RO-MAN, pages 993–1000, 2022

  21. [21]

    Spontaneous theory of mind for artificial intelligence

    Nikolos Gurney, David V Pynadath, and Volkan Ustun. Spontaneous theory of mind for artificial intelligence. InInternational conference on human-computer interaction, pages 60–75. Springer, 2024

  22. [22]

    Joseph Y. Halpern. Causes and explanations: A structural-model approach. Part I: Causes.The British Journal for the Philosophy of Science, 56(4):843–887, 2005

  23. [23]

    Joseph Y. Halpern. Causes and explanations: A structural-model approach. Part II: Explanations.The British Journal for the Philosophy of Science, 56(4):889–911, 2005

  24. [24]

    Submentalizing: I am not really reading your mind.Perspectives on Psychological Science, 9(2):131–143, 2014

    Cecilia Heyes. Submentalizing: I am not really reading your mind.Perspectives on Psychological Science, 9(2):131–143, 2014

  25. [25]

    Re-evaluating theory of mind evaluation in large language models.Philosophical Transactions of the Royal Society B: Biological Sciences, 380(1932), 2025

    Jennifer Hu, Felix Sosa, and Tomer Ullman. Re-evaluating theory of mind evaluation in large language models.Philosophical Transactions of the Royal Society B: Biological Sciences, 380(1932), 2025. 19

  26. [26]

    Theoryofmindasinversereinforcementlearning.Current Opinion in Behavioral Sciences, 29:105–110, 2019

    JulianJara-Ettinger. Theoryofmindasinversereinforcementlearning.Current Opinion in Behavioral Sciences, 29:105–110, 2019

  27. [27]

    Evaluating large language models in theory of mind tasks.Proceedings of the National Academy of Sciences, 121(45):e2405460121, 2024

    Michal Kosinski. Evaluating large language models in theory of mind tasks.Proceedings of the National Academy of Sciences, 121(45):e2405460121, 2024

  28. [28]

    Leslie, Ori Friedman, and Tim P

    Alan M. Leslie, Ori Friedman, and Tim P. German. Core mechanisms in ‘theory of mind’.Trends in Cognitive Sciences, 8(12):528–533, 2004

  29. [29]

    Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources.Behavioral and brain sciences, 43:e1, 2020

    Falk Lieder and Thomas L Griffiths. Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources.Behavioral and brain sciences, 43:e1, 2020

  30. [30]

    Hot-cold empathy gaps and medical decision making.Health psychology, 24(4S):S49, 2005

    George Loewenstein. Hot-cold empathy gaps and medical decision making.Health psychology, 24(4S):S49, 2005

  31. [31]

    MIT Press, Cambridge, MA, 1982

    David Marr.Vision: A Computational Investigation into the Human Representation and Processing of Visual Information. MIT Press, Cambridge, MA, 1982

  32. [32]

    Miller.Ex Machina: Coevolving Machines and the Origins of the Social Uni- verse

    John H. Miller.Ex Machina: Coevolving Machines and the Origins of the Social Uni- verse. SFI Press, 2022

  33. [33]

    Deep interpretable modelsof theory of mind

    Ifeoma Oguntola, Daniel Hughes, and KatiaSycara. Deep interpretable modelsof theory of mind. InProceedings of the 30th IEEE RO-MAN, pages 657–664, 2021

  34. [34]

    Human–robot collaborations in smart manufacturing environments: review and outlook.Sensors, 23(12):5663, 2023

    Uqba Othman and Erfu Yang. Human–robot collaborations in smart manufacturing environments: review and outlook.Sensors, 23(12):5663, 2023

  35. [35]

    Cambridge University Press, Cambridge, 2nd edition, 2009

    Judea Pearl.Causality: Models, Reasoning, and Inference. Cambridge University Press, Cambridge, 2nd edition, 2009

  36. [36]

    Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4):515–526, 1978

    David Premack and Guy Woodruff. Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4):515–526, 1978

  37. [37]

    Ef- fectiveness of teamwork-level interventions through decision-theoretic reasoning in a minecraft search-and-rescue task

    David V Pynadath, Nikolos Gurney, Sarah Kenny, Rajay Kumar, Stacy C Marsella, Haley Matuszak, Hala Mostafa, Pedro Sequeira, Volkan Ustun, and Peggy Wu. Ef- fectiveness of teamwork-level interventions through decision-theoretic reasoning in a minecraft search-and-rescue task. InProceedings of the 2023 International Conference on Autonomous Agents and Multi...

  38. [38]

    Pynadath, Nikolos Gurney, Siena Kenny, and Rajesh Kumar

    David V. Pynadath, Nikolos Gurney, Siena Kenny, and Rajesh Kumar. Effectiveness of teamwork-level interventions through decision-theoretic reasoning in a Minecraft search- and-rescue task. InProceedings of the ..., 2023

  39. [39]

    Pynadath and Stacy C

    David V. Pynadath and Stacy C. Marsella. Psychsim: Modeling theory of mind with decision-theoretic agents. InIJCAI, volume 5, pages 1181–1186, 2005

  40. [40]

    Craig Hopkins, Geoffrey Bird, and Cecilia Heyes

    Idalmis Santiesteban, Caroline Catmur, S. Craig Hopkins, Geoffrey Bird, and Cecilia Heyes. Avatars and arrows: implicit mentalizing or domain-general processing?Journal of Experimental Psychology: Human Perception and Performance, 40(3):929–937, 2014. 20

  41. [41]

    Large language models fail on trivial alterations to theory-of-mind tasks

    Tomer Ullman. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399, 2023

  42. [42]

    ToM2C: Target-oriented multi-agent communication and cooperation with theory of mind.arXiv preprint arXiv:2111.09189, 2021

    Yuanfei Wang, Fan Zhong, Jianye Xu, and Yaodong Wang. ToM2C: Target-oriented multi-agent communication and cooperation with theory of mind.arXiv preprint arXiv:2111.09189, 2021

  43. [43]

    Wellman, David Cross, and Julanne Watson

    Henry M. Wellman, David Cross, and Julanne Watson. Meta-analysis of theory-of-mind development: The truth about false belief.Child development, 72(3):655–684, 2001

  44. [44]

    Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception.Cognition, 13(1):103–113, 1983

    Heinz Wimmer and Josef Perner. Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception.Cognition, 13(1):103–113, 1983

  45. [45]

    Game theory of mind.PLoS compu- tational biology, 4(12):e1000254, 2008

    Wako Yoshida, Ray J Dolan, and Karl J Friston. Game theory of mind.PLoS compu- tational biology, 4(12):e1000254, 2008. 21