Pith. sign in

REVIEW 2 major objections 4 minor 40 references

This paper argues that a user's felt sense of an AI's inner life — perceived mind — is governed by dimensional completeness, the joint expression of four first-person stances, and is separable from raw task intelligence, testable at matched

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 22:02 UTC pith:6RRYK4MK

load-bearing objection Useful, honest framework paper whose core testability claim has an unresolved identification gap; worth reading for the decomposition, not for evidence. the 2 major comments →

arxiv 2607.15883 v1 pith:6RRYK4MK submitted 2026-07-17 cs.HC cs.AI

Perceived AGI: Believability as Dimensional Completeness, Not Capability

classification cs.HC cs.AI
keywords perceived mindbelievabilitymind perceptionanthropomorphismconversational AIlarge language modelscompanion AIinitiative and cadence
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Large language models answer well but feel flat in sustained conversation; the paper's thesis is that the missing ingredient is not intelligence but 'dimensional completeness' — whether the agent expresses a small set of first-person stances humans read as evidence of a mind. Four stances are proposed: time, truth, entropy, and love, each defined as a behavioral stance an engineer can build on the same underlying model, not as a capability. The paper claims that at matched, measured task performance, a user's perceived mind — attributed inner life — increases with completeness, and that the joint effect should exceed the sum of single cues. It identifies an observable behavior layer, initiative and cadence, through which the stances surface, gives concrete emulation paths (one prototyped), states six falsifiable predictions, and draws a firm boundary: the object is inferable interiority, not machine consciousness. All comparative claims are predictions; no human-subjects data appear here.

Core claim

The central discovery claim is a hypothesis about attribution, stated in the paper's own terms: at matched capability — operationalized as a matched base model with task performance measured, not assumed fixed — perceived mind increases with dimensional completeness, the joint presence of four first-person stances. Each dimension is defined by the fixed-model test: it must be exhibitable on a fixed underlying model, more or less, without changing what the agent knows. Time is endogenous rhythm and continuity; truth is epistemic humility and visible belief revision; entropy is identity-consistent variation rather than randomness; love is partiality toward a specific user and a persisting rela

What carries the argument

The carrying machinery is the four-stance decomposition plus the fixed-model test. A 'first-person stance' is a behavioral disposition an agent exhibits — not a competency — and each passes the fixed-model test if the same base model can be made to exhibit more or less of it. Time (endogenous rhythm and continuity), truth (epistemic humility and revision), entropy (structured identity-consistent variation), and love (partiality and relationship persistence) are the four; the behavior layer of initiative and cadence is the observable surface through which they reach the user. The matched-base-model protocol prevents any perceived-mind effect from being explained by a smarter model, and the co

Load-bearing premise

The load-bearing premise, stated as the fixed-model test in Section 3, is that each of the four stances can be independently built on the same base model without changing capability, content quality, or interaction length; the paper itself notes that only time currently has a prototype, so this premise is untested.

What would settle it

A pre-registered matched-model study with a factorial ablation of the four stances would settle it: if the dimensionally complete agent does not outrate the best single-stance agent and does not exceed an additive main-effects prediction, completeness fails; separately, if a proactive agent is rated no higher than an otherwise identical reactive one, the behavior-layer flagship fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Believability becomes an engineerable target separable from capability: a product could raise perceived mind by adding stances or behavior-layer features without upgrading the model.
  • The default reactive chatbot is predicted to sit at a 'capable-but-flat' ceiling; even minimal changes such as occasional double replies or unprompted messages are hypothesized to move perceived mind.
  • P3 and P4, which test initiative and cadence directly, are pre-registrable now, so the behavior-layer half of the framework can be checked before the other dimensions are built.
  • If P6 holds, dimensional completeness is a genuine interaction effect — the whole of the four stances exceeds the sum of its parts — not just four useful cues.
  • The framework's acceptance test doubles as an ethical constraint: any believability effect that collapses when its mechanism is disclosed is treated as concealment-dependent and abandoned.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Extension: If perceived mind is separable from measured capability, capability benchmarks and companion-product roadmaps may be optimizing different things; a perceived-mind scale alongside task benchmarks would make the trade-off visible.
  • Extension: The paper's own order of operations suggests a cheap first falsification: run the pre-registrable P3/P4 comparison on a matched base model before building the truth, entropy, and love ablations, since failure there removes the behavioral flagship.
  • Extension: The transparency-collapse test could generalize as an audit instrument for any anthropomorphic AI feature: cross any believability feature with a disclosure factor and ask whether the perceived-mind advantage survives.
  • Extension: If the four stances turn out to load onto the established agency/experience axes, then the framework's contribution would reduce to engineering guidance rather than a new psychometric structure; the planned factor analysis is the point where that would be decided.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper proposes a conceptual framework for why capable LLM interlocutors feel 'flat' in sustained conversation. It argues that perceived mind—attributed inner life—is driven not by task intelligence but by 'dimensional completeness': the joint expression of four first-person stances (time, truth, entropy, love) through an observable behavior layer of initiative and cadence. Each stance is defined behaviorally rather than as a benchmark competency, and a 'fixed-model test' is offered as a guard against circularity. The paper reports no human-subjects data, explicitly labels its central comparative claims as predictions rather than findings, and lists six falsifiable predictions (P1–P6) for a future preregistered study, separating those that are currently pre-registrable (P3, P4) from those awaiting operationalization. An ethics section specifies disclosure, user control, and vulnerable-user safeguards.

Significance. If the framework holds, it would provide an engineerable decomposition of believability that is partially orthogonal to measured task performance, with direct implications for companion-AI design and evaluation. The paper's main strengths are its unusually honest status ledger—distinguishing pre-registrable predictions from conjectures—and its explicit attempt to design around the capability confound rather than assume it away. It also offers a concrete, testable complementarity prediction (P6) and a substantive ethics section with a falsifiable transparency-collapse test. The contribution is best understood as a research agenda or position paper, not an empirical result.

major comments (2)
  1. [§5 Identification strategy; Table 3] The §2 central claim requires 'matched capability' to make the completeness-not-capability contrast meaningful, and §5 admits capability is 'not metaphysically fixed.' The P3/P4 estimand is the total effect of an initiative/cadence package; task performance, content volume, and rated competence are only 'matched or measured,' with content volume explicitly post-treatment. Since stance emulations add memory, initiative, cadence, and relationship state, they plausibly alter task-relevant performance, making measured capability a post-treatment mediator. A total-effect estimand cannot separate 'dimensional completeness' from 'capability'; descriptive recording or sensitivity analyses is not an identification strategy. Specify the primary estimand and assumptions (e.g., a surface-only control with capability fixed by construction, or formal mediation with sensitivity bounds), or restrict the
  2. [§3 Fixed-model test; Table 2] The fixed-model test (§3) establishes only exhibitability, not capability-orthogonality. Each emulation path adds task-relevant machinery: time uses oscillators and 'hierarchical modulation of retrieval and style' [4]; truth changes utterance content; entropy changes response selection; love requires durable user-specific memory and relationship state. Two agents on the same base model can therefore differ in what they know or in measured task performance. The assertion that time can differ 'without differing in what they know' is not demonstrated for these implementations. Without a companion-domain capability metric or a surface-only manipulation that provably leaves content and capability untouched, the fixed-model test cannot support the claim that stances are separable from capability.
minor comments (4)
  1. [Title and Abstract] The title and abstract present 'Believability as Dimensional Completeness' as a thesis; the abstract does clearly say the claims are predictions, but adding 'Hypothesis' or 'Proposal' to the title would reduce the risk of overreading.
  2. [§3 'Why these four?'] The two selection filters—'distinct inference humans draw about minds' and 'recognizable failure mode'—are acknowledged as subjective. It would strengthen the paper to cite empirical evidence for each filter or label them more explicitly as design choices rather than psychological findings.
  3. [§3 Time prototype; §4 Anima Felix] The time prototype and deployed features are author-reported and the paper says so. This is acceptable for a feasibility illustration, but the statements should be kept clearly separate from any claim of empirical validation; consider a boxed 'author-attested, not independently verified' note.
  4. [§5 Dependent measure] The 'perceived mind' rating battery is described only as adapted from dimensional mind-perception scales. Since the paper stipulates a unidimensional construct while citing a dimensional literature, the planned factor analysis should be described in enough detail that a reader can see how the single primary outcome would be formed.

Circularity Check

2 steps flagged

Core completeness prediction largely restates the selection criterion for the four stances; time exhibitability leans on an author-reported self-citation. The framework is otherwise honest, abstains from data, and retains independent content in complementarity and matched-capability tests.

specific steps
  1. self definitional [Section 2 (Thesis) + Section 3 (Why these four?) + Table 3 P1/P2]
    "'Completeness is likewise stipulative: it names the joint presence of the four stances of Section 3' (Sec. 2); 'each dimension names a distinct inference humans draw about minds—continuity, honesty, character, attachment—and each corresponds to a recognizable failure mode' (Sec. 3); P1: 'A dimensionally complete agent is rated higher than an ablated one.'"

    The predictor (dimensional completeness) is stipulated as the joint presence of stances that were selected precisely because each 'names a distinct inference humans draw about minds'—i.e., because each is a known cue of the same attribution that defines the outcome (perceived mind). P1 and P2 therefore unpack the selection rule rather than derive an independent prediction. The paper is candid that the cues are not newly discovered, and its genuinely independent content lies in P6 (complementarity) and in the matched-capability separation, so this is partial, not total, circularity.

  2. self citation load bearing [Section 3, 'Time' paragraph (citation [4])]
    "a companion paper specifies exactly such an architecture—coupled oscillators, zeitgeber entrainment, hierarchical modulation of retrieval and style [4]—and an author-reported prototype implements its oscillator, entrainment, and modulation core. We claim engineerability for time in the same restricted register as the deployment example of Section 4: an author-attested build, not an independently verified result."

    The exhibitability claim for the time stance—'the prototype demonstrates oscillator-level feasibility of rhythmicity and entrainment'—rests on a same-author preprint/prototype that is explicitly not independently verified. This is a self-citation used to support a present claim about buildability, but it is transparent and is not used to support any perceived-mind effect. It is a minor load-bearing self-citation rather than a full circular chain.

full rationale

This is a conceptual framework that reports no human-subjects data and explicitly labels its comparative claims as predictions, not findings. The main circularity concern is that the four dimensions are selected because they name inferences humans draw about minds, so predictions P1 and P2 are close to restatements of that selection criterion. However, the paper acknowledges that the cues are not undiscovered, frames the taxonomy as an engineering basis, and locates its genuinely falsifiable novelty in P6 (complementarity) and in the matched-capability, measured-task-performance design. No fitted parameters are renamed as predictions; the time prototype is an author-reported same-author build, cited transparently and not used as evidence of any effect. The separable-capability claim and complementarity prediction can fail and have independent content, so the argument is not forced. Score reflects one partial self-definitional overlap and one minor self-citation rather than a fully circular derivation.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 3 invented entities

The framework rests on untested psychological assumptions about how people attribute minds, on the feasibility of implementing the four stances independently, and on the identifiability of matched-model experiments. The only working prototype is author-reported, so the ledger is dominated by domain assumptions and planned-but-unbuilt operationalizations.

axioms (5)
  • domain assumption Humans infer mind from a small set of first-person stances in sustained conversation.
    Foundation of the framework, stated in Section 2; referenced to mind-perception literature but not empirically established for conversational LLMs.
  • domain assumption Believability is separable from task intelligence: a less capable system can feel more present.
    Central contrast from Section 1; supported by anecdote and the believable-agents tradition, but not by controlled comparison.
  • ad hoc to paper All four stances pass the fixed-model test — each can be exhibited on a fixed base model independently of capability.
    Defined in Section 3. Only time has an author-reported prototype; truth, entropy, and love are planned. The ablations in Section 5 depend on this assumption.
  • domain assumption Unprompted action (initiative) is difficult to explain except by positing an internal source, so it is a strong signal of mind.
    Stated in Section 4. The cited prior virtual-agent study [3] found self-motivated action did not improve believability, so this is fragile and load-bearing for P3.
  • ad hoc to paper A matched-base-model manipulation can be achieved without confounding by capability or content volume.
    Section 5 acknowledges stance machinery adds affordances and that realized content volume may differ; the protocol measures rather than removes these. The claim's identifiability is an assumption until tested.
invented entities (3)
  • The four first-person stances (time, truth, entropy, love) no independent evidence
    purpose: The proposed decomposition of perceived mind into engineerable behavioral stances.
    New theoretical constructs. Only time has a prototype, and it is author-reported; none has been shown to affect perceived mind.
  • Behavior layer (initiative and cadence) no independent evidence
    purpose: The observable surface through which the stances are expressed; the subject of predictions P3 and P4.
    Partially deployed in Anima Felix, but the paper makes no claim that these features improve perceived mind; no controlled evidence exists.
  • Perceived mind as a unidimensional construct no independent evidence
    purpose: The single pinned dependent variable of the framework.
    Defined by stipulation in Section 2 and Table 1; no validated measurement instrument is supplied.

pith-pipeline@v1.3.0-alltime-deepseek · 11090 in / 9510 out tokens · 100897 ms · 2026-08-01T22:02:27.719667+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Perceived AGI: Believability as Dimensional Completeness, Not Capability." pith.science (2026). https://pith.science/paper/6RRYK4MK

@misc{pith2026260715883,
  author       = {Pith},
  title        = {Pith review of: Perceived AGI: Believability as Dimensional Completeness, Not Capability},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6RRYK4MK}},
  note         = {Machine review of arXiv:2607.15883}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large language models are broadly capable, yet in sustained one-to-one conversation they still read as flat: competent, responsive, and somehow not quite the presence of a mind. We hypothesize that a central missing ingredient is not more capability but dimensional completeness. We propose that the believability of an artificial interlocutor -- the degree to which a user attributes an inner life to it, which we call perceived mind -- is governed by whether the agent expresses a small set of first-person stances that humans use as evidence of mind, and that this is separable from task intelligence. We name four such dimensions -- time, truth, entropy, and love -- each defined as a behavioral stance rather than a benchmark competency, each with a human analog and a concrete emulation path; the time dimension already has an author-reported prototype. We identify an observable behavior layer -- initiative (unprompted action) and cadence (the shape and timing of turns) -- through which the stances surface in conversation, both partially realized as deployed features in a production companion application. We state six falsifiable predictions that a later pre-registered study will test, separating those that are pre-registrable now from those that remain conjectures pending operationalization. This is a conceptual framework: it reports no human-subjects data, and its central comparative claims are predictions, not findings. Throughout we hold a firm boundary -- the object is inferrable interiority, not interiority; this is perception engineering, not a theory of machine consciousness -- and we treat the resulting attachment and manipulation risks as load-bearing rather than incidental.

Figures

Figures reproduced from arXiv: 2607.15883 by Sebastian Cochinescu.

Figure 1
Figure 1. Figure 1: The framework. Four first-person stances—time, truth, entropy, love—surface through an observable behavior layer (initiative, cadence, and content cues), from which the user infers a mind. The dependent construct is perceived mind; conditions share a matched base model, so any effect must survive measured task performance and rated competence. The predictions of Section 5 test the arrows: ablating stances … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

40 extracted references · 2 linked inside Pith

  1. [1]

    The role of emotion in believable agents.Communications of the ACM, 37(7):122–125, 1994

    Joseph Bates. The role of emotion in believable agents.Communications of the ACM, 37(7):122–125, 1994

  2. [2]

    Bickmore and Rosalind W

    Timothy W. Bickmore and Rosalind W. Picard. Establishing and maintaining long-term human- computer relationships.ACM Transactions on Computer-Human Interaction, 12(2):293–327, 2005

  3. [3]

    What makes virtual agents believable?Con- nection Science, 28(1):83–108, 2016

    Anton Bogdanovych, Tomas Trescak, and Simeon Simoff. What makes virtual agents believable?Con- nection Science, 28(1):83–108, 2016

  4. [4]

    SCN-LLM: An SCN-inspired temporal architecture for rhythmic entrainment and adaptive retrieval in large language models

    Sebastian Cochinescu. SCN-LLM: An SCN-inspired temporal architecture for rhythmic entrainment and adaptive retrieval in large language models. Preprint, ResearchGate publication 392556656, 2025. Revised version in preparation for arXiv

  5. [5]

    The Collected Works of Samuel Taylor Coleridge, Bollingen Series LXXV, Vol

    Samuel Taylor Coleridge.Biographia Literaria, or Biographical Sketches of My Literary Life and Opin- ions. The Collected Works of Samuel Taylor Coleridge, Bollingen Series LXXV, Vol. 7. Princeton University Press, Princeton, NJ, 1983. J. Engell and W. J. Bate (eds.); original work published 1817

  6. [6]

    Clara Colombatto and Stephen M. Fleming. Folk psychological attributions of consciousness to large language models.Neuroscience of Consciousness, 2024(1):niae013, 2024

  7. [7]

    Proactive conversational AI: A comprehensive survey of advancements and opportunities.ACM Transactions on Information Systems, 43(3):1–45, 2025

    Yang Deng, Lizi Liao, Wenqiang Lei, Grace Hui Yang, Wai Lam, and Tat-Seng Chua. Proactive conversational AI: A comprehensive survey of advancements and opportunities.ACM Transactions on Information Systems, 43(3):1–45, 2025

  8. [8]

    Dennett.The Intentional Stance

    Daniel C. Dennett.The Intentional Stance. MIT Press, Cambridge, MA, 1987

  9. [9]

    Cacioppo

    Nicholas Epley, Adam Waytz, and John T. Cacioppo. On seeing human: A three-factor theory of anthropomorphism.Psychological Review, 114(4):864–886, 2007

  10. [10]

    The ethics of advanced AI assistants.arXiv preprint arXiv:2404.16244, 2024

    Iason Gabriel, Arianna Manzini, Geoff Keeling, Lisa Anne Hendricks, Verena Rieser, et al. The ethics of advanced AI assistants.arXiv preprint arXiv:2404.16244, 2024. Google DeepMind

  11. [11]

    Ulrich Gnewuch, Stefan Morana, Marc T. P. Adam, and Alexander Maedche. Opposing effects of response time in human–chatbot interaction.Business & Information Systems Engineering, 64(6):773– 791, 2022

  12. [12]

    Gray, Kurt Gray, and Daniel M

    Heather M. Gray, Kurt Gray, and Daniel M. Wegner. Dimensions of mind perception.Science, 315(5812):619, 2007

  13. [13]

    Developing a scale for measuring the believability of virtual agents

    Siqi Guo, Nicoletta Adamo, and Christos Mousas. Developing a scale for measuring the believability of virtual agents. InICAT-EGVE 2023 – International Conference on Artificial Reality and Telexistence and Eurographics Symposium on Virtual Environments, pages 45–52. The Eurographics Association, 2023. 10

  14. [14]

    Richard Wohl

    Donald Horton and R. Richard Wohl. Mass communication and para-social interaction: Observations on intimacy at a distance.Psychiatry, 19(3):215–229, 1956

  15. [15]

    Principles of mixed-initiative user interfaces

    Eric Horvitz. Principles of mixed-initiative user interfaces. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’99), pages 159–166, 1999

  16. [16]

    AI in your mind: Counterbalancing perceived agency and experience in human-ai interaction

    Angel Hsing-Chi Hwang and Andrea Stevenson Won. AI in your mind: Counterbalancing perceived agency and experience in human-ai interaction. InExtended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (CHI EA ’22), 2022

  17. [17]

    Kalman and Sheizaf Rafaeli

    Yoram M. Kalman and Sheizaf Rafaeli. Online pauses and silence: Chronemic expectancy violations in written computer-mediated communication.Communication Research, 38(1):54–69, 2011

  18. [18]

    Karimova

    Gulnara Z. Karimova. Not in our image: Rethinking anthropomorphism in expert chatbot design.AI & Society, 41(1):611–628, 2026. First published online 2025

  19. [19]

    Evaluating large language models in theory of mind tasks.Proceedings of the National Academy of Sciences, 121(45):e2405460121, 2024

    Micha l Kosinski. Evaluating large language models in theory of mind tasks.Proceedings of the National Academy of Sciences, 121(45):e2405460121, 2024

  20. [20]

    Linnea Laestadius, Andrea Bishop, Michael Gonzalez, Diana Illenˇ c ´ ık, and Celeste Campos-Castillo. Too human and not human enough: A grounded theory analysis of mental health harms from emotional dependence on the social chatbot Replika.New Media & Society, 26(10):5923–5941, 2024

  21. [21]

    Universal intelligence: A definition of machine intelligence.Minds and Machines, 17(4):391–444, 2007

    Shane Legg and Marcus Hutter. Universal intelligence: A definition of machine intelligence.Minds and Machines, 17(4):391–444, 2007

  22. [22]

    Levinson and Francisco Torreira

    Stephen C. Levinson and Francisco Torreira. Timing in turn-taking and its implications for processing models of language.Frontiers in Psychology, 6:731, 2015

  23. [23]

    Mind perception of artificial intelligence: A systematic review.Information Processing & Management, 63(7):104820, 2026

    Yingming Li, Xin Qin, Kai Chi Yam, Kurt Gray, Xiang Cheng, Junqi Wen, and Xiaoyang Lai. Mind perception of artificial intelligence: A systematic review.Information Processing & Management, 63(7):104820, 2026

  24. [24]

    Bertram F. Malle. How many dimensions of mind perception really are there? InProceedings of the 41st Annual Meeting of the Cognitive Science Society (CogSci 2019), pages 2268–2274, Montreal, QC, 2019

  25. [25]

    An oz-centric review of interactive drama and believable agents

    Michael Mateas. An oz-centric review of interactive drama and believable agents. InArtificial Intel- ligence Today: Recent Trends and Developments, volume 1600 ofLecture Notes in Computer Science, pages 297–328. Springer, Berlin, Heidelberg, 1999

  26. [26]

    MacDorman, and Norri Kageki

    Masahiro Mori, Karl F. MacDorman, and Norri Kageki. The uncanny valley [from the field].IEEE Robotics & Automation Magazine, 19(2):98–100, 2012

  27. [27]

    Machines and mindlessness: Social responses to computers.Journal of Social Issues, 56(1):81–103, 2000

    Clifford Nass and Youngme Moon. Machines and mindlessness: Social responses to computers.Journal of Social Issues, 56(1):81–103, 2000

  28. [28]

    O’Brien, Carrie J

    Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23), 2023

  29. [29]

    Exploring relationship development with social chat- bots: A mixed-method study of Replika.Computers in Human Behavior, 140:107600, 2023

    Iryna Pentina, Tyler Hancock, and Tianling Xie. Exploring relationship development with social chat- bots: A mixed-method study of Replika.Computers in Human Behavior, 140:107600, 2023

  30. [30]

    CSLI Publications and Cambridge University Press, Stanford, CA and New York, 1996

    Byron Reeves and Clifford Nass.The Media Equation: How People Treat Computers, Television, and New Media Like Real People and Places. CSLI Publications and Cambridge University Press, Stanford, CA and New York, 1996

  31. [31]

    Schegloff, and Gail Jefferson

    Harvey Sacks, Emanuel A. Schegloff, and Gail Jefferson. A simplest systematics for the organization of turn-taking for conversation.Language, 50(4):696–735, 1974. 11

  32. [32]

    Role play with large language models.Nature, 623(7987):493–498, 2023

    Murray Shanahan, Kyle McDonell, and Laria Reynolds. Role play with large language models.Nature, 623(7987):493–498, 2023

  33. [33]

    My chatbot com- panion – a study of human-chatbot relationships.International Journal of Human-Computer Studies, 149:102601, 2021

    Marita Skjuve, Asbjørn Følstad, Knut Inge Fostervold, and Petter Bae Brandtzæg. My chatbot com- panion – a study of human-chatbot relationships.International Journal of Human-Computer Studies, 149:102601, 2021

  34. [34]

    James W. A. Strachan, Dalila Albergo, Giulia Borghini, Oriana Pansardi, Eugenio Scaliti, Saurabh Gupta, Krati Saxena, Alessandro Rufo, Stefano Panzeri, Guido Manzi, Michael S. A. Graziano, and Cristina Becchio. Testing theory of mind in large language models and humans.Nature Human Be- haviour, 8(7):1285–1295, 2024

  35. [35]

    Szczuka, Lisa M¨ uhl, and Tanja Schneeberger

    Jessica M. Szczuka, Lisa M¨ uhl, and Tanja Schneeberger. Intimacy by design: Definition, state of research, and interdisciplinary research agenda on intimate human-ai interactions.AI & Society, 2026

  36. [36]

    Williams, John Omerod, and Eliza Bliss-Moreau

    Kallie Tzelios, Lisa A. Williams, John Omerod, and Eliza Bliss-Moreau. Evidence of the unidimensional structure of mind perception.Scientific Reports, 12:18978, 2022

  37. [37]

    Large language models fail on trivial alterations to theory-of-mind tasks

    Tomer Ullman. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399, 2023

  38. [38]

    Weisz, and Mei Si

    Qiaosi Wang, Joel Wester, Marvin Pafla, Minha Lee, Justin D. Weisz, and Mei Si. ToMinHAI at CUI ’2025: Theory of mind in human-cui interaction. InProceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25), 2025

  39. [39]

    Cacioppo, and Nicholas Epley

    Adam Waytz, John T. Cacioppo, and Nicholas Epley. Who sees human? the stability and importance of individual differences in anthropomorphism.Perspectives on Psychological Science, 5(3):219–232, 2010

  40. [40]

    Dweck, and Ellen M

    Kara Weisman, Carol S. Dweck, and Ellen M. Markman. Rethinking people’s conceptions of mental life.Proceedings of the National Academy of Sciences, 114(43):11374–11379, 2017. 12