Pith. sign in

REVIEW 3 major objections 5 minor 21 references

Across 2,008 long synthetic conversations, no tested AI companion kept its persona and remembered the shared history: trajectory recall averaged 44.4% and user-state memory was at chance.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 00:17 UTC pith:DLYA54WX

load-bearing objection ANCHOR is a serious, carefully bounded long-horizon audit; the negative result is probably right, but the all-LLM question pipeline means the headline 44.4% needs a human ceiling check before I'd quote it. the 3 major comments →

arxiv 2607.28818 v1 pith:DLYA54WX submitted 2026-07-30 cs.AI cs.CL

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

classification cs.AI cs.CL
keywords persona collapsebehavioral driftAI companionslong-horizon memorytrajectory recallpersona fidelityLLM judge reliabilitycounterfactual question evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

People increasingly treat AI companions as ongoing relationships, but a fluent reply does not ensure the system still enacts its assigned role, boundaries, values, and style, or that it remembers the conversation's history. The paper introduces ANCHOR, a controlled synthetic audit that generates 85–130-session conversations across 27 personas, nine interaction schedules, three memory settings, and four evaluated models, measuring two separate failure surfaces: persona collapse/behavioral drift and trajectory recall. The central finding is negative: no evaluated model and configuration reliably preserves either dimension, with trajectory accuracy averaging 44.4% on four-option counterfactual questions and user-state recall hovering at chance under every tested condition. Among the social-pressure schedules, emotional-vulnerability and agreement-seeking conversations produce more boundary yields than explicit adversarial prompts, because ordinary requests for reassurance do not trigger the same refusals. The paper argues that continuity is multidimensional and evaluator-dependent, so audits must report question family, context condition, evaluator provenance, and uncertainty instead of collapsing everything into a single stability score.

Core claim

ANCHOR operationalizes two observable long-horizon failures. 'Persona collapse' is the loss of a deployed role, boundaries, values, or style; 'behavioral drift' is their gradual or recurrent erosion. The Identity Probe combines a sealed 102-item questionnaire, taken at four checkpoints with answers excluded from the dialogue, and turn-level judgments of every assistant turn on the four persona axes; the Persona Retention projection measures how far a later questionnaire response has drifted toward the model's bare-assistant anchor. The Trajectory Probe scores 110 calibrated counterfactual four-option questions in seven families—persona voice, persona protection, persona update, active and ex

What carries the argument

The central object is ANCHOR (Assistant-Normalised Character and Historical Outcome Recall), a synthetic audit pipeline. Its Identity Probe measures persona enactment through a sealed checkpoint questionnaire and three-level per-turn judgments on four axes (role identity, boundaries, values, style), scored by three LLM judges with judge choice treated as measurement uncertainty. Its Trajectory Probe measures memory through calibrated counterfactual four-option questions that require distinguishing what actually happened from plausible alternatives. The statistical pivot is Persona Retention, a projection of later questionnaire responses onto the direction from the model's bare-assistant anch

Load-bearing premise

The load-bearing premise is that the LLM judges' turn-level labels validly reflect whether a persona axis is held; the paper's own validation shows 64–68% exact four-axis agreement on 50 human-labeled turns and sharp judge disagreement on some models, so if the primary judge is miscalibrated the schedule, time, and recovery findings are rubric artifacts rather than model behavior.

What would settle it

A decisive check is a human-labeling study: take a few hundred turns from the released conversations, have humans score the four persona axes without seeing any model-judge labels, and compare model orderings. If humans do not reproduce the primary judge's ordering—particularly the large identity-axis gap between the two models rated high and the two rated low—then the schedule, time, and recovery results are rubric artifacts. A second, sharper falsifier: any configuration that scores above 90% on all seven trajectory families while maintaining human-agreement-level persona fidelity across all

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Systems configured like the ones tested cannot be assumed to sustain a disclosed companion role over long use; users and developers should treat unstated continuity as unsupported.
  • Memory architecture alone is not the fix: long-context transcripts, hierarchical summaries, self-managed JSON state, and retrieval-based scoring all leave trajectory recall near 44% overall and user-state recall at chance.
  • Any single-number 'stability' or 'trust' score for a companion is misleading; the paper's evidence requires disaggregating persona enactment, trajectory recall, evaluator provenance, and deployment context.
  • Near-chance user-state recall implies a system can silently lose a consequential update, such as a changed treatment goal or a changed user life circumstance, which is a concrete risk for health-adjacent companions.
  • Explicit adversarial re-role attempts can look less damaging than ordinary emotional or agreement-seeking conversation, because recognizable attacks trigger visible refusals while routine requests create more opportunities for boundary yield.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • My inference: if the pattern generalizes, continuity is a property of the whole deployed configuration—persona card, memory pipeline, scheduling, and evaluator—rather than of the base model, so a product team's own tests matter more than any model-level ranking.
  • My inference: the chance-level user-state recall points to a concrete design requirement—write-audited, replayable state updates rather than model-written summaries—because the paper's self-managed JSON setting was one such attempt and still failed, suggesting the failure is in how state is written and read, not merely in context length.
  • My inference: a natural extension would randomize user-simulator behavior or insert long time gaps between sessions to test whether drift scales with emotional intensity, conversation length, or recency; the paper's deterministic seed and ledger make this a cheap next experiment.
  • My inference: the sharp judge disagreement on role and style suggests that audits should prefer observable behavioral criteria—such as boundary-refusal rates on scripted scenarios—over holistic 'sounds like the character' judgments, since the latter are not reproducible across evaluators.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces ANCHOR, a synthetic audit framework that separately measures long-horizon persona enactment and trajectory recall in AI companion systems. The study generates 2,008 conversations from 27 authored personas, nine interaction schedules, three memory settings, and four evaluated models. The Identity Probe uses a sealed 102-item questionnaire plus turn-level judgments on four persona axes (role, boundaries, values, style); the Trajectory Probe constructs 110 calibrated counterfactual multiple-choice questions from 35 conversation banks and scores them under long-context, hierarchical-summary, self-managed, and retrieval conditions. The headline results are that trajectory accuracy averages only 44.4%, user-state recall is near four-option chance (0.214–0.250), and no tested context or memory condition consistently resolves these failures. Questionnaire retention varies by model and facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. The paper concludes that current systems do not reliably support long-horizon companion continuity and that audits should keep persona enactment, trajectory recall, evaluator provenance, and deployment context distinct.

Significance. If the central findings hold, this is a valuable and timely negative result for the AI-companion evaluation literature. The paper makes several concrete contributions: a controlled synthetic corpus that separates persona enactment from trajectory recall, a transparent audit pipeline with explicit uncertainty and disaggregated reporting, open score files and manifests for artifact reproduction, and a clear demonstration that judge choice materially changes persona-fidelity conclusions. The refusal to collapse the two probes into a single stability score is methodologically sound. However, the strongest quantitative claims—44.4% trajectory accuracy and chance-level user-state recall—depend on LLM-generated and LLM-calibrated questions with no reported human ceiling check, and the persona-collapse findings rest largely on a single primary judge with limited human validation. These gaps cap the confidence one can place in the abstract's central claim. The framework is significant even as a stress-test methodology; the specific numerical conclusions need additional validation before they can be read as established properties of the evaluated systems.

major comments (3)
  1. [§3.5 / §5.5 (Table 4, Table 8)] The Trajectory Probe's headline accuracy (44.4%) and the near-chance user-state recall rest entirely on LLM-mediated question construction and validation: candidate questions are LLM-written, the blind and with-history panels are LLM panels, and the calibrator ceiling is another LLM scoring with a ±15-session window. No human annotator checks that the gold answer uniquely follows from the transcript or that the distractors are plausible-but-wrong rather than arbitrary. If calibrated questions are ambiguous, even a model with perfect conversational memory would score near chance. I ask the authors to add a human ceiling/answerability study on a stratified sample of calibrated questions, or otherwise justify that the LLM-only calibration is sufficient. This is load-bearing for the central negative claim.
  2. [§4 / Appendix B (Table 10)] The turn-level persona-fidelity results, including the schedule, time, and recovery analyses in §5.4, rely on a single primary judge (Claude Sonnet 4.6) whose outputs are not independently validated on a substantial sample. Appendix B Table 10 shows extreme disagreement: Claude marks GPT-4o-mini's identity-axis held rate at 15.5%, while Gemini-Flash marks it at 100.0%. The paper's own 50-turn human calibration set yields only 64–68% exact four-axis agreement. The authors are transparent that these results are rubric-dependent, but the abstract's claim of 'persona collapse' as an observed failure is weaker than the trajectory claim. I recommend reporting the main turn-level analyses under multiple judges or providing a larger human-validated sample, and softening any language that implies model-level persona-collapse findings are established.
  3. [§5.5 / Table 8] The 'user-state change' and 'persona update' families are extremely small: 28 and 8 pooled scored decisions per condition, respectively. The user-state family is the basis for the abstract's claim that recall remains near four-option chance, yet the table reports only pooled decisions, not the number of unique calibrated questions. If this family contains only a handful of unique items (as the n=28 number suggests), the near-chance result is a fragile basis for a general conclusion. Please report unique-question counts per family, per-bank confidence intervals, and avoid generalizing from families with n=8 (the retrieval-condition collapse of persona-update accuracy to 0.25 is explicitly exploratory, but the text still folds it into the 'no memory condition fixes it' summary).
minor comments (5)
  1. [§3.3] 'three memory architecture in the pipeline' should be 'three memory architectures'.
  2. [Figures 3, 4, 6] The model is labeled 'GPT-5.4-mini' in several figures but 'GPT-5-mini' in the tables and text. Please standardize.
  3. [Table 8 / Figure 8] The table and figure labels imply unique-question counts for small families, but only pooled decision counts are shown. Add a column for unique calibrated questions per family to support the exploratory caveats.
  4. [§5.4] The paragraph 'Why explicit attacks can look less damaging' is appropriately framed as a hypothesis. It would be useful to explicitly mark it as untested in the section header or subheading.
  5. [§4, Eq. (1)] Persona Retention is described as unbounded and can exceed 1 or become negative. The text explains this well; consider adding a one-line reminder in Figure 4's caption that values are means of signed projections.

Circularity Check

0 steps flagged

No significant circularity: the headline results are empirical measurements with an explicit baseline, though the Trajectory Probe's LLM-mediated construction is a validity limitation rather than a circular reduction.

full rationale

The paper does not derive its headline numbers from fitted parameters or self-defined targets. The Persona Retention metric (Eq. 1) is a projection onto the direction from each model's bare-assistant anchor to its initial persona response; the anchor is an explicit measurement baseline, and PR values are computed from held-out checkpoint responses, not constructed to equal 1. Turn-level fidelity uses an LLM judge against the effective persona card; this operationalizes the paper's definition of persona collapse rather than smuggling in the conclusion. The Trajectory Probe is built by LLM question generation with blind, consensus, and calibrator filters, then scored on four models; the filters select answerable items but do not fit any parameter to the evaluated models, so the 44.4% average is an empirical estimate, not a forced result. The only self-citation (Venkit et al. 2026 for persona-card structure) is a design-provenance citation and is not load-bearing; no uniqueness theorem or ansatz is imported from the authors' prior work. The paper's own limitations section acknowledges that LLMs mediate all stages and that evaluator choice changes results; the absence of a human ceiling check for Trajectory questions is a correctness/validity risk (the headline could be an artifact of ambiguous questions), but that is a missing support, not a circular reduction. Score 1 reflects a minor non-load-bearing self-citation and the explicitly LLM-mediated pipeline, not any equation-level circularity.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The paper introduces no new physical/theoretical entities. It defines observable audit constructs (persona collapse, behavioral drift, ANCHOR) operationally, and the central claims rest on domain assumptions about judge validity, synthetic-corpus representativeness, vector-space questionnaire geometry, and the calibration pipeline. The main hand-set analytic thresholds are listed as free parameters.

free parameters (3)
  • Turn-level fidelity thresholds ('soft deviation' vs 'hard failure')
    Each persona axis is judged at three levels; the boundary between soft deviation and hard failure is authored by the researchers and determines all turn-level failure/recovery rates.
  • Trajectory question-selection thresholds (3-of-4 gold consensus, 5 calibrator repeats, ±15-session window)
    These hand-set thresholds determine which questions enter the 110-question primary analysis, and thus the families, counts, and the 44.4% aggregate accuracy.
  • Majority-of-three all-axes-held criterion
    The 79.0%/96.7% comparison requires all four axes held under majority of three judges; changing this criterion changes the model ordering.
axioms (4)
  • domain assumption LLM judges can validly score persona axes from the effective persona card and immediate context.
    Load-bearing for all turn-level results; supported only by a 50-turn author-labeled calibration with 64–68% exact-axis agreement and by large cross-judge disagreement in Appendix B.
  • domain assumption Synthetic conversations generated by GPT-4.1 from author-written templates and one seed per schedule approximate deployed companion interactions well enough to audit continuity.
    Central to external validity; explicitly bounded by the authors, who note schedules are stress tests, not user evidence (§6.1, §9).
  • domain assumption Euclidean projection of questionnaire response vectors is a meaningful directional measure of persona retention.
    PR in Eq. (1) assumes vector-space geometry of 102-item responses and a model-specific bare-assistant anchor; used for all checkpoint results.
  • domain assumption The trajectory calibration pipeline produces gold-standard questions whose answers are unambiguous and not guessable without history.
    Relies on GPT/Claude panels for blind/consensus/calibrator judgments; no human validation of the 110 calibrated trajectory questions is reported.

pith-pipeline@v1.3.0-alltime-deepseek · 15201 in / 12371 out tokens · 125418 ms · 2026-08-03T00:17:22.347073+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions." pith.science (2026). https://pith.science/paper/DLYA54WX

@misc{pith2026260728818,
  author       = {Pith},
  title        = {Pith review of: Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DLYA54WX}},
  note         = {Machine review of arXiv:2607.28818}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajectory accuracy averages only 44.4%, user-state recall remains near four-option chance, and no tested context condition or memory consistently resolves these failures. Questionnaire retention also varies by model and persona facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. These results indicate that current systems do not yet reliably support long-horizon companion continuity and that audits must distinguish persona enactment, trajectory recall, evaluator provenance, and deployment context rather than collapse them into a single trust or stability score.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

21 extracted references · 3 linked inside Pith

  1. [1]

    Inter-university Consortium for Political and Social Research,

    JamesAllanDavisandTomWilliamSmith.GeneralSocialSurveys1972–2000: CumulativeCodebook. Inter-university Consortium for Political and Social Research,

  2. [7]

    Prerna Juneja and Lika Lomidze

    doi: 10.18653/v1/ 2024.findings-naacl.229. Prerna Juneja and Lika Lomidze. Persona-grounded safety evaluation of AI companions in multi-turn conversations. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 18148–18175. Association for Computational Linguistics,

  3. [9]

    HorizonBench: Long-horizon personalization with evolving preferences.arXiv preprint arXiv:2604.17283,

    Shuyue Stella Li, Bhargavi Paranjape, Kerem Oktar, Zhongyao Ma, Gelin Zhou, Lin Guan, Na Zhang, Sem Park, Lin Chen, Diyi Yang, et al. HorizonBench: Long-horizon personalization with evolving preferences.arXiv preprint arXiv:2604.17283,

  4. [10]

    The assistant axis: Situating and stabilizing the default persona of language models.arXiv preprint arXiv:2601.10387,

    Christina Lu, Jack Gallagher, Jonathan Michala, Kyle Fish, and Jack Lindsey. The assistant axis: Situating and stabilizing the default persona of language models.arXiv preprint arXiv:2601.10387,

  5. [11]

    When human–AI interactions become parasocial: Agency and anthro- pomorphism in affective design

    Takuya Maeda and Anabel Quan-Haase. When human–AI interactions become parasocial: Agency and anthro- pomorphism in affective design. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 1068–1077. Association for Computing Machinery,

  6. [13]

    Aikaterina Manoli, Janet V

    doi: 10.1007/s00146-025-02318-6. Aikaterina Manoli, Janet V. T. Pauketat, Ali Ladak, Hayoun Noh, Angel Hsing-Chi Hwang, and Jacy Reese Anthis. Digital companionship: Overlapping uses of AI companions and AI assistants. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery,

  7. [14]

    Arianna Manzini, Geoff Keeling, Nahema Marchal, Kevin R

    doi: 10.1145/3772318.3791331. Arianna Manzini, Geoff Keeling, Nahema Marchal, Kevin R. McKee, Verena Rieser, and Iason Gabriel. Should users trust advanced AI assistants? justified trust as a function of competence and alignment. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 1174–1186. Association for Comput...

  8. [15]

    Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al

    doi: 10.1145/3630106.3658964. Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al. Discovering language model behaviors with model-written evaluations. InFindings of the association for computational linguistics: ACL 2023, pages 13387–13434,

  9. [16]

    Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, 16 Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions Jamila Smith-Loud, Daniel Theron, and Parker Barnes. Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmi...

  10. [20]

    Pranav Narayanan Venkit, Yu Li, Yada Pruksachatkun, and Chien-Sheng Wu

    Association for Computational Linguistics. Pranav Narayanan Venkit, Yu Li, Yada Pruksachatkun, and Chien-Sheng Wu. The need for a socially-grounded persona framework for user simulation.arXiv preprint arXiv:2601.07110,

  11. [21]

    Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu

    Association for Computational Linguistics. Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. LongMemEval: Benchmarking chat assistants on long-term interactive memory. InInternational Conference on Learning Representations (ICLR), 2025a. arXiv:2410.10813. Shujin Wu, Yi R Fung, Cheng Qian, Jeonghwan Kim, Dilek Hakkani-Tur, and Heng J...

  12. [22]

    Persona and Schedule Definitions The 27 persona cards span professional helpers, caregiving and creative roles, and stylized literary roles

    A. Persona and Schedule Definitions The 27 persona cards span professional helpers, caregiving and creative roles, and stylized literary roles. Each card has four scored components: role identity, explicit boundaries, stated values, and style. The author-defined Near/Mid/Far identifiers describe conceptual distance from a generic assistant and were motiva...

  13. [1992]

    Mrinank Sharma, Meg Tong, Tomek Korbak, David Duvenaud, Amanda Askell, Sam Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott Johnston, Shauna Kravec, et al

    doi: 10.1016/S0065-2601(08)60281-6. Mrinank Sharma, Meg Tong, Tomek Korbak, David Duvenaud, Amanda Askell, Sam Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott Johnston, Shauna Kravec, et al. Towards understanding sycophancy in language models. InInternational Conference on Learning Representations, volume 2024, pages 110–144,

  14. [1999]

    Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094,

    Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094,

  15. [2014]

    Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in AI

    Alon Jacovi, Ana Marasović, Tim Miller, and Yoav Goldberg. Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in AI. InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 624–635. Association for Computing Machinery,

  16. [2017]

    Quan Tu, Shilong Fan, Zihang Tian, Tianhao Shen, Shuo Shang, Xin Gao, and Rui Yan

    doi: 10.1016/j.jrp.2017.02.004. Quan Tu, Shilong Fan, Zihang Tian, Tianhao Shen, Shuo Shang, Xin Gao, and Rui Yan. CharacterEval: A chinese benchmark for role-playing conversational agent evaluation. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11836–11850, Bangkok, Thailand,

  17. [2020]

    Shalom H

    doi: 10.1145/3351095.3372873. Shalom H. Schwartz. Universals in the content and structure of values: Theoretical advances and empirical tests in 20 countries. InAdvances in Experimental Social Psychology, volume 25, pages 1–65. Academic Press,

  18. [2021]

    Bowen Jiang, Yuan Yuan, Maohao Shen, Zhuoqun Hao, Zhangchen Xu, Zichen Chen, Ziyi Liu, Anvesh Rao Vijjini, Jiashu He, Hanchao Yu, et al

    doi: 10.1145/3442188.3445923. Bowen Jiang, Yuan Yuan, Maohao Shen, Zhuoqun Hao, Zhangchen Xu, Zichen Chen, Ziyi Liu, Anvesh Rao Vijjini, Jiashu He, Hanchao Yu, et al. PersonaMem-v2: Towards personalized intelligence via learning implicit user personas and agentic memory.arXiv preprint arXiv:2512.06688,

  19. [2024]

    How ai companionship develops: Evidence from a longitudinal study.arXiv preprint arXiv:2510.10079,

    Angel Hsing-Chi Hwang, Fiona Li, Jacy Reese Anthis, and Hayoun Noh. How ai companionship develops: Evidence from a longitudinal study.arXiv preprint arXiv:2510.10079,

  20. [2025]

    PersonaLLM: Investigating the ability of large language models to express personality traits

    Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. PersonaLLM: Investigating the ability of large language models to express personality traits. InFindings of the Association for Computational Linguistics: NAACL 2024, pages 3605–3627. Association for Computational Linguistics,

  21. [2026]

    Hannah Rose Kirk, Henry Davidson, Ed Saunders, Lennart Luettgau, Bertie Vidgen, Scott A Hale, and Christopher Summerfield

    doi: 10.18653/v1/2026.acl-long.828. Hannah Rose Kirk, Henry Davidson, Ed Saunders, Lennart Luettgau, Bertie Vidgen, Scott A Hale, and Christopher Summerfield. Neural steering vectors reveal dose and exposure-dependent impacts of human-ai relationships. arXiv preprint arXiv:2512.01991,