REVIEW 3 major objections 5 minor 21 references
Across 2,008 long synthetic conversations, no tested AI companion kept its persona and remembered the shared history: trajectory recall averaged 44.4% and user-state memory was at chance.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Across 2,008 synthetic long-horizon conversations, four AI companion models all fail to reliably preserve persona and trajectory memory; trajectory accuracy averages only 44.4%.
T0 review reviewed 2026-08-03 challenge →
load-bearing objection ANCHOR is a serious, carefully bounded long-horizon audit; the negative result is probably right, but the all-LLM question pipeline means the headline 44.4% needs a human ceiling check before I'd quote it. the 3 major comments →
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
ANCHOR operationalizes two observable long-horizon failures. 'Persona collapse' is the loss of a deployed role, boundaries, values, or style; 'behavioral drift' is their gradual or recurrent erosion. The Identity Probe combines a sealed 102-item questionnaire, taken at four checkpoints with answers excluded from the dialogue, and turn-level judgments of every assistant turn on the four persona axes; the Persona Retention projection measures how far a later questionnaire response has drifted toward the model's bare-assistant anchor. The Trajectory Probe scores 110 calibrated counterfactual four-option questions in seven families—persona voice, persona protection, persona update, active and ex
What carries the argument
The central object is ANCHOR (Assistant-Normalised Character and Historical Outcome Recall), a synthetic audit pipeline. Its Identity Probe measures persona enactment through a sealed checkpoint questionnaire and three-level per-turn judgments on four axes (role identity, boundaries, values, style), scored by three LLM judges with judge choice treated as measurement uncertainty. Its Trajectory Probe measures memory through calibrated counterfactual four-option questions that require distinguishing what actually happened from plausible alternatives. The statistical pivot is Persona Retention, a projection of later questionnaire responses onto the direction from the model's bare-assistant anch
Load-bearing premise
The load-bearing premise is that the LLM judges' turn-level labels validly reflect whether a persona axis is held; the paper's own validation shows 64–68% exact four-axis agreement on 50 human-labeled turns and sharp judge disagreement on some models, so if the primary judge is miscalibrated the schedule, time, and recovery findings are rubric artifacts rather than model behavior.
What would settle it
A decisive check is a human-labeling study: take a few hundred turns from the released conversations, have humans score the four persona axes without seeing any model-judge labels, and compare model orderings. If humans do not reproduce the primary judge's ordering—particularly the large identity-axis gap between the two models rated high and the two rated low—then the schedule, time, and recovery results are rubric artifacts. A second, sharper falsifier: any configuration that scores above 90% on all seven trajectory families while maintaining human-agreement-level persona fidelity across all
If this is right
- Systems configured like the ones tested cannot be assumed to sustain a disclosed companion role over long use; users and developers should treat unstated continuity as unsupported.
- Memory architecture alone is not the fix: long-context transcripts, hierarchical summaries, self-managed JSON state, and retrieval-based scoring all leave trajectory recall near 44% overall and user-state recall at chance.
- Any single-number 'stability' or 'trust' score for a companion is misleading; the paper's evidence requires disaggregating persona enactment, trajectory recall, evaluator provenance, and deployment context.
- Near-chance user-state recall implies a system can silently lose a consequential update, such as a changed treatment goal or a changed user life circumstance, which is a concrete risk for health-adjacent companions.
- Explicit adversarial re-role attempts can look less damaging than ordinary emotional or agreement-seeking conversation, because recognizable attacks trigger visible refusals while routine requests create more opportunities for boundary yield.
Where Pith is reading between the lines
- My inference: if the pattern generalizes, continuity is a property of the whole deployed configuration—persona card, memory pipeline, scheduling, and evaluator—rather than of the base model, so a product team's own tests matter more than any model-level ranking.
- My inference: the chance-level user-state recall points to a concrete design requirement—write-audited, replayable state updates rather than model-written summaries—because the paper's self-managed JSON setting was one such attempt and still failed, suggesting the failure is in how state is written and read, not merely in context length.
- My inference: a natural extension would randomize user-simulator behavior or insert long time gaps between sessions to test whether drift scales with emotional intensity, conversation length, or recency; the paper's deterministic seed and ledger make this a cheap next experiment.
- My inference: the sharp judge disagreement on role and style suggests that audits should prefer observable behavioral criteria—such as boundary-refusal rates on scripted scenarios—over holistic 'sounds like the character' judgments, since the latter are not reproducible across evaluators.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ANCHOR, a synthetic audit framework that separately measures long-horizon persona enactment and trajectory recall in AI companion systems. The study generates 2,008 conversations from 27 authored personas, nine interaction schedules, three memory settings, and four evaluated models. The Identity Probe uses a sealed 102-item questionnaire plus turn-level judgments on four persona axes (role, boundaries, values, style); the Trajectory Probe constructs 110 calibrated counterfactual multiple-choice questions from 35 conversation banks and scores them under long-context, hierarchical-summary, self-managed, and retrieval conditions. The headline results are that trajectory accuracy averages only 44.4%, user-state recall is near four-option chance (0.214–0.250), and no tested context or memory condition consistently resolves these failures. Questionnaire retention varies by model and facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. The paper concludes that current systems do not reliably support long-horizon companion continuity and that audits should keep persona enactment, trajectory recall, evaluator provenance, and deployment context distinct.
Significance. If the central findings hold, this is a valuable and timely negative result for the AI-companion evaluation literature. The paper makes several concrete contributions: a controlled synthetic corpus that separates persona enactment from trajectory recall, a transparent audit pipeline with explicit uncertainty and disaggregated reporting, open score files and manifests for artifact reproduction, and a clear demonstration that judge choice materially changes persona-fidelity conclusions. The refusal to collapse the two probes into a single stability score is methodologically sound. However, the strongest quantitative claims—44.4% trajectory accuracy and chance-level user-state recall—depend on LLM-generated and LLM-calibrated questions with no reported human ceiling check, and the persona-collapse findings rest largely on a single primary judge with limited human validation. These gaps cap the confidence one can place in the abstract's central claim. The framework is significant even as a stress-test methodology; the specific numerical conclusions need additional validation before they can be read as established properties of the evaluated systems.
major comments (3)
- [§3.5 / §5.5 (Table 4, Table 8)] The Trajectory Probe's headline accuracy (44.4%) and the near-chance user-state recall rest entirely on LLM-mediated question construction and validation: candidate questions are LLM-written, the blind and with-history panels are LLM panels, and the calibrator ceiling is another LLM scoring with a ±15-session window. No human annotator checks that the gold answer uniquely follows from the transcript or that the distractors are plausible-but-wrong rather than arbitrary. If calibrated questions are ambiguous, even a model with perfect conversational memory would score near chance. I ask the authors to add a human ceiling/answerability study on a stratified sample of calibrated questions, or otherwise justify that the LLM-only calibration is sufficient. This is load-bearing for the central negative claim.
- [§4 / Appendix B (Table 10)] The turn-level persona-fidelity results, including the schedule, time, and recovery analyses in §5.4, rely on a single primary judge (Claude Sonnet 4.6) whose outputs are not independently validated on a substantial sample. Appendix B Table 10 shows extreme disagreement: Claude marks GPT-4o-mini's identity-axis held rate at 15.5%, while Gemini-Flash marks it at 100.0%. The paper's own 50-turn human calibration set yields only 64–68% exact four-axis agreement. The authors are transparent that these results are rubric-dependent, but the abstract's claim of 'persona collapse' as an observed failure is weaker than the trajectory claim. I recommend reporting the main turn-level analyses under multiple judges or providing a larger human-validated sample, and softening any language that implies model-level persona-collapse findings are established.
- [§5.5 / Table 8] The 'user-state change' and 'persona update' families are extremely small: 28 and 8 pooled scored decisions per condition, respectively. The user-state family is the basis for the abstract's claim that recall remains near four-option chance, yet the table reports only pooled decisions, not the number of unique calibrated questions. If this family contains only a handful of unique items (as the n=28 number suggests), the near-chance result is a fragile basis for a general conclusion. Please report unique-question counts per family, per-bank confidence intervals, and avoid generalizing from families with n=8 (the retrieval-condition collapse of persona-update accuracy to 0.25 is explicitly exploratory, but the text still folds it into the 'no memory condition fixes it' summary).
minor comments (5)
- [§3.3] 'three memory architecture in the pipeline' should be 'three memory architectures'.
- [Figures 3, 4, 6] The model is labeled 'GPT-5.4-mini' in several figures but 'GPT-5-mini' in the tables and text. Please standardize.
- [Table 8 / Figure 8] The table and figure labels imply unique-question counts for small families, but only pooled decision counts are shown. Add a column for unique calibrated questions per family to support the exploratory caveats.
- [§5.4] The paragraph 'Why explicit attacks can look less damaging' is appropriately framed as a hypothesis. It would be useful to explicitly mark it as untested in the section header or subheading.
- [§4, Eq. (1)] Persona Retention is described as unbounded and can exceed 1 or become negative. The text explains this well; consider adding a one-line reminder in Figure 4's caption that values are means of signed projections.
Circularity Check
No significant circularity: the headline results are empirical measurements with an explicit baseline, though the Trajectory Probe's LLM-mediated construction is a validity limitation rather than a circular reduction.
full rationale
The paper does not derive its headline numbers from fitted parameters or self-defined targets. The Persona Retention metric (Eq. 1) is a projection onto the direction from each model's bare-assistant anchor to its initial persona response; the anchor is an explicit measurement baseline, and PR values are computed from held-out checkpoint responses, not constructed to equal 1. Turn-level fidelity uses an LLM judge against the effective persona card; this operationalizes the paper's definition of persona collapse rather than smuggling in the conclusion. The Trajectory Probe is built by LLM question generation with blind, consensus, and calibrator filters, then scored on four models; the filters select answerable items but do not fit any parameter to the evaluated models, so the 44.4% average is an empirical estimate, not a forced result. The only self-citation (Venkit et al. 2026 for persona-card structure) is a design-provenance citation and is not load-bearing; no uniqueness theorem or ansatz is imported from the authors' prior work. The paper's own limitations section acknowledges that LLMs mediate all stages and that evaluator choice changes results; the absence of a human ceiling check for Trajectory questions is a correctness/validity risk (the headline could be an artifact of ambiguous questions), but that is a missing support, not a circular reduction. Score 1 reflects a minor non-load-bearing self-citation and the explicitly LLM-mediated pipeline, not any equation-level circularity.
Axiom & Free-Parameter Ledger
free parameters (3)
- Turn-level fidelity thresholds ('soft deviation' vs 'hard failure')
- Trajectory question-selection thresholds (3-of-4 gold consensus, 5 calibrator repeats, ±15-session window)
- Majority-of-three all-axes-held criterion
axioms (4)
- domain assumption LLM judges can validly score persona axes from the effective persona card and immediate context.
- domain assumption Synthetic conversations generated by GPT-4.1 from author-written templates and one seed per schedule approximate deployed companion interactions well enough to audit continuity.
- domain assumption Euclidean projection of questionnaire response vectors is a meaningful directional measure of persona retention.
- domain assumption The trajectory calibration pipeline produces gold-standard questions whose answers are unambiguous and not guessable without history.
Cite this review
Pith. "Pith review of Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions." pith.science (2026). https://pith.science/paper/DLYA54WX
@misc{pith2026260728818,
author = {Pith},
title = {Pith review of: Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions},
year = {2026},
howpublished = {\url{https://pith.science/paper/DLYA54WX}},
note = {Machine review of arXiv:2607.28818}
}
read the original abstract
As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajectory accuracy averages only 44.4%, user-state recall remains near four-option chance, and no tested context condition or memory consistently resolves these failures. Questionnaire retention also varies by model and persona facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. These results indicate that current systems do not yet reliably support long-horizon companion continuity and that audits must distinguish persona enactment, trajectory recall, evaluator provenance, and deployment context rather than collapse them into a single trust or stability score.
Reference graph
Works this paper leans on
-
[1]
Inter-university Consortium for Political and Social Research,
JamesAllanDavisandTomWilliamSmith.GeneralSocialSurveys1972–2000: CumulativeCodebook. Inter-university Consortium for Political and Social Research,
2000
-
[7]
Prerna Juneja and Lika Lomidze
doi: 10.18653/v1/ 2024.findings-naacl.229. Prerna Juneja and Lika Lomidze. Persona-grounded safety evaluation of AI companions in multi-turn conversations. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 18148–18175. Association for Computational Linguistics,
doi:10.18653/v1/ 2024
-
[9]
Shuyue Stella Li, Bhargavi Paranjape, Kerem Oktar, Zhongyao Ma, Gelin Zhou, Lin Guan, Na Zhang, Sem Park, Lin Chen, Diyi Yang, et al. HorizonBench: Long-horizon personalization with evolving preferences.arXiv preprint arXiv:2604.17283,
-
[10]
Christina Lu, Jack Gallagher, Jonathan Michala, Kyle Fish, and Jack Lindsey. The assistant axis: Situating and stabilizing the default persona of language models.arXiv preprint arXiv:2601.10387,
-
[11]
When human–AI interactions become parasocial: Agency and anthro- pomorphism in affective design
Takuya Maeda and Anabel Quan-Haase. When human–AI interactions become parasocial: Agency and anthro- pomorphism in affective design. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 1068–1077. Association for Computing Machinery,
2024
-
[13]
doi: 10.1007/s00146-025-02318-6. Aikaterina Manoli, Janet V. T. Pauketat, Ali Ladak, Hayoun Noh, Angel Hsing-Chi Hwang, and Jacy Reese Anthis. Digital companionship: Overlapping uses of AI companions and AI assistants. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery,
-
[14]
Arianna Manzini, Geoff Keeling, Nahema Marchal, Kevin R
doi: 10.1145/3772318.3791331. Arianna Manzini, Geoff Keeling, Nahema Marchal, Kevin R. McKee, Verena Rieser, and Iason Gabriel. Should users trust advanced AI assistants? justified trust as a function of competence and alignment. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 1174–1186. Association for Comput...
arXiv 2024
-
[15]
doi: 10.1145/3630106.3658964. Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al. Discovering language model behaviors with model-written evaluations. InFindings of the association for computational linguistics: ACL 2023, pages 13387–13434,
arXiv 2023
-
[16]
Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, 16 Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions Jamila Smith-Loud, Daniel Theron, and Parker Barnes. Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmi...
2020
-
[20]
Pranav Narayanan Venkit, Yu Li, Yada Pruksachatkun, and Chien-Sheng Wu
Association for Computational Linguistics. Pranav Narayanan Venkit, Yu Li, Yada Pruksachatkun, and Chien-Sheng Wu. The need for a socially-grounded persona framework for user simulation.arXiv preprint arXiv:2601.07110,
-
[21]
Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu
Association for Computational Linguistics. Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. LongMemEval: Benchmarking chat assistants on long-term interactive memory. InInternational Conference on Learning Representations (ICLR), 2025a. arXiv:2410.10813. Shujin Wu, Yi R Fung, Cheng Qian, Jeonghwan Kim, Dilek Hakkani-Tur, and Heng J...
-
[22]
Persona and Schedule Definitions The 27 persona cards span professional helpers, caregiving and creative roles, and stylized literary roles
A. Persona and Schedule Definitions The 27 persona cards span professional helpers, caregiving and creative roles, and stylized literary roles. Each card has four scored components: role identity, explicit boundaries, stated values, and style. The author-defined Near/Mid/Far identifiers describe conceptual distance from a generic assistant and were motiva...
2026
-
[1992]
doi: 10.1016/S0065-2601(08)60281-6. Mrinank Sharma, Meg Tong, Tomek Korbak, David Duvenaud, Amanda Askell, Sam Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott Johnston, Shauna Kravec, et al. Towards understanding sycophancy in language models. InInternational Conference on Learning Representations, volume 2024, pages 110–144,
-
[1999]
Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094,
Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094,
-
[2014]
Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in AI
Alon Jacovi, Ana Marasović, Tim Miller, and Yoav Goldberg. Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in AI. InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 624–635. Association for Computing Machinery,
2021
-
[2017]
Quan Tu, Shilong Fan, Zihang Tian, Tianhao Shen, Shuo Shang, Xin Gao, and Rui Yan
doi: 10.1016/j.jrp.2017.02.004. Quan Tu, Shilong Fan, Zihang Tian, Tianhao Shen, Shuo Shang, Xin Gao, and Rui Yan. CharacterEval: A chinese benchmark for role-playing conversational agent evaluation. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11836–11850, Bangkok, Thailand,
- [2020]
-
[2021]
doi: 10.1145/3442188.3445923. Bowen Jiang, Yuan Yuan, Maohao Shen, Zhuoqun Hao, Zhangchen Xu, Zichen Chen, Ziyi Liu, Anvesh Rao Vijjini, Jiashu He, Hanchao Yu, et al. PersonaMem-v2: Towards personalized intelligence via learning implicit user personas and agentic memory.arXiv preprint arXiv:2512.06688,
-
[2024]
How ai companionship develops: Evidence from a longitudinal study.arXiv preprint arXiv:2510.10079,
Angel Hsing-Chi Hwang, Fiona Li, Jacy Reese Anthis, and Hayoun Noh. How ai companionship develops: Evidence from a longitudinal study.arXiv preprint arXiv:2510.10079,
-
[2025]
PersonaLLM: Investigating the ability of large language models to express personality traits
Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. PersonaLLM: Investigating the ability of large language models to express personality traits. InFindings of the Association for Computational Linguistics: NAACL 2024, pages 3605–3627. Association for Computational Linguistics,
2024
-
[2026]
doi: 10.18653/v1/2026.acl-long.828. Hannah Rose Kirk, Henry Davidson, Ed Saunders, Lennart Luettgau, Bertie Vidgen, Scott A Hale, and Christopher Summerfield. Neural steering vectors reveal dose and exposure-dependent impacts of human-ai relationships. arXiv preprint arXiv:2512.01991,
arXiv 2026
This paper was first reviewed by deepseek-v4-flash on August 3, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.