{"id":"e243ba0c-76cb-4512-b453-b7db77f5962f","arxiv_id":"2607.28818","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Across 2,008 synthetic long-horizon conversations, four AI companion models all fail to reliably preserve persona and trajectory memory; trajectory accuracy averages only 44.4%.","lead":"An audit of AI companion chatbots shows that across 2,008 long synthetic conversations, none of four tested models reliably keeps its assigned persona or remembers the shared conversation history. The result suggests users cannot yet count on companion continuity, and that evaluation must separate role fidelity, memory, and evaluator effects rather than use one stability score.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Trajectory Probe's 44.4% headline rests on LLM-written questions with no human ceiling check; if the questions are ambiguous, the central negative claim could be an artifact.","rationale":"The reader's weakest assumption focused on the turn-level persona-fidelity judge, citing large cross-judge disagreements (§5.4, Table 10). That is a legitimate concern, but the abstract's strongest numerical claims—44.4% trajectory accuracy and user-state recall near chance—come from the Trajectory Probe, not the persona judge. Both probes depend on LLM evaluation, yet the Trajectory Probe has no human validation at all. If the trajectory questions are ambiguous or the gold answers are not uniquely recoverable, the paper's central negative conclusion could be an artifact of its own LLM-mediated question pipeline. This is not a question of consensus or style; it is a question of whether the benchmark measures what it claims. The paper does include a 'with-history' LLM panel requiring 3-of-4 agreement, but that only shows that some LLMs can be made to agree, not that a competent human reader would answer the same way. Human annotation of the trajectory questions is a concrete, feasible check. The paper is otherwise careful: it discloses private test sets, reports judge dependence, and avoids overclaiming. Therefore the appropriate verdict is CONDITIONAL rather than REJECT or UNCHANGED: the core benchmark needs one additional validity check before the headline result can be accepted as evidence about AI companions rather than about LLM-generated evaluation questions.","tokens_in":15453,"tokens_out":5664,"duration_ms":60784,"concrete_test":"Run a human ceiling study: give 3–5 human annotators the full transcripts (without the event ledger) for a stratified sample of the 110 calibrated trajectory questions, ask them to select the correct option, and compare to the gold answers. Pre-register a passing bar: e.g., at least 80% human accuracy overall and at least 50% on the user-state family; per-question, require at least 3/5 annotators to choose the gold answer. If human accuracy is near 44% or at chance on user-state items, the Trajectory Probe lacks construct validity and the central claim is unsupported. If humans score high, the low model accuracy is a genuine memory failure.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim's quantitative core is the Trajectory Probe: 44.4% average accuracy and user-state recall at chance (§5.5). But every stage of the Trajectory Probe is mediated by LLMs: questions and distractor options are LLM-written, the blind and with-history panels are LLM panels, and the 'calibrator ceiling' is another LLM scoring the question with ±15 sessions (§3.5, Table 4). No human annotator ever validates that the 110 calibrated questions are answerable from the transcript, that the gold answer uniquely follows, or that the distractors are plausible-but-wrong rather than arbitrary. If the questions are ambiguous or under-specified, even a model with perfect conversational memory would score near chance. This is especially acute for the 'user-state change' family (n=28 scored decisions per condition, likely only 7 unique questions), which is the basis for the 'near four-option chance' claim, and for 'persona update' (n=8 per condition), whose retrieval-condition collapse to 0.25 is presented as exploratory but is still part of the 'no memory condition fixes it' conclusion. The paper is transparent about LLM mediation in the persona-judge results, but it does not disclose any human ceiling check for the Trajectory Probe; this is a missing support for the strongest abstract claim. Without such a check, the headline accuracy could be an artifact of the question-generation/validation pipeline rather than a property of the evaluated companions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ANCHOR, a synthetic audit framework that separately measures long-horizon persona enactment and trajectory recall in AI companion systems. The study generates 2,008 conversations from 27 authored personas, nine interaction schedules, three memory settings, and four evaluated models. The Identity Probe uses a sealed 102-item questionnaire plus turn-level judgments on four persona axes (role, boundaries, values, style); the Trajectory Probe constructs 110 calibrated counterfactual multiple-choice questions from 35 conversation banks and scores them under long-context, hierarchical-summary, self-managed, and retrieval conditions. The headline results are that trajectory accuracy averages only 44.4%, user-state recall is near four-option chance (0.214–0.250), and no tested context or memory condition consistently resolves these failures. Questionnaire retention varies by model and facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. The paper concludes that current systems do not reliably support long-horizon companion continuity and that audits should keep persona enactment, trajectory recall, evaluator provenance, and deployment context distinct.","tokens_in":15870,"tokens_out":4810,"duration_ms":49584,"significance":"If the central findings hold, this is a valuable and timely negative result for the AI-companion evaluation literature. The paper makes several concrete contributions: a controlled synthetic corpus that separates persona enactment from trajectory recall, a transparent audit pipeline with explicit uncertainty and disaggregated reporting, open score files and manifests for artifact reproduction, and a clear demonstration that judge choice materially changes persona-fidelity conclusions. The refusal to collapse the two probes into a single stability score is methodologically sound. However, the strongest quantitative claims—44.4% trajectory accuracy and chance-level user-state recall—depend on LLM-generated and LLM-calibrated questions with no reported human ceiling check, and the persona-collapse findings rest largely on a single primary judge with limited human validation. These gaps cap the confidence one can place in the abstract's central claim. The framework is significant even as a stress-test methodology; the specific numerical conclusions need additional validation before they can be read as established properties of the evaluated systems.","major_comments":[{"comment":"The Trajectory Probe's headline accuracy (44.4%) and the near-chance user-state recall rest entirely on LLM-mediated question construction and validation: candidate questions are LLM-written, the blind and with-history panels are LLM panels, and the calibrator ceiling is another LLM scoring with a ±15-session window. No human annotator checks that the gold answer uniquely follows from the transcript or that the distractors are plausible-but-wrong rather than arbitrary. If calibrated questions are ambiguous, even a model with perfect conversational memory would score near chance. I ask the authors to add a human ceiling/answerability study on a stratified sample of calibrated questions, or otherwise justify that the LLM-only calibration is sufficient. This is load-bearing for the central negative claim.","section":"§3.5 / §5.5 (Table 4, Table 8)"},{"comment":"The turn-level persona-fidelity results, including the schedule, time, and recovery analyses in §5.4, rely on a single primary judge (Claude Sonnet 4.6) whose outputs are not independently validated on a substantial sample. Appendix B Table 10 shows extreme disagreement: Claude marks GPT-4o-mini's identity-axis held rate at 15.5%, while Gemini-Flash marks it at 100.0%. The paper's own 50-turn human calibration set yields only 64–68% exact four-axis agreement. The authors are transparent that these results are rubric-dependent, but the abstract's claim of 'persona collapse' as an observed failure is weaker than the trajectory claim. I recommend reporting the main turn-level analyses under multiple judges or providing a larger human-validated sample, and softening any language that implies model-level persona-collapse findings are established.","section":"§4 / Appendix B (Table 10)"},{"comment":"The 'user-state change' and 'persona update' families are extremely small: 28 and 8 pooled scored decisions per condition, respectively. The user-state family is the basis for the abstract's claim that recall remains near four-option chance, yet the table reports only pooled decisions, not the number of unique calibrated questions. If this family contains only a handful of unique items (as the n=28 number suggests), the near-chance result is a fragile basis for a general conclusion. Please report unique-question counts per family, per-bank confidence intervals, and avoid generalizing from families with n=8 (the retrieval-condition collapse of persona-update accuracy to 0.25 is explicitly exploratory, but the text still folds it into the 'no memory condition fixes it' summary).","section":"§5.5 / Table 8"}],"minor_comments":[{"comment":"'three memory architecture in the pipeline' should be 'three memory architectures'.","section":"§3.3"},{"comment":"The model is labeled 'GPT-5.4-mini' in several figures but 'GPT-5-mini' in the tables and text. Please standardize.","section":"Figures 3, 4, 6"},{"comment":"The table and figure labels imply unique-question counts for small families, but only pooled decision counts are shown. Add a column for unique calibrated questions per family to support the exploratory caveats.","section":"Table 8 / Figure 8"},{"comment":"The paragraph 'Why explicit attacks can look less damaging' is appropriately framed as a hypothesis. It would be useful to explicitly mark it as untested in the section header or subheading.","section":"§5.4"},{"comment":"Persona Retention is described as unbounded and can exceed 1 or become negative. The text explains this well; consider adding a one-line reminder in Figure 4's caption that values are means of signed projections.","section":"§4, Eq. (1)"}],"recommendation":"major_revision","confidential_remarks":"The central contribution—a transparent, disaggregated audit of long-horizon companion continuity—is strong and worth publishing. My main concern is that the headline trajectory numbers and the persona-collapse narrative both depend on LLM-mediated judgment with no human ceiling validation for the Trajectory Probe and a very small human calibration for the turn-level judge. These are fixable with additional validation and more careful framing; they are not fatal to the framework. If the authors add a human-answerability check on the calibrated trajectory questions and either broaden the judge validation or temper the persona-specific claims, I would support acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. ANCHOR is a real contribution: it separates persona enactment from trajectory recall over 85–130 session synthetic conversations, uses 27 personas and nine schedules, and reports evaluator provenance instead of collapsing everything into one stability score. The PR metric and turn-level judge give complementary evidence, and the paper is careful about nested units and small question families. The central negative result—no evaluated model/configuration reliably preserves both persona and trajectory—is supported by a large corpus, with trajectory accuracy mostly above chance but far from reliable across families and context conditions.\n\nThe new piece is the dual-probe design and the transparent reporting. They distinguish questionnaire retention from turn-level behavior, show large inter-judge disagreement, and refuse to average the probes into a single score. That is worth crediting.\n\nThe main soft spot is the one the stress-test note flags. Every stage of the Trajectory Probe is LLM-mediated: question writing, blind and with-history panels, and the calibrator ceiling. There is no human validation that the 110 questions are answerable from the transcript, that the gold answer is uniquely forced, or that the distractors are plausible rather than arbitrary. That means the 44.4% headline is conditional on the calibration pipeline. If the questions are ambiguous, a model with perfect memory could still score near chance. The calibration step—an LLM answering from the relevant window—reduces that risk, but it is not the same as human judgment. This does not sink the paper, but it caps confidence in the exact numbers and should be acknowledged more directly. The same goes for the tiny user-state-change family (about seven questions) and the persona-update family (two questions); they are labeled exploratory, but they feed the abstract's 'no context condition fixes it' claim.\n\nThe turn-level findings are more sound. The paper explicitly labels schedule, time, and recovery results as primary-judge findings, and the inter-judge disagreement is dramatic (e.g., GPT-4o-mini identity axis at 15.5% with one judge and 100% with another). That is honest and correctly limits what can be claimed.\n\nI largely agree with your read. The protocol is solid enough for peer review, and the central argument holds as a bounded audit result. The missing human ceiling check is a revision-level fix, not a rejection-level one. I'd send it out, and in revision I'd add a human sanity check on a sample of trajectory questions and soften the abstract until it is in.","headline":"ANCHOR is a serious, carefully bounded long-horizon audit; the negative result is probably right, but the all-LLM question pipeline means the headline 44.4% needs a human ceiling check before I'd quote it.","tokens_in":16330,"tokens_out":4387,"would_cite":true,"duration_ms":46311,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Across 2,008 long synthetic conversations, no tested AI companion kept its persona and remembered the shared history: trajectory recall averaged 44.4% and user-state memory was at chance.","keywords":["persona collapse","behavioral drift","AI companions","long-horizon memory","trajectory recall","persona fidelity","LLM judge reliability","counterfactual question evaluation"],"falsifier":"A decisive check is a human-labeling study: take a few hundred turns from the released conversations, have humans score the four persona axes without seeing any model-judge labels, and compare model orderings. If humans do not reproduce the primary judge's ordering—particularly the large identity-axis gap between the two models rated high and the two rated low—then the schedule, time, and recovery results are rubric artifacts. A second, sharper falsifier: any configuration that scores above 90% on all seven trajectory families while maintaining human-agreement-level persona fidelity across all","tokens_in":15391,"feed_emoji":"🎭","tokens_out":7530,"duration_ms":75056,"temperature":0.7,"pith_summary":"People increasingly treat AI companions as ongoing relationships, but a fluent reply does not ensure the system still enacts its assigned role, boundaries, values, and style, or that it remembers the conversation's history. The paper introduces ANCHOR, a controlled synthetic audit that generates 85–130-session conversations across 27 personas, nine interaction schedules, three memory settings, and four evaluated models, measuring two separate failure surfaces: persona collapse/behavioral drift and trajectory recall. The central finding is negative: no evaluated model and configuration reliably preserves either dimension, with trajectory accuracy averaging 44.4% on four-option counterfactual questions and user-state recall hovering at chance under every tested condition. Among the social-pressure schedules, emotional-vulnerability and agreement-seeking conversations produce more boundary yields than explicit adversarial prompts, because ordinary requests for reassurance do not trigger the same refusals. The paper argues that continuity is multidimensional and evaluator-dependent, so audits must report question family, context condition, evaluator provenance, and uncertainty instead of collapsing everything into a single stability score.","feed_headline":"AI companions forget who they are and what happened","feed_subtitle":"A 2,008-conversation audit finds no tested model holds its persona or recalls the shared history.","key_machinery":"The central object is ANCHOR (Assistant-Normalised Character and Historical Outcome Recall), a synthetic audit pipeline. Its Identity Probe measures persona enactment through a sealed checkpoint questionnaire and three-level per-turn judgments on four axes (role identity, boundaries, values, style), scored by three LLM judges with judge choice treated as measurement uncertainty. Its Trajectory Probe measures memory through calibrated counterfactual four-option questions that require distinguishing what actually happened from plausible alternatives. The statistical pivot is Persona Retention, a projection of later questionnaire responses onto the direction from the model's bare-assistant anch","core_discovery":"ANCHOR operationalizes two observable long-horizon failures. 'Persona collapse' is the loss of a deployed role, boundaries, values, or style; 'behavioral drift' is their gradual or recurrent erosion. The Identity Probe combines a sealed 102-item questionnaire, taken at four checkpoints with answers excluded from the dialogue, and turn-level judgments of every assistant turn on the four persona axes; the Persona Retention projection measures how far a later questionnaire response has drifted toward the model's bare-assistant anchor. The Trajectory Probe scores 110 calibrated counterfactual four-option questions in seven families—persona voice, persona protection, persona update, active and ex","pith_inferences":["My inference: if the pattern generalizes, continuity is a property of the whole deployed configuration—persona card, memory pipeline, scheduling, and evaluator—rather than of the base model, so a product team's own tests matter more than any model-level ranking.","My inference: the chance-level user-state recall points to a concrete design requirement—write-audited, replayable state updates rather than model-written summaries—because the paper's self-managed JSON setting was one such attempt and still failed, suggesting the failure is in how state is written and read, not merely in context length.","My inference: a natural extension would randomize user-simulator behavior or insert long time gaps between sessions to test whether drift scales with emotional intensity, conversation length, or recency; the paper's deterministic seed and ledger make this a cheap next experiment.","My inference: the sharp judge disagreement on role and style suggests that audits should prefer observable behavioral criteria—such as boundary-refusal rates on scripted scenarios—over holistic 'sounds like the character' judgments, since the latter are not reproducible across evaluators."],"forward_implications":["Systems configured like the ones tested cannot be assumed to sustain a disclosed companion role over long use; users and developers should treat unstated continuity as unsupported.","Memory architecture alone is not the fix: long-context transcripts, hierarchical summaries, self-managed JSON state, and retrieval-based scoring all leave trajectory recall near 44% overall and user-state recall at chance.","Any single-number 'stability' or 'trust' score for a companion is misleading; the paper's evidence requires disaggregating persona enactment, trajectory recall, evaluator provenance, and deployment context.","Near-chance user-state recall implies a system can silently lose a consequential update, such as a changed treatment goal or a changed user life circumstance, which is a concrete risk for health-adjacent companions.","Explicit adversarial re-role attempts can look less damaging than ordinary emotional or agreement-seeking conversation, because recognizable attacks trigger visible refusals while routine requests create more opportunities for boundary yield."],"fun_headline_variants":["AI companions drift: no model holds persona or history","2,008 chats show AI companions forget their role","Persona collapse and drift: no AI companion passes audit","AI companions fail long-horizon persona and memory tests","No AI companion keeps its persona or your shared history"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the LLM judges' turn-level labels validly reflect whether a persona axis is held; the paper's own validation shows 64–68% exact four-axis agreement on 50 human-labeled turns and sharp judge disagreement on some models, so if the primary judge is miscalibrated the schedule, time, and recovery findings are rubric artifacts rather than model behavior.","fun_headline_variants_meta":{"raw":{"variants":["AI companions drift: no model holds persona or history","2,008 chats show AI companions forget their role","Persona collapse and drift: no AI companion passes audit","AI companions fail long-horizon persona and memory tests","No AI companion keeps its persona or your shared history"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000537,"raw_usage":{"total_tokens":2430,"prompt_tokens":776,"completion_tokens":1654,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":1576}},"tokens_in":520,"tokens_out":1654,"duration_ms":11918,"temperature":1.0,"reasoning_tokens":1576,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T00:17:22.347073+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check is a human-labeling study: take a few hundred turns from the released conversations, have humans score the four persona axes without seeing any model-judge labels, and compare model orderings. If humans do not reproduce the primary judge's ordering—particularly the large identity-axis gap between the two models rated high and the two rated low—then the schedule, time, and recovery results are rubric artifacts. A second, sharper falsifier: any configuration that scores above 90% on all seven trajectory families while maintaining human-agreement-level persona fidelity across all","supporting_citations":[],"review_version":1}