Pith. sign in

REVIEW 4 major objections 5 minor 49 references

LLMs have no representation of the person beyond the prompt; marking six categories of ignorance makes them ask before they advise.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 02:38 UTC pith:J3GERGFU

load-bearing objection The memory-blind-spot result is solid and worth citing; the six-dimension schema's specific effect is not established, and the paper's own ablation undercuts its central claim. the 4 major comments →

arxiv 2607.14250 v1 pith:J3GERGFU submitted 2026-07-15 cs.CL

The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt

classification cs.CL
keywords severance problempersonal AI assistantsstructured ignoranceunknown-unknownssycophancyhallucinationclarifying questionsuser context
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that personal AI assistants fail not because they lack facts about the user, but because they lack a representation of what kinds of facts about a person might be missing. This is the Severance Problem. The proposed fix, the Severance Schema, is a prompt block that names six dimensions of a person's life—body, time, stakes, history, roles, inner life—and marks each as known or [unknown]. Across five model families, the schema roughly doubled how often models recognized missing person-context, cut harmful advice and sycophancy by more than half on open-weight models, and reduced the hallucination that memory alone introduced.

Core claim

The paper's central claim is that a model cannot reason about what it does not know about a person unless it has an explicit inventory of the categories in which such knowledge might live. The Severance Schema provides that inventory: six dimensions of person-context, each either filled from memory or marked [unknown], placed in the system prompt as an 'outie mental model.' Empirically, this structured ignorance converts unknown-unknowns into known-unknowns: models with the schema ask clarifying questions before answering, rather than extrapolating from incomplete context. Across five model families, unknown-awareness roughly doubles; harmful advice and sycophancy fall by more than half on o

What carries the argument

The Severance Schema: a prompt-level 'outie mental model' that lists six dimensions of person-context—physicality, temporality, consequences, continuity, multiplicity, and interiority—with each slot marked known or [unknown]. It works by making the model's knowledge boundary explicit, converting unknown-unknowns into known-unknowns so the model can seek missing information. The paper isolates its effect with a four-condition comparison (no-schema, schema, memory, schema+memory) and an awareness plane that separates information about the person from awareness of the knowledge boundary.

Load-bearing premise

The result depends on the six dimensions of the Severance Schema being the right, sufficiently complete categories of person-context, and on the evaluation's ground truth being independent of the intervention—the paper's synthetic profiles and per-scenario claims are built from exactly those six categories, and the paper's limitations note that domain-specific and long-term effects remain untested.

What would settle it

Build an advisory benchmark whose decision-flip facts are deliberately chosen to lie outside the six schema dimensions (for example, legal jurisdiction, cultural norms, or a seventh unlisted category). If the schema no longer doubles unknown-awareness or fails to surface those missing facts, the effect is benchmark-conformant rather than a general awareness of missing person-context.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • If the schema is correct, a short structured-ignorance block in the system prompt is a cheap, model-agnostic safety intervention for personal AI assistants.
  • Memory-based personalization has a structural blind spot: more user facts suppress question-asking and raise hallucination, so memory should be paired with an explicit representation of what remains unknown.
  • The schema's benefit persists even when all six dimensions are populated, so marking unknowns remains valuable as user data accumulates.
  • In multi-turn interaction, the schema converts clarifying questions into measurable gains in calibration and usefulness, suggesting that question-asking functions as a retrieval mechanism, not just a conversational nicety.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the mechanism generalizes, the same scaffold could be adapted to domain-specific advisory settings (medical, legal, financial), but the six dimensions would need to be re-derived from that domain's own unknowns; the paper's limitations leave this untested.
  • Because the benchmark's decision-flip claims are drawn from the same six categories the schema names, an independent test using missing facts outside those categories would determine whether the gains reflect a general awareness capability or alignment to the benchmark's ontology.
  • A natural further test is a human-subject study of whether schema-driven clarification improves users' actual decisions and trust, not just judge-scored response quality; the paper itself flags long-term effects as an open question.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper identifies the 'Severance Problem': LLM-based personal assistants lack an explicit representation of what they do not know about the user beyond the prompt. The authors propose the 'Severance Schema', a prompt-level scaffold that lists six dimensions of person-context (physicality, temporality, consequences, continuity, multiplicity, interiority) and marks each as known or [unknown]. Across five model families and a synthetic benchmark of 10 profiles and 30 advice-seeking scenarios, they report that adding the schema roughly doubles unknown-awareness, halves harmful advice and sycophancy on open-weight models, and reduces hallucination rates in the Memory condition. They also report ablations (Structure Only vs. Semantic Only), a multi-turn extension, and human validation of the LLM judge.

Significance. If the central claim holds, the Severance Schema is a simple, model-agnostic, prompt-level intervention with substantial practical value for personal AI safety: it reduces sycophancy, harmful advice, and hallucination without retraining. The paper's strengths include the evaluation across five model families, paired cluster-bootstrap confidence intervals, a validation judge, a human agreement study, and a clear conceptual framing of first-order vs. second-order unawareness. The multi-turn result — that schema-asking converts into downstream usefulness gains — is a genuinely interesting and falsifiable finding. However, the empirical support for the specific claim that the six dimensions are the active ingredient is weakened by the benchmark–intervention circularity and by the ablation results, which show that neutral slots or semantic lists alone recover most of the unknown-awareness gain.

major comments (4)
  1. [§3.1, Table 5; App. E] The ablation does not support the claim that 'neither the [unknown] slot format alone nor simply naming the person-context categories recovers the full effect.' Table 5 shows Structure Only gives unknown-awareness 5.64 vs. 5.84 for the full schema, and Semantic Only gives 5.79 vs. 5.84. Both ablations recover roughly 95–99% of the gain over No Schema (2.56). Without a generic-uncertainty control (e.g., a non-person 'what don't I know about this decision?' prompt) or a statistical comparison among ablations, the data support the weaker conclusion that any explicit unknown-slot structure suffices. The §3.1 text overstates the ablation's discriminating power.
  2. [§3.1, App. B/C/D; Fig. 7; Judge Prompt] The evaluation's ground truth is not independent of the intervention. The 10 profiles are organized along exactly the six schema dimensions (App. B), the per-scenario good-answer/decision-flip claims are grouped and color-coded by those same six dimensions (App. C, Fig. 7), and the judge prompt explicitly instructs the evaluator to reward responses that surface 'the SPECIFIC decision-flip claims listed above' and counts mentions_unknowns only if they match those claims. The schema prompt names exactly those dimensions. Consequently, the unknown-awareness gain is at least partly a benchmark-alignment artifact: the model is rewarded for generating the categories the benchmark was built to expect. The paper needs an independent benchmark — ideally with claims authored by human annotators who are not given the six-dimension taxonomy — or at minimum a control condition that uses a different,
  3. [§3.3, App. H] The multi-turn usefulness claim rests on a retrieval step that extracts clarifying questions and fills the corresponding schema slots from the ground-truth profile. This procedure is reasonable, but the retrieval model (Claude Haiku) uses the full profile, which is also authored in the six-dimension ontology. The second-turn improvement therefore partly reflects the alignment between the schema's question-asking and the benchmark's pre-structured retrieval keys. The paper should report how many retrieved answers actually correspond to decision-flip vs. good-answer claims, and how often the retrieval injects information the model did not ask for, to establish that the clarification itself drives the gain.
  4. [§3.2, Table 8; App. G] The claim that the schema's advantage persists at 100% information is used to argue that 'the schema's dimensions remain beneficial even after they have been populated.' But at 100% fill, both conditions receive identical information; the only remaining difference is the prompt format (bullets vs. filled schema slots). The unknown-awareness gap at 100% (+1.82) could simply reflect the schema's instruction to 'name what you don't know' rather than any deeper structural benefit. The paper should include a condition that receives the same information with a non-schema prompt that also explicitly encourages flagging residual unknowns, to separate the checklist effect from the generic instruction effect.
minor comments (5)
  1. [§2, p.4] The phrase 'these six dimensions are a general-purpose starting set; they can be expanded or refined' suggests the taxonomy is ad hoc. Please provide a principled criterion for selecting or extending dimensions; otherwise the reader cannot judge completeness or generalizability.
  2. [§3.1, Table 1] Several individual-model confidence intervals for harm reduction include zero (e.g., Gemma, Qwen3 in the Schema−No Schema contrast; Sonnet, Gemma, Qwen3 in the Schema+Mem−Mem contrast). The text says 'reducing harmful advice and sycophancy by more than half on the open-weight models' — clarify that this is not uniformly significant at the individual-model level, and rely on the pooled/bootstrap evidence for the general claim.
  3. [App. D, Judge Prompt, calibration rubric] The calibration rubric says a score of 5 requires 'the unknowns it flags or asks about are the SPECIFIC decision-flip claims listed above.' This embeds the benchmark's claim labels into the quality metric. Even with condition blinding, this makes calibration partly a claim-recall metric rather than an independent measure of personalized calibration.
  4. [App. F, Table 6] The Sonnet-4 validation judge yields much higher hallucination rates across all conditions (e.g., 9.0% for No Schema vs. GPT-5.2's 0.0% on Claude Sonnet 4). The paper acknowledges differences but does not discuss whether the safety metric conclusions are judge-dependent. A sentence explaining the magnitude and direction of this discrepancy would help readers assess robustness.
  5. [Various] Minor typographical and formatting issues: 'uncertainty' missing in §1 ('over known-unknowns'), missing space after 'as the' on p.2, and 'Mieleszczenko-Kowszewicz et al.' is overlong inline; consider shortening or using 'et al.' consistently.

Circularity Check

1 steps flagged

Unknown-awareness is measured against a benchmark whose profiles and claim labels are built from the same six-dimension ontology the schema injects; the paper's own ablation shows generic structure/semantics recover most of the gain, so the central claim is substantially benchmark-aligned.

specific steps
  1. self definitional [§3 Data / App. B / App. C Fig. 7 / App. D Judge Prompt]
    "Each scenario is manually annotated with two sets of claims... The claims are visible only to the LLM judge and allow it to ground the unknown-unknowns with labels; without these annotations, the judge would have to infer which missing factors mattered... Figure 7: rows are the 52 claim labels, grouped and color-coded by the six dimensions (physicality, temporality, consequences, continuity, multiplicity, interiority)."

    The Severance Schema prompt (App. A) instructs the model to maintain an OUTIE MENTAL MODEL with exactly these six dimensions (BODY, TIMELINE, STAKES, HISTORY, ROLES, INNER LIFE). The benchmark's ground truth is not independent: App. B states each synthetic profile is organized along the six person-context high-level dimensions of §2, and App. C groups every good-answer and decision-flip claim by those same six dimensions. The judge is given those claims and is told that a 5/5 calibration response must flag 'the SPECIFIC decision-flip claims listed above,' and mentions_unknowns is counted only against the scenario's decision-flip and good-answer claims. Thus a model that echoes the schema's category names is rewarded by construction. Blinding the judge to condition prevents prompt-condition

full rationale

The paper's central empirical claim is that the six-dimension Severance Schema roughly doubles unknown-awareness and halves harm/sycophancy. The unknown-awareness result is not independent of the intervention: the benchmark's synthetic profiles (App. B) and every scenario's good-answer/decision-flip claim labels (App. C, Fig. 7) are organized along exactly the six schema dimensions, and the LLM judge (App. D) is given those claims and instructed to reward responses that surface 'the SPECIFIC decision-flip claims listed above.' The schema prompt injects the same six dimensions. So the model is being scored on its ability to name the categories the benchmark was built to expect; blinding the judge to condition does not remove this ontology-level alignment. The safety metrics (harmful advice, sycophancy, hallucination) are less tightly coupled to the category names and are cross-validated by a second judge and a human pairwise study, so those results retain independent content. The paper's own ablation (App. E, Table 5) further undercuts the claim that the combination of structure and semantics is necessary: Structure Only (5.64) and Semantic Only (5.79) recover nearly all of the full schema's unknown-awareness (5.84), while No Schema is 2.56. No self-citation chain is load-bearing. On balance, the central unknown-awareness claim is partially circular (the measurement ontology is the intervention ontology), but the core safety results are not; score 6.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 1 invented entities

The central claim rests on a hand-chosen taxonomy (six dimensions), a benchmark constructed from that taxonomy, and LLM-judge labels. The schema is the intervention and the test is partly written in its language, which is the main circularity burden. No numeric constants are fitted to data.

free parameters (3)
  • Six Severance Schema dimensions (physicality, temporality, consequences, continuity, multiplicity, interiority)
    Hand-chosen taxonomy in §2; the central claim depends on this ontology, and no independent argument or data selects these six over others.
  • Per-scenario good-answer/decision-flip claim sets
    Manually written labels (App. C, Fig. 7) define what counts as awareness and safety; they are the benchmark's ground truth and directly shape the scores.
  • Synthetic person profiles (10)
    Profiles in App. B are synthetic and organized along the same six dimensions; all model responses are scored against these constructed profiles.
axioms (3)
  • domain assumption LLM-as-judge ratings (GPT-5.2, with Sonnet 4 validation) correspond to true response quality and safety.
    Used for all metrics; human validation in Tab. 4 covers only calibration and mentions-unknowns, not hallucination, harm, sycophancy, or usefulness.
  • ad hoc to paper The six-dimension taxonomy sufficiently covers person-context.
    §2 introduces the dimensions; App. B/C construct profiles and claims using them, so completeness is assumed rather than externally demonstrated.
  • domain assumption Meta-ignorance taxonomy from Smithson applies to language-model behavior.
    §2 imports the known-unknown/unknown-unknown distinction as an explanation of model behavior without direct evidence of internal representation.
invented entities (1)
  • Severance Schema (structured-ignorance prompt scaffold) no independent evidence
    purpose: Provide an explicit inventory of six person-context dimensions, each marked known or [unknown], so the model asks about what it lacks rather than extrapolating.
    The benchmark used to validate the schema is built from the same six dimensions (App. B, C; Fig. 7), so there is no external, schema-free benchmark that could falsify the taxonomy.

pith-pipeline@v1.3.0-alltime-deepseek · 22847 in / 13601 out tokens · 137271 ms · 2026-08-02T02:38:23.630358+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt." pith.science (2026). https://pith.science/paper/J3GERGFU

@misc{pith2026260714250,
  author       = {Pith},
  title        = {Pith review of: The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/J3GERGFU}},
  note         = {Machine review of arXiv:2607.14250}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Personal AI assistants have attracted significant interest for their potential to enhance everyday life by automating routine tasks, supporting consequential decisions, and assisting with everyday personal matters. Yet despite rapid recent technical advances, these assistants continue to exhibit undesirable behaviors, such as sycophancy, overconfidence, and hallucination. We argue that these failures stem from a fundamental limitation: language models lack an explicit representation of the person beyond the context they are given, which we term as the \textbf{Severance Problem}. Even with rich personal context and strong commonsense reasoning capabilities from the backbone model, current AI assistants fail to represent what remains unknown about the user. We propose a simple solution: incorporating structured ignorance into the language model context via the \textbf{Severance Schema}, which explicitly outlines dimensions along which the model lacks knowledge about the user, including physicality, temporality, consequences, continuity, multiplicity, and interiority. Empirically, across five model families, with the Severance Schema, the assistant consistently reduces sycophancy, harmful advice, and hallucination. Notably, models with the schema ask clarifying questions when information about the user is missing, rather than confidently extrapolating from incomplete user information.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

49 extracted references · 10 linked inside Pith

  1. [1]

    CLAMBER: A benchmark of identifying and clarifying ambiguous information needs in large language models

    Tong Zhang, Peixin Qin, Yang Deng, Chen Huang, Wenqiang Lei, Junhong Liu, Dingnan Jin, Hongru Liang, and Tat-Seng Chua. CLAMBER: A benchmark of identifying and clarifying ambiguous information needs in large language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024. doi: 10.18653/v1/2024.acl-lon...

  2. [2]

    CLAM: Selective clarification for ambiguous questions with generative language models.arXiv preprint arXiv:2212.07769, 2022

    Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. CLAM: Selective clarification for ambiguous questions with generative language models.arXiv preprint arXiv:2212.07769, 2022

  3. [3]

    Patil, Ion Stoica, and Joseph E

    Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. MemGPT: Towards LLMs as operating systems.arXiv preprint arXiv:2310.08560, 2024

  4. [4]

    Mem0: Building production-ready ai agents with scalable long-term memory.arXiv preprint arXiv:2504.19413, 2025

    Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready ai agents with scalable long-term memory.arXiv preprint arXiv:2504.19413, 2025

  5. [5]

    Memory and new controls for ChatGPT

    OpenAI. Memory and new controls for ChatGPT. OpenAI blog, 2024. URL https: //openai.com/index/memory-and-new-controls-for-chatgpt/

  6. [6]

    Smithson.Ignorance and Uncertainty: Emerging Paradigms

    M. Smithson.Ignorance and Uncertainty: Emerging Paradigms. Springer-Verlag, New York, 1989. 11

  7. [7]

    System card: Claude opus 4 and claude sonnet 4

    Anthropic. System card: Claude opus 4 and claude sonnet 4. Tech- nical report, Anthropic, May 2025. URL https://www.cdn.anthropic.com/ 4263b940cabb546aa0e3283f35b686f4f3b2ff47.pdf

  8. [8]

    Llama 3.3 70B Instruct model card

    Meta. Llama 3.3 70B Instruct model card. Technical report, Meta, December 2024. URL https://www.llama.com/docs/model-cards-and-prompt-formats/llama3_3/

  9. [9]

    Deepseek-v3 technical report.arXiv preprint arXiv:2412.19437, 2024

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report.arXiv preprint arXiv:2412.19437, 2024

  10. [10]

    Gemma 4 model card

    Google DeepMind. Gemma 4 model card. Technical report, Google, April 2026. URL https://ai.google.dev/gemma/docs/core/model_card_4

  11. [11]

    Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025

  12. [12]

    Update to GPT-5 system card: GPT-5.2

    OpenAI. Update to GPT-5 system card: GPT-5.2. Technical report, OpenAI, December

  13. [13]

    Language models (mostly) know what they know.arXiv preprint arXiv:2207.05221, 2022

    Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. Language models (mostly) know what they know.arXiv preprint arXiv:2207.05221, 2022

  14. [14]

    Teaching models to express their uncertainty in words.Transactions on Machine Learning Research, 2022

    Stephanie Lin, Jacob Hilton, and Owain Evans. Teaching models to express their uncertainty in words.Transactions on Machine Learning Research, 2022. ISSN 2835-8856. URL https://openreview.net/forum?id=8s8K2UZGTZ

  15. [15]

    Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024

    Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024

  16. [16]

    Do large language models know what they don’t know? InFindings of the Association for Computational Linguistics, 2023

    Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. Do large language models know what they don’t know? InFindings of the Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.findings-acl.551. URLhttps: //aclanthology.org/2023.findings-acl.551/

  17. [17]

    Knowledge of knowledge: Exploring known-unknowns uncertainty with large language models

    Alfonso Amayuelas, Kyle Wong, Liangming Pan, Wenhu Chen, and William Yang Wang. Knowledge of knowledge: Exploring known-unknowns uncertainty with large language models. InFindings of the Association for Computational Linguistics, pages 6416–6432, 2024. doi: 10.18653/v1/2024.findings-acl.383. URL https://aclanthology.org/2024.findings-acl. 383/

  18. [18]

    Do llms estimate uncer- tainty well in instruction-following? InInternational Conference on Learning Representations (ICLR), volume 2025, pages 95951–95974, 2025

    Juyeon Heo, Miao Xiong, Christina Heinze-Deml, and Jaya Narain. Do llms estimate uncer- tainty well in instruction-following? InInternational Conference on Learning Representations (ICLR), volume 2025, pages 95951–95974, 2025

  19. [19]

    Evaluating large language models in theory of mind tasks.Proceedings of the National Academy of Sciences, 121(45):e2405460121, 2024

    Michal Kosinski. Evaluating large language models in theory of mind tasks.Proceedings of the National Academy of Sciences, 121(45):e2405460121, 2024. doi: 10.1073/pnas.2405460121. URLhttps://www.pnas.org/doi/abs/10.1073/pnas.2405460121

  20. [20]

    Large language models fail on trivial alterations to theory-of-mind tasks

    Tomer Ullman. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399, 2023. 12

  21. [21]

    Neural theory-of-mind? on the limits of social intelligence in large LMs

    Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi. Neural theory-of-mind? on the limits of social intelligence in large LMs. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022. doi: 10.18653/v1/2022.emnlp-main.248. URL https://aclanthology.org/2022.emnlp-main.248/

  22. [22]

    Me, myself, and AI: The situational awareness dataset (SAD) for LLMs

    Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hariharan, Mikita Balesni, Jérémy Scheurer, Marius Hobbhahn, Alexander Meinke, and Owain Evans. Me, myself, and AI: The situational awareness dataset (SAD) for LLMs. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024

  23. [23]

    Tell me about yourself: LLMs are aware of their learned behaviors

    Jan Betley, Xuchan Bao, Martín Soto, Anna Sztyber-Betley, James Chua, and Owain Evans. Tell me about yourself: LLMs are aware of their learned behaviors. InThe Thirteenth International Conference on Learning Representations, 2025

  24. [24]

    On targeted manipulation and deception when optimizing LLMs for user feedback

    Marcus Williams, Micah Carroll, Adhyyan Narang, Constantin Weisser, Brendan Murphy, and Anca Dragan. On targeted manipulation and deception when optimizing LLMs for user feedback. InThe Thirteenth International Conference on Learning Representations, 2025

  25. [25]

    Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin DURMUS, Zac Hatfield-Dodds, Scott R Johnston, Shauna M Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. Towards understanding sycophancy in language models. InThe Twelfth Internatio...

  26. [26]

    Wiktoria Mieleszczenko-Kowszewicz, Dawid Płudowski, Filip Kołodziejczyk, Jakub Świstak, Julian Sienkiewicz, and Przemysław Biecek. The dark patterns of personalized persuasion in large language models: Exposing persuasive linguistic features for big five personality traits in LLMs responses.arXiv preprint arXiv:2411.06008, 2024

  27. [27]

    P. S. Park, S. Goldstein, A. O’Gara, M. Chen, and D. Hendrycks. AI deception: A survey of examples, risks, and potential solutions.Patterns, 5(5), 2024

  28. [28]

    Gui, Tianyi Peng, Daniel J

    Olivier Toubia, George Z. Gui, Tianyi Peng, Daniel J. Merlau, Ang Li, and Haozhe Chen. Database report: Twin-2K-500: A data set for building digital twins of over 2,000 people based on their answers to over 500 questions.Marketing Science, 44(6):1446–1455, 2025. doi: 10.1287/mksc.2025.0262. URLhttps://doi.org/10.1287/mksc.2025.0262

  29. [29]

    O’Brien, Carrie J

    Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. InIn the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23), 2023

  30. [30]

    Generative agent simulations of 1,000 people.arXiv preprint arXiv:2411.10109, 52, 2024

    Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Mered- ith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. Generative agent simulations of 1,000 people.arXiv preprint arXiv:2411.10109, 52, 2024

  31. [31]

    How far are LLMs from being our digital twins? a benchmark for persona-based behavior chain simulation

    Rui Li, Heming Xia, Xinfeng Yuan, Qingxiu Dong, Lei Sha, Wenjie Li, and Zhifang Sui. How far are LLMs from being our digital twins? a benchmark for persona-based behavior chain simulation. InFindings of the Association for Computational Linguistics: ACL 2025, 2025. doi: 10.18653/v1/2025.findings-acl.813. URL https://aclanthology.org/2025.findings-acl. 813/

  32. [32]

    KnowU-Bench: Towards interactive, proactive, and personalized mobile agent evaluation.arXiv preprint arXiv:2604.08455, 2026

    Tongbo Chen, Zhengxi Lu, Zhan Xu, Guocheng Shao, Shaohan Zhao, Fei Tang, Yong Du, Kaitao Song, Yizhou Liu, Yuchen Yan, et al. KnowU-Bench: Towards interactive, proactive, and personalized mobile agent evaluation.arXiv preprint arXiv:2604.08455, 2026. 13

  33. [33]

    LaMP: When large language models meet personalization

    Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. LaMP: When large language models meet personalization. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024. doi: 10.18653/v1/2024.acl-long.399. URLhttps://aclanthology.org/2024.acl-long.399/

  34. [34]

    PersonalLLM: Tailoring LLMs to individual preferences

    Thomas P Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li, and Hongseok Namkoong. PersonalLLM: Tailoring LLMs to individual preferences. InThe Thirteenth International Conference on Learning Representations, 2025

  35. [35]

    Whose opinions do language models reflect? InInternational conference on machine learning, 2023

    Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? InInternational conference on machine learning, 2023

  36. [36]

    Uncertainty of thoughts: Uncertainty-aware planning enhances information seeking in LLMs

    Zhiyuan Hu, Chumin Liu, Xidong Feng, Yilun Zhao, See-Kiong Ng, Anh Tuan Luu, Junxian He, Pang Wei Koh, and Bryan Hooi. Uncertainty of thoughts: Uncertainty-aware planning enhances information seeking in LLMs. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  37. [37]

    Bradley Knox, and Eunsol Choi

    Michael JQ Zhang, W. Bradley Knox, and Eunsol Choi. Modeling future conversation turns to teach LLMs to ask clarifying questions. InThe Thirteenth International Conference on Learning Representations, 2025

  38. [38]

    Li, Alex Tamkin, Noah Goodman, and Jacob Andreas

    Belinda Z. Li, Alex Tamkin, Noah Goodman, and Jacob Andreas. Eliciting human pref- erences with language models. InThe Thirteenth International Conference on Learning Representations, 2025

  39. [39]

    STar- GATE: Teaching language models to ask clarifying questions

    Chinmaya Andukuri, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah Goodman. STar- GATE: Teaching language models to ask clarifying questions. InFirst Conference on Language Modeling (COLM), 2024

  40. [40]

    Artificial intelligence, values, and alignment.Minds and machines, 30(3): 411–437, 2020

    Iason Gabriel. Artificial intelligence, values, and alignment.Minds and machines, 30(3): 411–437, 2020

  41. [41]

    LLM alignment should go beyond harmlessness-helpfulness and incorporate human agency.Cognitive Computation, 18(1):26, 2026

    Usman Naseem, Tanmoy Chakraborty, Kai-Wei Chang, Mark Dras, Preslav Nakov, Nanyun Peng, and Soujanya Poria. LLM alignment should go beyond harmlessness-helpfulness and incorporate human agency.Cognitive Computation, 18(1):26, 2026

  42. [42]

    The benefits, risks and bounds of personalizing the alignment of large language models to individuals.Nature Machine Intelligence, 6(4):383–392, 2024

    Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A Hale. The benefits, risks and bounds of personalizing the alignment of large language models to individuals.Nature Machine Intelligence, 6(4):383–392, 2024

  43. [43]

    Discovering language model behaviors with model-written evaluations

    Ethan Perez, Sam Ringer, Kamile Lukosiute, et al. Discovering language model behaviors with model-written evaluations. InFindings of the Association for Computational Linguistics: ACL 2023, 2023. doi: 10.18653/v1/2023.findings-acl.847. URLhttps://aclanthology.org/ 2023.findings-acl.847/

  44. [44]

    Syco- phantic ai decreases prosocial intentions and promotes dependence.Science, 391(6792): eaec8352, 2026

    Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, and Dan Jurafsky. Syco- phantic ai decreases prosocial intentions and promotes dependence.Science, 391(6792): eaec8352, 2026

  45. [45]

    Your agent, their asset: A real-world safety analysis of openclaw

    Zijun Wang, Haoqin Tu, Letian Zhang, Hardy Chen, Juncheng Wu, Xiangyan Liu, Zhenlong Yuan, Tianyu Pang, Michael Qizhe Shieh, Fengze Liu, Zeyu Zheng, Huaxiu Yao, Yuyin Zhou, and Cihang Xie. Your agent, their asset: A real-world safety analysis of openclaw. In2nd Workshop on Compositional Learning: Safety, Interpretability, and Agents, 2026. 14

  46. [46]

    Openagentsafety: A comprehensive framework for evaluating real-world AI agent safety

    Sanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang, Nouha Dziri, Graham Neubig, and Maarten Sap. Openagentsafety: A comprehensive framework for evaluating real-world AI agent safety. InThe Fourteenth International Conference on Learning Representations, 2026

  47. [47]

    Quantifying language models’ sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting

    Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language models’ sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting. InThe Twelfth International Conference on Learning Representations, 2024

  48. [48]

    I want to start training for a marathon. How should I begin?

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-judge with MT-bench and chatbot arena. InThirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. 15 A Prompts The f...

  49. [2025]

    URL https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944f8d/ oai_5_2_system-card.pdf