Pith. sign in

REVIEW 3 major objections 4 minor 48 references

Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Personality-conditioned LLM agents shift after life events, but their evolution tracks the mean of human personality dynamics, not its shape.

desk verdict A careful, reusable benchmark showing LLM personas respond to life events but with weak directional fidelity and a universal retirement reversal; the quantitative heterogeneity-collapse claim rests on a human-comparison mismatch that the authors should fix. read the letter →

arxiv 2608.06485 v1 pith:MQPMOWDW submitted 2026-08-06 cs.CL cs.AIcs.SI

classification cs.CLcs.AIcs.SI
keywords personality-conditionedLLMagentspersonalityevolutionBigFiveBFI-44majorlifeeventsBFI-Adaptpsychometricevaluationbenchmarking
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether LLM agents given stable personas, called PC-Agents, evolve their Big Five personality profiles after major life events the way people do. The authors ran 100 demographically varied personas through a baseline BFI-44 questionnaire, a first-person reflection on one of 11 life events, and a post-event BFI-44 questionnaire, across 11 main models plus three open-weight extensions. They find that event-induced trait shifts are measurable and repeatable, but weakly tied to the specific event, usually too small in magnitude, insensitive to gender and cultural-region prompts, and much more uniform across personas than human samples. The paper introduces BFI-Adapt, a benchmark that scores whether event-induced changes go in the directions human longitudinal studies document, and uses it to rank models. The central conclusion is that current PC-Agents simulate the mean of human personality dynamics, but not its shape.

What carries the argument

The load-bearing instrument is the paired BFI-44 protocol: a persona answers all 44 items before and after a first-person life-event reflection, giving a per-trait change $\Delta$. A noise boundary $\varepsilon=0.1$ classifies each persona's movement as up, down, or neutral, and directional match is assessed against a prior matrix of expected human changes for 27 of 55 event-trait pairs, taken primarily from Specht (2017). At the item level, linearly weighted Cohen's $\kappa$ measures rating stability between the two administrations, and DCR measures whether the items that change move mostly in one direction. BFI-Adapt combines these per pair, $$\mathrm{BFI\text{-}Adapt} = \frac{1}{|C|}\sum_{(e,t)\in C}\max(0,\kappa_{e,t})\,S_{e,t}\,\mathbb{1}^{\mathrm{dir}}_{e,t}$$ with $S_{e,t}=2\,\mathrm{DCR}_{e,t}-1$, rewarding reliability, one-sided systematic movement, and agreement with the human direction. The human reference envelope, standardized change $|d|\in[0.05,0.20]$ and within-trait SD $\sigma\in[0.5,0.8]$, is converted through $\sigma\approx0.7$ into a raw-change band $\Delta\in[0.035,0.14]$ that anchors the magnitude and dispersion diagnoses.

What would settle it

Recompute the event-conditional change-score standard deviation from raw longitudinal human data for the same 11 events; if the empirical values cluster near the LLM median of 0.19 rather than $\sigma\in[0.5,0.8]$, the heterogeneity-collapse finding would be an artifact of the benchmark. Alternatively, if any current LLM, under a multi-session protocol, shows the documented post-retirement Conscientiousness decline or raises its per-cell change SD above about 0.4, the central claim that PC-Agents only simulate the mean would be directly weakened.

Watch

Extended reading notes

Core claim

The paper reports four empirical patterns. First, movement is indiscriminate: agents move at similar rates for event-trait pairs with and without documented human change directions, with within-model median differences below 0.05 and nine of eleven models within 10 percentage points on high-movement mass. Second, direction and magnitude are miscalibrated: when personas move, only 11.0%-16.4% of responses fall inside the human effect-size band, the median absolute change lies below that band, and the documented post-retirement decline in Conscientiousness is reversed by every model (median match 11.5%), while social events such as marriage, divorce, and childbirth stay near chance. Third, demographic shape is flat: no gender or world-region moderator survives multiple-testing correction, and the median across-strata standard deviation of trait change is 0.044 BFI units. Fourth, individual shape is compressed: the median within-cell standard deviation of change is 0.19 versus a human within-trait SD envelope of 0.5-0.8, a three- to four-fold collapse. Validation checks show the shifts exceed no-event retest noise, keep their event-trait structure under independent paraphrases, and partly persist after unrelated dialogue, while convergence with scenario-based decisions is limited and model-dependent.

Load-bearing premise

The load-bearing premise is that the human reference envelope, a within-trait SD of $\sigma\in[0.5,0.8]$ BFI units and standardized change $|d|\in[0.05,0.20]$ from meta-analytic studies, is the right benchmark for these 11 events and for change-score dispersion; the paper itself calls these ranges coarse.

Editorial extensions

If this is right

  • Lifelong agents that aim to stay psychologically plausible cannot be validated by static persona fidelity alone; trajectory-level criteria are needed to catch agents that stay in character yet evolve implausibly.
  • Any current PC-Agent used in social simulation or role-play involving retirees will systematically invert the human pattern of post-retirement Conscientiousness decline, so simulations of aging populations should treat that output as uncalibrated.
  • Model choice has a measurable effect on trajectory quality: BFI-Adapt spans 0.071-0.348 among the 11 API models, and models fail in different ways, some under-shifting and some reversing or overshooting expected changes.
  • Because event-conditioned shifts exceed no-event retest noise and survive paraphrases and intervening dialogue, the measured trajectories are stable response patterns of current models rather than one-off prompt artifacts.
  • Prompting gender and cultural region does not reproduce human demographic moderation in event-driven trait change; reproducing that structure would require additional mechanisms.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A caveat the paper itself raises: if the human envelope (within-trait SD of 0.5-0.8 and absolute standardized change of 0.05-0.20) is not the right benchmark for these 11 events, the magnitude and heterogeneity findings would be overstated; the retirement reversal and near-chance social-event directions are less dependent on that envelope and are the firmer core.
  • A testable extension is to determine whether the compressed dispersion is architectural or an artifact of the single-reflection protocol; a multi-session design with richer idiosyncratic life histories could reveal whether persona-level spread grows toward the human envelope.
  • The low and model-dependent correlation between BFI trait changes and scenario-based decisions suggests that shape might be more visible in concrete choices than in self-report inventories; a behavior-anchored version of the benchmark could rank models differently.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper investigates whether LLM agents conditioned on synthetic personas show personality change after major life events, using paired BFI-44 administrations before and after event exposure. The design covers 11 life events, 100 controlled personas (2 genders × 5 cultural regions × 10 archetypes), and 11 base models for the main four-axis analysis, with three open-weight models added to a 14-model leaderboard. Four research questions examine the existence of change above retest noise (RQ1), directional and magnitude agreement with human longitudinal priors (RQ2), demographic moderation (RQ3), and persona-level dispersion of change (RQ4). The paper introduces BFI-Adapt, a composite combining item-level reliability, within-pair directional consistency, and agreement with the Specht (2017) prior matrix, and validates the measurement pipeline through no-event retests, independent paraphrases, scenario-decision convergence, and short-range retention. The central finding is that current PC-Agents move after life events but in ways that are weakly event-specific, poorly calibrated in magnitude, demographically flat, and compressed in dispersion, leading the authors to conclude that agents simulate the mean but not the shape of human personality dynamics.

Significance. The paper is a carefully executed and transparent empirical study. Its strengths include the controlled factorial persona design, the use of external human-change priors rather than priors fitted to model outputs, the explicit handling of item-level reliability and directional consistency, and a thorough validation suite that separates event-conditioned signal from retest noise, paraphrase instability, behavioral non-convergence, and short-range decay. The BFI-Adapt benchmark is a reusable resource, and the promised release of code, scenarios, and per-model logs supports reproducibility. If the interpretive issues below are resolved, the four-axis diagnostic would be a useful template for evaluating the psychological plausibility of LLM personas. I also credit the authors for stating limitations in Section 8, although two of those limitations are more consequential for the headline claim than the current framing suggests.

major comments (3)
  1. [§4.5, Fig. 6] The headline claim that PC-Agents 'simulate the mean but not its shape' rests in part on RQ4, which reports a three- to four-fold compression of persona-level change dispersion (median σ_LLM = 0.19 versus a human envelope of 0.5–0.8). However, σ_LLM is the standard deviation of within-person change scores Δ, while the cited human envelope from Bühler et al. (2024) and Roberts et al. (2006) is a trait-level between-person SD used to standardize mean-level change. These are different quantities: with the high rank-order stability of personality traits, the SD of person-level change scores can be substantially smaller than the trait SD for the same sample. The Section 8 note that the envelope is a 'coarse meta-analytic range' does not address this construct mismatch. The manuscript should either benchmark σ_LLM against the SD of human change scores for comparable events and intervals, or explicitly restrict the RQ4 claim to 'compressed relative to baseline trait SD' and temper the 'not its shape' conclusion accordingly.
  2. [§4.3, §8] The RQ2 magnitude calibration compares immediate post-event BFI-44 changes to a standardized mean-change band derived from longitudinal meta-analyses spanning months to years. The measurement here is a single self-report immediately after an event notification and a short reflection, which Section 8 acknowledges is far short of the multi-year horizons of human panels. This timescale mismatch is not merely a caveat: the low in-range rate (11.0–16.4%) and the 'under-shift' interpretation may reflect the difference between an instantaneous response and a long-term adaptation target. The authors should either justify the mapping from immediate model output to longitudinal human change, or reframe the magnitude results as exploratory and remove them from the central 'mean but not shape' conclusion.
  3. [§3.1, §4.4] RQ3's null demographic result is presented as evidence that PC-Agents lack demographic shape, but the design has limited sensitivity: each gender × cultural-region stratum contains only 10 personas, the cultural manipulation is a single continent label in the system prompt, and stratum medians are based on 10 observations. The Fisher tests on binary match/mismatch may have low power to detect moderation. The conclusion should be softened to 'no detectable demographic moderation under this coarse manipulation' rather than stated as 'demographics do not measurably change the post-event BFI item ratings.'
minor comments (4)
  1. [Fig. 1 vs §4.3 and Fig. 4] The Pearson correlation between κ and DCR is reported as +0.226 in the Figure 1 overview and as +0.455 in Section 4.3 and Figure 4; please reconcile these values.
  2. [Abstract and §4.1] The paper fluctuates between an 11-model main grid and a 14-model leaderboard; the abstract should state clearly that the four-axis analysis covers 11 models and that the BFI-Adapt leaderboard extends to 14.
  3. [Appendix E] The 'chance baseline implied by ε = 0.1 on a 5-point Likert scale' for the RQ1 movement tests is never derived; please state the null model used for pct_moved.
  4. [Eq. (5)] The BFI-Adapt direction indicator uses the pair's dominant DCR direction rather than the modal persona-level direction; please explain why item-level dominance is the appropriate level for directional fidelity, since the two can disagree when item-level changes are concentrated in a few items.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims are anchored to external human longitudinal priors and validated with independent control conditions.

full rationale

The paper's central derivation is self-contained against external human psychology references rather than against its own fitted quantities. The directional priors for the 11 life events are taken from Specht (2017), Bühler et al. (2024), Roberts et al. (2006), Boyce et al. (2015), and Schwaba and Bleidorn (2019), which are independent meta-analytic and longitudinal sources. The match rate, DCR, kappa, and BFI-Adapt composite are computed from paired BFI-44 item ratings of LLM personas; none of these quantities is fitted to force a headline outcome. The no-event retest, independent paraphrase, scenario-decision, and delayed-measurement conditions provide external robustness checks whose thresholds come from the models' own retest floors rather than from circular definitions. Self-citations such as AnnaAgent and GenPT appear in related-work and motivation contexts, but the load-bearing evidence for the 'simulate the mean, not its shape' conclusion comes from external human priors and model response measurements. The RQ4 comparison of LLM change-score SD against a human within-trait SD envelope may involve a construct mismatch between trait-level dispersion and change-score dispersion, and the paper itself calls the envelope a 'coarse meta-analytic range'; however, this is a benchmark-validity concern, not a circular derivation. No equation in the paper reduces to its own inputs, and no fitted parameter is renamed as a prediction.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central measurements rest on three hand-set calibration constants (epsilon, sigma, human SD envelope) and on the assumption that BFI-44 self-reports from LLMs validly reflect personality constructs. The directional priors are external to the model outputs, so they are assumptions about ground truth rather than free parameters fitted to the data.

free parameters (3)
  • Noise threshold epsilon = 0.1
    Direction classification boundary in Eq. 1; chosen as the smallest one-item step of BFI-44 trait means (1/n_O=0.10). It determines pct_moved and all match rates, so it directly shapes the existence and direction findings.
  • Representative BFI trait SD sigma = approximately 0.7
    Used to convert standardized human effect sizes |d| in [0.05, 0.20] into the raw Likert reference band [0.035, 0.14] for magnitude calibration (Section 4.3). Stated as representative, from Bühler et al. (2024); not model-fitted.
  • Human within-trait SD envelope = 0.5 to 0.8
    Comparison target for RQ4 heterogeneity; if this envelope does not represent change-score SDs for the 11 events, the heterogeneity-collapse finding is overstated.
assumptions (4)
  • domain assumption BFI-44 measures Big Five traits faithfully when administered to LLMs in character
    The entire pipeline assumes inventory responses reflect a stable latent personality construct in the model. Prior static-assessment work supports it, but it remains an interpretive assumption.
  • domain assumption The expected-direction table from Specht (2017) is correct for the 11 life events
    The benchmark's directional-fidelity score treats these priors as ground truth; some rows have meta-analytic support (job entry, unemployment, retirement, illness) but the full 55-pair table is adopted verbatim from one source.
  • domain assumption Human longitudinal effect sizes and SDs are applicable as point benchmarks
    The magnitude band and heterogeneity envelope are taken as fixed targets for all event-trait pairs, ignoring documented heterogeneity across events and traits; the authors acknowledge these are coarse ranges.
  • standard math Ordered Likert item responses can be treated as interval with linear weights
    The linearly weighted Cohen's kappa and trait averaging assume equal-interval scaling; standard in psychometrics but not universally valid.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events." pith.science (2026). https://pith.science/paper/MQPMOWDW

@misc{pith2026260806485,
  author       = {Pith},
  title        = {Pith review of: Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MQPMOWDW}},
  note         = {Machine review of arXiv:2608.06485}
}
read the original abstract

Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts. Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology. Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event-trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples. To enable systematic comparison, we introduce BFI-Adapt, a reusable benchmark for scoring the directional fidelity of event-induced personality change, and use it to rank 14 models. A validation suite shows that the measured shifts exceed no-event retest noise, remain stable under independently paraphrased prompts, exhibit limited and model-dependent convergence with scenario-based behavioral choices, and persist across intervening unrelated dialogue. Together, these checks establish the measured trajectories as robust event-conditioned response patterns. Our results suggest that current PC-Agents simulate the mean of human personality dynamics, but not its shape.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

48 extracted references · 28 canonical work pages

  1. [1]

    and Naumann, Laura P

    John, Oliver P. and Naumann, Laura P. and Soto, Christopher J. , title =. Handbook of Personality: Theory and Research , editor =. 2008 , publisher =

  2. [2]

    and Walton, Kate E

    Roberts, Brent W. and Walton, Kate E. and Viechtbauer, Wolfgang , title =. Psychological Bulletin , volume =. 2006 , doi =

  3. [3]

    , title =

    Specht, Jule and Egloff, Boris and Schmukle, Stefan C. , title =. Journal of Personality and Social Psychology , volume =. 2011 , doi =

  4. [4]

    and Lucas, Richard E

    Bleidorn, Wiebke and Hopwood, Christopher J. and Lucas, Richard E. , title =. Journal of Personality , volume =. 2018 , doi =

  5. [5]

    Denissen, Jaap J. A. and Luhmann, Maike and Chung, Joanne M. and Bleidorn, Wiebke , title =. Journal of Personality and Social Psychology , volume =. 2019 , doi =

  6. [6]

    and Wood, Alex M

    Boyce, Christopher J. and Wood, Alex M. and Daly, Michael and Sedikides, Constantine , title =. Journal of Applied Psychology , volume =. 2015 , doi =

  7. [7]

    Journal of Personality and Social Psychology , volume =

    Schwaba, Ted and Bleidorn, Wiebke , title =. Journal of Personality and Social Psychology , volume =. 2019 , doi =

  8. [8]

    2020 , journal =

    Trajectories of Big Five Personality Traits: A Coordinated Analysis of 16 Longitudinal Samples , author =. 2020 , journal =

Show all 48 references
  1. [9]

    , title =

    Rammstedt, Beatrice and John, Oliver P. , title =. Journal of Research in Personality , volume =. 2007 , doi =

  2. [10]

    and John, Oliver P

    Soto, Christopher J. and John, Oliver P. , title =. Journal of Research in Personality , volume =. 2017 , doi =

  3. [11]

    and Costa, Paul T

    McCrae, Robert R. and Costa, Paul T. , title =. American Psychologist , volume =. 1997 , doi =

  4. [12]

    and Demetre, James D

    Robinson, Oliver C. and Demetre, James D. and Corney, Roslyn , title =. Personality and Individual Differences , volume =. 2010 , doi =

  5. [13]

    , title =

    Luhmann, Maike and Hofmann, Wilhelm and Eid, Michael and Lucas, Richard E. , title =. Journal of Personality and Social Psychology , volume =. 2012 , doi =

  6. [14]

    and Cai, Carrie J

    Park, Joon Sung and O'Brien, Joseph C. and Cai, Carrie J. and Morris, Meredith Ringel and Liang, Percy and Bernstein, Michael S. , title =. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST) , pages =. 2023 , doi =

  7. [15]

    2026 , eprint=

    LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals , author=. 2026 , eprint=

  8. [16]

    P ersona LLM : Investigating the Ability of Large Language Models to Express Personality Traits

    Jiang, Hang and Zhang, Xiajie and Cao, Xubo and Breazeal, Cynthia and Roy, Deb and Kabbara, Jad. P ersona LLM : Investigating the Ability of Large Language Models to Express Personality Traits. Findings of the Association for Computational Linguistics: NAACL 2024. 2024. doi:10...

  9. [17]

    , title =

    Huang, Jen-tse and Wang, Wenxuan and Li, Eric John and Lam, Man Ho and Ren, Shujie and Yuan, Youliang and Jiao, Wenxiang and Tu, Zhaopeng and Lyu, Michael R. , title =. The Twelfth International Conference on Learning Representations , year =

  10. [18]

    G en PT : Beyond Self-Report for Reliable LLM Psychometrics via Generative Projective Testing

    Wang, Ming and Wu, Shuang and Wang, Bixuan and Lin, Lu and Chen, Yuxin and Yang, Xiaocui and Wang, Daling and Feng, Shi and Zhang, Yifei and Sun, Yufan. G en PT : Beyond Self-Report for Reliable LLM Psychometrics via Generative Projective Testing. Proceedings of the 64th Annua...

  11. [19]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL) , pages =

    Wang, Xintao and Xiao, Yunze and Huang, Jen-tse and Yuan, Siyu and Xu, Rui and Guo, Haoran and Tu, Quan and Fei, Yaying and Leng, Ziang and Wang, Wei and Chen, Jiangjie and Li, Cheng and Xiao, Yanghua , title =. Proceedings of the 62nd Annual Meeting of the Association for Com...

  12. [20]

    Wang, Noah and Peng, Z.y. and Que, Haoran and Liu, Jiaheng and Zhou, Wangchunshu and Wu, Yuhan and Guo, Hongcheng and Gan, Ruitong and Ni, Zehao and Yang, Jian and Zhang, Man and Zhang, Zhaoxiang and Ouyang, Wanli and Xu, Ke and Huang, Wenhao and Fu, Jie and Peng, Junran. R ol...

  13. [21]

    Capturing Minds, Not Just Words: Enhancing Role-Playing Language Models with Personality-Indicative Data

    Ran, Yiting and Wang, Xintao and Xu, Rui and Yuan, Xinfeng and Liang, Jiaqing and Xiao, Yanghua and Yang, Deqing. Capturing Minds, Not Just Words: Enhancing Role-Playing Language Models with Personality-Indicative Data. Findings of the Association for Computational Linguistics...

  14. [22]

    arXiv preprint arXiv:2401.01275 , year =

    Tu, Quan and Fan, Shilong and Tian, Zihang and Yan, Rui , title =. arXiv preprint arXiv:2401.01275 , year =. doi:10.48550/arXiv.2401.01275 , url =

  15. [23]

    arXiv preprint arXiv:2307.00184 , year=

    Personality Traits in Large Language Models , author=. arXiv preprint arXiv:2307.00184 , year=. 2307.00184 , archivePrefix=

  16. [24]

    Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =

    Xing, Jiayi and Niu, Tong and Srivastava, Shachi , title =. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =. 2025 , doi =

  17. [25]

    Computational Linguistics , volume =

    Zheng, Jinyu and Wang, Xinyu and Hosio, Simo and Xu, Xuhai and Lee, Lik-Hang , title =. Computational Linguistics , volume =. 2025 , doi =

  18. [26]

    and He, Lin and Xu, Xin , title =

    Wang, Yongqi and Zhao, Junwei and Ones, Deniz S. and He, Lin and Xu, Xin , title =. Scientific Reports , volume =. 2025 , doi =

  19. [27]

    arXiv preprint arXiv:2407.11484 , year =

    Chen, Ning and Wang, Yu and Deng, Yang and Li, Jing , title =. arXiv preprint arXiv:2407.11484 , year =

  20. [28]

    arXiv preprint arXiv:2509.04794 , year =

    Handa, Gaurish and Wu, Zhijin and Koshiyama, Adriano and Treleaven, Philip , title =. arXiv preprint arXiv:2509.04794 , year =

  21. [29]

    Artificial Intelligence Review , author =

    Validation is the central challenge for generative social simulation: a critical review of. Artificial Intelligence Review , author =. 2025 , pages =. doi:10.1007/s10462-025-11412-6 , abstract =

  22. [30]

    , title =

    Wilson, Edwin B. , title =. Journal of the American Statistical Association , volume =. 1927 , doi =

  23. [31]

    Journal of the Royal Statistical Society: Series B , volume =

    Benjamini, Yoav and Hochberg, Yosef , title =. Journal of the Royal Statistical Society: Series B , volume =. 1995 , doi =

  24. [32]

    Life Events and Personality Change: A Systematic Review and Meta-Analysis , journal =

    B. Life Events and Personality Change: A Systematic Review and Meta-Analysis , journal =. 2024 , doi =

  25. [33]

    Personality Testing of Large Language Models: Limited Temporal Stability, But Highlighted Prosociality , journal =

    Bodro. Personality Testing of Large Language Models: Limited Temporal Stability, But Highlighted Prosociality , journal =. 2024 , doi =

  26. [34]

    arXiv preprint arXiv:2602.00016 , year =

    Yu, Jiongchi and Ma, Yuhan and Zhang, Xiaoyu and Wang, Junjie and Hu, Qiang and Shen, Chao and Xie, Xiaofei , title =. arXiv preprint arXiv:2602.00016 , year =

  27. [35]

    Michael , title =

    Han, Pengrui and Kocielnik, Rafal and Song, Peiyang and Debnath, Ramit and Mobbs, Dean and Anandkumar, Anima and Alvarez, R. Michael , title =. arXiv preprint arXiv:2509.03730 , year =. doi:10.48550/arXiv.2509.03730 , url =

  28. [36]

    A Systematic Analysis of the Impact of Persona Steering on

    Chen, Jiaqi and Wang, Ming and Xie, Tingna and Feng, Shi and Liu, Yongkang , booktitle =. A Systematic Analysis of the Impact of Persona Steering on. 2026 , url =

  29. [37]

    A nna A gent: Dynamic Evolution Agent System with Multi-Session Memory for Realistic Seeker Simulation

    Wang, Ming and Wang, Peidong and Wu, Lin and Yang, Xiaocui and Wang, Daling and Feng, Shi and Chen, Yuxin and Wang, Bixuan and Zhang, Yifei. A nna A gent: Dynamic Evolution Agent System with Multi-Session Memory for Realistic Seeker Simulation. Findings of the Association for ...

  30. [38]

    From Pattern Recognizers to Personalized Companions: a Survey of Large Language Models in Mental Health , year=

    Hu, He and Zhou, Yucheng and Wang, Qianning and Zou, Yingjian and Ma, Chiyuan and Si, Juzheng and Liu, Jianzhuang and Yu, Zitong and Cui, Laizhong and Ma, Fei and Tian, Qi , journal=. From Pattern Recognizers to Personalized Companions: a Survey of Large Language Models in Men...

  31. [39]

    ACM Comput

    Mou, Xinyi and Ding, Xuanwen and He, Qi and Wang, Liang and Liang, Jingcong and Zhang, Xinnong and Sun, Libo and Lin, Jiayu and Zhou, Jie and Xuanjing, Huang and Wei, Zhongyu , title =. ACM Comput. Surv. , month = apr, articleno =. 2026 , issue_date =. doi:10.1145/3800683 , abstract =

  32. [40]

    Personality development in reaction to major life events , editor =

    Jule Specht , keywords =. Personality development in reaction to major life events , editor =. Personality Development Across the Lifespan , publisher =. 2017 , isbn =. doi:10.1016/B978-0-12-804674-6.00021-1 , url =

  33. [41]

    2026 , eprint=

    OpenAI GPT-5 System Card , author=. 2026 , eprint=

  34. [42]

    2024 , eprint=

    GPT-4 Technical Report , author=. 2024 , eprint=

  35. [43]

    2025 , eprint=

    Gemini: A Family of Highly Capable Multimodal Models , author=. 2025 , eprint=

  36. [44]

    DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence , author=

  37. [45]

    2026 , howpublished=

    MiMo-V2.5-Pro , author=. 2026 , howpublished=

  38. [46]

    2026 , eprint=

    Kimi K2: Open Agentic Intelligence , author=. 2026 , eprint=

  39. [47]

    2025 , eprint=

    GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models , author=. 2025 , eprint=

  40. [48]

    2025 , eprint=

    Qwen3 Technical Report , author=. 2025 , eprint=

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.