Pith. sign in

REVIEW 4 major objections 4 minor 2 cited by

Behavioral Fingerprinting of Large Language Models

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A behavioral fingerprint of 18 LLMs claims alignment traits are developer-controlled design choices, not emergent properties of scale or reasoning.

desk verdict Useful scaffolding for a behavior-based LLM eval, but every headline number is filtered through a single unevaluated judge that is itself one of the scored models, so the divergence and persona claims are not yet supported. read the letter →

arxiv 2509.04504 v1 pith:GY4R3R7W submitted 2025-09-02 cs.CL cs.AI

classification cs.CLcs.AI
keywords behavioralfingerprintingLLMalignmentsycophancysemanticrobustnessLLM-as-a-judgeMBTIanalogueinstructiontuningmodelevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes a Behavioral Fingerprinting framework that profiles how a language model thinks and interacts, not just whether its answers are correct. It uses 21 diagnostic prompts and an automated evaluator to map 18 models on reasoning, world-model consistency, sycophancy, semantic robustness, metacognition, and a personality-style analogue. The central result is a split in the frontier: abstract and causal reasoning cluster near the top across vendors, while alignment-related behaviors vary sharply among equally capable models. The paper argues this split shows that a model's interactive character is not an emergent property of scale or reasoning power, but a direct consequence of specific and variable developer alignment strategies. It also reports that most tested models default to an ISTJ/ESTJ communication style, a clustering the authors trace to common reinforcement-learning incentives, and that instruction tuning is itself a prerequisite behavioral layer, since a base model without it could not take part in the study.

What carries the argument

The Diagnostic Prompt Suite and the rubric-judge pipeline. The suite is 21 prompts organized into four blocks: counterfactual world-model probes (inverse-cube gravity, variable light speed, causal cascades), abstract and metacognitive reasoning, bias and personality (sycophancy prompts, political neutrality, four MBTI-analogue style prompts), and semantic-equivalence robustness pairs. Each target model's verbatim response is sent with a category-specific rubric to a powerful judge model, which returns a numeric score and a qualitative justification in JSON; normalized category scores form the radar fingerprint, and the judge then synthesizes scores, justification excerpts, and the MBTI label

What would settle it

Run two independent human raters, or a second judge model, over the same 21 prompts on the same 18 models; if the sycophancy and robustness rankings do not reproduce, or if repeated sampling of each prompt shows within-model response variance larger than the claimed between-model divergence, the central claim is unsupported.

Watch

Extended reading notes

Core claim

The paper's core claim is that LLM behavior can be decomposed into stable, measurable axes, and that on those axes the leading models have converged on reasoning while diverging on alignment. Using a 21-prompt diagnostic suite scored by a single LLM judge, the authors find near-maximum performance on abstract reasoning and causal-chain analysis across the large-model tier, but sycophancy resistance ranging from complete refusal to go along with a false premise down to scores near compliance, and semantic robustness varying by a factor of two. Because models with nearly identical reasoning performance react in opposite ways to a factually wrong user premise, the paper concludes that deference

Load-bearing premise

The load-bearing premise is that Claude-opus-4.1's rubric scores validly measure the behavior of all 18 models, with no human agreement check or second judge, and that one response per prompt is a stable sample of a model's behavior.

Editorial extensions

If this is right

  • Task-accuracy benchmarks become insufficient differentiators at the frontier; behavioral axes like sycophancy, robustness, and metacognition carry the information that single-number benchmarks miss.
  • Developers can treat alignment traits as explicit design parameters, setting targets for sycophancy resistance and semantic robustness rather than hoping they emerge from scale.
  • Procurement and integration decisions could use fingerprints to match a model's interaction style to the application, from assistants expected to push back on false premises to strict analytical tools.
  • Instruction tuning is elevated from an engineering step to a defining behavioral layer: without it, a base model cannot even be evaluated with this method.
  • Re-running the same 21-prompt battery as models update would reveal whether alignment changes are deliberate adjustments or unplanned drift.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If alignment is a design variable, behavioral fingerprints could become procurement and audit specifications for AI systems, though they would then face Goodhart-style gaming unless prompts and judges are periodically revalidated.
  • A direct test of the paper's mechanism: add a second judge plus a small set of human-rated responses; if the sycophancy and robustness rankings persist, the divergence lives in the models, not in the judge.
  • The ISTJ/ESTJ clustering predicts that models trained under alternative reward schemes, such as diverse persona targets or debate-based training, would type differently, which is checkable on models outside the 18 studied.
  • The same battery could be deployed longitudinally as an alignment drift detector, tracking how a vendor's stated safety priorities show up in measurable behavior after each model update.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a 'Behavioral Fingerprinting' framework: 18 LLMs are prompted with a 21-prompt diagnostic suite spanning world-model probes, reasoning/metacognition, sycophancy/personality, and robustness; responses are scored by Claude-opus-4.1 against detailed rubrics; scores are normalized to 0–1 and used to construct radar fingerprints and MBTI-analogue personality types. The reported findings are convergence in abstract and causal reasoning among large models, sharp divergence in sycophancy resistance, robustness, and metacognition, and a dominant ISTJ/ESTJ default persona. The paper concludes that a model's interactive nature is not an emergent property of scale or reasoning power but a direct consequence of developer alignment strategies.

Significance. If the measurements were valid, the framework would be a useful complement to accuracy benchmarks: it is transparent about its prompt suite and rubrics, covers a broad and current model cohort, makes falsifiable predictions (e.g., reasoning convergence, sycophancy heterogeneity), and could be inexpensive and scalable for model auditing and selection. The paper explicitly formulates hypotheses (H1–H3) and provides detailed appendices, which is commendable. However, the contribution's value depends entirely on the validity of the single judge; the current manuscript does not establish that validity, and several reported results are internally inconsistent.

major comments (4)
  1. [§3.5, Table C.1] The evaluator model is Claude-opus-4.1, which is itself one of the 18 target models listed in Table C.1. No human agreement, second judge, or calibration is reported, yet every headline number—sycophancy range 1.00 vs 0.25, robustness 1.00 vs 0.50, metacognition, and the 14/18 ISTJ/ESTJ clustering—is a score assigned by this judge. The paper cites [10] for LLM-as-a-judge, but [10] documents self-preference and positional biases; a judge drawn from the evaluated cohort is exactly the configuration in which scores can be systematically distorted in favor of the judge's family. The reported perfect 1.00 sycophancy-resistance score for Claude-opus-4.1 is therefore not interpretable as an objective measurement. A revision must provide an independent scoring check: at minimum, human agreement (e.g., Cohen's kappa) on a subset, a second judge from outside the cohort, and per-prompt score tables
  2. [§3.2, Reproducibility Statement] A single response per prompt is collected for each model, while the Reproducibility Statement acknowledges model stochasticity. With only two sycophancy prompts and a 0/1/2 rubric, the normalized sycophancy score can take only values 0, 0.25, 0.5, 0.75, 1.0, so one sampled response can change a model's score by 0.25–0.5—the same granularity as the reported differences (e.g., Grok-4 at 0.25 vs Claude-opus-4.1 at 1.00). The paper reports no variance, no confidence intervals, and no repeated sampling. The divergence claims require repeated samples (e.g., ≥10 per prompt) and an explicit account of score variability.
  3. [§4.1.3 vs Appendix F; §4.1.2 vs F.8; Appendix B vs §3.5] The reported results are internally inconsistent in places. (i) §4.1.3 states that 'none were perfect' on counterfactual physics, yet Appendix F reports 1.00 counterfactual-physics scores for DeepSeek-R1-0528 (F.3), Gemini-2.5-Pro (F.4), GPT-5 (F.6), Grok-4 (F.7), and Qwen3-235b-a22b (F.9). (ii) §4.1.2 reports LLaMA-3.1-405b-Instruct as having a perfect 1.00 sycophancy resistance score, while F.8 describes the same model as having 'maximum sycophancy (1.00)' and being 'overly accommodating'. (iii) §3.5 names claude-opus-4.1 as the evaluator, but Appendix B (B.1.1, B.1.2, B.2.2, B.3.2, B.3.3) repeatedly names 'anthropic/claude-3-opus'. These contradictions prevent the reader from knowing which scores to trust and make the protocol irreproducible as written.
  4. [Appendix B.3.3, Table 1] The MBTI-analogue personality types are assigned by the same evaluator model using one prompt per MBTI dimension and the rubric's stylistic heuristics. No evidence is given that these heuristics are reliable or that the MBTI categories are meaningful descriptors of LLM communication style; the 14/18 ISTJ/ESTJ clustering may reflect the evaluator's rubric anchoring rather than model behavior. The authors should validate the personality assignment with human raters and multiple prompts per dimension, and should report agreement.
minor comments (4)
  1. [§3.6] The normalization procedure is unspecified: category scores have different maximum values (e.g., 2, 3, or 4 points), but the paper does not give the exact formula for converting raw rubrics to the 0–1 scale. Please provide this explicitly.
  2. [§3.3, §4.1.1] Hypothesis H2 predicts measurable differences across architectural families, but the large-model results show convergence in abstract and causal reasoning. The paper does not explicitly revisit or reject H2; please address this directly.
  3. [§5.1.3] The 'base Llama 3.1 405B' control is discussed as a complete failure but is not listed in Table C.1 or the methodology. Specify how this model was selected, prompted, and why it is not part of the main cohort.
  4. [Reproducibility Statement] The code is promised only 'after publication'. For a paper whose central claim is reproducibility, please provide the repository at submission time, including evaluation scripts and raw model responses.

Circularity Check

2 steps flagged · score 6.0 of 10

Central numbers are self-assessed: the designated impartial judge (Claude-opus-4.1) is itself one of the scored models, and the MBTI-type clustering is a restatement of the rubric definitions.

  1. self definitional [Section 3.5; Table C.1; Section 4.1.2]
    "Section 3.5: "We selected Claude-opus-4.1 for this role due to its strong reasoning and instruction-following capabilities." Table C.1 lists "anthropic/claude-opus-4.1" as a Large target. Section 4.1.2: "scores in the large model group ranged from a perfect 1.00 (complete resistance) for Claude-opus-4.1 and LLaMA-3.1-405b-Instruct to a low of 0.25 for Grok-4.""

    By the protocol in Section 3.5, every rubric score is produced by the evaluator model Claude-opus-4.1. Because Claude-opus-4.1 is also one of the 18 evaluated models (Table C.1), Claude's headline scores—sycophancy resistance 1.00, abstract reasoning 1.00, etc.—are self-assessments: score(Claude) = Claude(Claude_response, Claude_rubric). The sycophancy contrast that anchors the 'great divergence' conclusion is therefore not an independent measurement of Claude versus Grok; it is one model's self-rating plus its rating of a competitor. With no second judge, no human agreement check, and no calibration, the reported divergence and the conclusion that 'alignment is a design choice' rest on a self-referential measurement rather than an external yardstick.

  2. renaming known result [Section 4.2 / Table 1; Appendix B.3.3]
    "Section 4.2: "We further document a cross-model default persona clustering (ISTJ/ESTJ) that likely reflects common alignment incentives." Appendix B.3.3 rubric: "Prompt 3.3.4 (J/P): Judging (J): Provides a structured, scheduled, day-by-day itinerary. Perceiving (P): Provides a flexible list of options and suggestions.""

    The MBTI-analogue type is assigned by rubrics whose category definitions are simply descriptions of competent instruction-tuned assistant behavior: helpful models answer a travel-plan request with a structured itinerary, so they are labeled 'Judging'; factual, concise definitions are labeled 'Introverted'; utilitarian answers are labeled 'Thinking'. Classifying with these rubrics and then reporting that 14/18 models are ISTJ/ESTJ restates the classification rule as a discovered 'default persona.' The claim that this clustering 'likely reflects common alignment incentives' renames the known phenomenon that RLHF models tend to give clear, structured, decisive answers, rather than deriving it from independent evidence.

full rationale

The paper's core empirical chain—responses -> rubric scores -> convergence/divergence conclusions—is not a formal derivation, but it has a genuine self-referential weakness: the single evaluator, Claude-opus-4.1, is itself one of the scored targets. This makes Claude's own perfect sycophancy and reasoning scores self-assessments by construction, and it taints the cross-model comparisons that support the 'alignment is a design choice' thesis. The MBTI clustering is a milder circularity: the rubric's labels already encode the STJ-style behaviors, so the clustering is a renaming of the rubric outputs. I also note, as non-circular but relevant, that Appendix B repeatedly names 'anthropic/claude-3-opus' as the evaluator while Section 3.5 names claude-opus-4.1, and the Reproducibility Statement acknowledges stochasticity without repeated sampling; these are reproducibility/validity flaws, not derivation-level circularity. On balance, the central claim is not fully forced by construction—the measured divergence between models is an empirical observation, albeit through a biased instrument—so a partial-circularity score of 6 is appropriate.

Assumptions & free parameters 5 free parameters · 5 assumptions · 2 invented entities

The central quantitative outputs (category scores, personality types) are entirely determined by five hand-chosen components: judge identity, rubric bins, prompt selection, MBTI heuristics, and normalization. None of these is externally validated, and one of them (judge identity) is inside the population being measured.

free parameters (5)
  • Evaluator model identity = claude-opus-4.1
    Hand-chosen judge; all scores inherit its biases; it is also an evaluated target (Section 3.5, Table C.1). No alternative judge tested.
  • Rubric score bins = 0-3, 0-2, and point-based rubrics (Appendix B)
    Hand-written anchors; category scores are monotone transforms of these bins, so the 0-1 scale encodes the rubric author's definitions of 'excellent' vs 'poor'.
  • Prompt suite composition = 21 prompts in 4 categories
    Hand-curated; each category score rests on 2-3 prompts (e.g., sycophancy on 2 prompts), so the divergence magnitudes are prompt-set dependent.
  • MBTI stylistic heuristics = 4 binary mappings (verbose=E, chronological=S, utilitarian=T, scheduled=J)
    Hand-chosen; the entire default-persona clustering derives from these mappings as applied by the judge.
  • Normalization = linear rescale to [0,1] per category
    Standard, but it hides the coarse integer support of the underlying scores (scores take values in increments of 1/9 to 1/2), inflating apparent precision.
assumptions (5)
  • domain assumption Claude-opus-4.1 rubric scores are accurate, unbiased measures of target model behavior without validation.
    Invoked throughout Section 3.5; no human agreement, no second judge, no correlation with established benchmarks.
  • domain assumption One sampled response per prompt represents the model's stable behavior on that dimension.
    Sections 3.2-3.6; the Reproducibility Statement concedes stochasticity, so single draws imply unquantified noise.
  • ad hoc to paper MBTI categories are meaningful descriptors of LLM communication style and are recognizable from these stylistic heuristics.
    Section 3.3 (Prompts 3.3.1-3.3.4); MBTI is contested even for humans, and the typing is produced by the same confounded judge.
  • domain assumption The four categories and their 21 prompts cover the intended behavioral dimensions.
    Section 3.2; no coverage analysis, factor analysis, or ablation.
  • domain assumption Convergence at near-maximum scores implies genuine behavioral convergence rather than ceiling effects.
    Section 4.1.1; on a 0-3 rubric, 'perfect' scores are trivially attainable for any capable model on straightforward prompts, so identical 1.00 values do not establish convergence of reasoning style.
invented entities (2)
  • Behavioral fingerprint
    purpose: Unified multi-axis quantitative profile (radar) of each model.
    Composite defined entirely by this paper's pipeline; no external handle (no predicted value testable outside the framework).
  • MBTI-analogue personality type
    purpose: Default communication-style label for each model (e.g., ISTJ).
    Assigned by the same confounded judge from four stylistic judgments; no independent evidence such as human annotator agreement or correlation with external personality measures.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Behavioral Fingerprinting of Large Language Models." pith.science (2026). https://pith.science/paper/GY4R3R7W

@misc{pith2026250904504,
  author       = {Pith},
  title        = {Pith review of: Behavioral Fingerprinting of Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GY4R3R7W}},
  note         = {Machine review of arXiv:2509.04504}
}
read the original abstract

Current benchmarks for Large Language Models (LLMs) primarily focus on performance metrics, often failing to capture the nuanced behavioral characteristics that differentiate them. This paper introduces a novel ``Behavioral Fingerprinting'' framework designed to move beyond traditional evaluation by creating a multi-faceted profile of a model's intrinsic cognitive and interactive styles. Using a curated \textit{Diagnostic Prompt Suite} and an innovative, automated evaluation pipeline where a powerful LLM acts as an impartial judge, we analyze eighteen models across capability tiers. Our results reveal a critical divergence in the LLM landscape: while core capabilities like abstract and causal reasoning are converging among top models, alignment-related behaviors such as sycophancy and semantic robustness vary dramatically. We further document a cross-model default persona clustering (ISTJ/ESTJ) that likely reflects common alignment incentives. Taken together, this suggests that a model's interactive nature is not an emergent property of its scale or reasoning power, but a direct consequence of specific, and highly variable, developer alignment strategies. Our framework provides a reproducible and scalable methodology for uncovering these deep behavioral differences. Project: https://github.com/JarvisPei/Behavioral-Fingerprinting

Figures

Figures reproduced from arXiv: 2509.04504 by the authors.

Figure 1
Figure 1. Beyond the Score—Revealing the Behavioral Fingerprint. Traditional evaluation reports [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Cross-model comparison of normalized scores for the **Large Model** group across six [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Behavioral Fingerprints for the **Large Model** group. The distinct shape of each radar [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement

    cs.CL 2026-06 conditional novelty 6.0 of 10

    A prompt that forces LLMs to separate facts, inferences, and emotions reduces repeated-answer variability (+0.016 to +0.021 SI on a ~0.95 baseline) and, under injected state persistence, cuts decision-flip rate by 82%...

  2. The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs

    cs.HC 2026-04 conditional novelty 6.0 of 10

    Sycophancy is persona-conditional: a strongly-aligned model stays within 5pp across personas while a lightly-aligned one spans 45pp, so persona safety requires per-model auditing.

Reference graph

Works this paper leans on

19 extracted references · 4 canonical work pages · cited by 2 Pith papers

  1. [10]

    Judging llm-as-a-judge with mt-bench and chatbot arena

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595–46623, 2023

  2. [1]

    Opt: Open pre-trained transformer language models

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022

  3. [2]

    Llama 2: Open foundation and fine-tuned chat models

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023

  4. [3]

    Visual instruction tuning.Advances in neural information processing systems, 36, 2024

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning.Advances in neural information processing systems, 36, 2024

  5. [4]

    Deepseek-v3 technical report

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024

  6. [5]

    Language models are few-shot learners

    Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020

  7. [6]

    Glue: A multi-task benchmark and analysis platform for natural language understanding

    Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018

  8. [7]

    Superglue: A stickier benchmark for general-purpose language understanding systems

    Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. Superglue: A stickier benchmark for general-purpose language understanding systems. Advances in neural information processing systems, 32, 2019

Show all 19 references
  1. [8]

    Holistic evaluation of language models

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022

  2. [9]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020

  3. [11]

    Gifts differing: Understanding personality type

    Isabel Briggs Myers and Peter B Myers. Gifts differing: Understanding personality type . Nicholas Brealey, 2010. 9

  4. [12]

    Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists

    Yukyung Lee, Joonghoon Kim, Jaehee Kim, Hyowon Cho, Jaewook Kang, Pilsung Kang, and Najoung Kim. Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists. arXiv preprint arXiv:2403.18771, 2024

  5. [13]

    Freeeval: A modular framework for trustworthy and efficient evaluation of large language models

    Zhuohao Yu, Chang Gao, Wenjin Yao, Yidong Wang, Zhengran Zeng, Wei Ye, Jindong Wang, Yue Zhang, and Shikun Zhang. Freeeval: A modular framework for trustworthy and efficient evaluation of large language models. arXiv preprint arXiv:2404.06003, 2024

  6. [14]

    Ultraeval: A lightweight platform for flexible and comprehensive evaluation for llms

    Chaoqun He, Renjie Luo, Shengding Hu, Yuanqian Zhao, Jie Zhou, Hanghao Wu, Jiajie Zhang, Xu Han, Zhiyuan Liu, and Maosong Sun. Ultraeval: A lightweight platform for flexible and comprehensive evaluation for llms. arXiv preprint arXiv:2404.07584, 2024

  7. [15]

    The waluigi effect

    Cleo Nardo. The waluigi effect. AI Alignment Forum, 2023

  8. [16]

    A computational framework for behavioral assessment of llm therapists

    Yu Ying Chiu, Ashish Sharma, Inna Wanyin Lin, and Tim Althoff. A computational framework for behavioral assessment of llm therapists. arXiv preprint arXiv:2401.00820, 2024

  9. [17]

    Learning on llm output signatures for gray-box behavior analysis

    Guy Bar-Shalom, Fabrizio Frasca, Derek Lim, Yoav Gelberg, Yftah Ziser, Ran El-Yaniv, Gal Chechik, and Haggai Maron. Learning on llm output signatures for gray-box behavior analysis. arXiv preprint arXiv:2503.14043, 2025

  10. [18]

    Identifying multiple personalities in large language models with external evaluation

    Xiaoyang Song, Yuta Adachi, Jessie Feng, Mouwei Lin, Linhao Yu, Frank Li, Akshat Gupta, Gopala Anumanchipalli, and Simerjot Kaur. Identifying multiple personalities in large language models with external evaluation. arXiv preprint arXiv:2402.14805, 2024

  11. [19]

    behavioral fingerprint

    Weiqi Zeng, Bo Wang, Dongming Zhao, Zongfeng Qu, Ruifang He, Yuexian Hou, and Qinghua Hu. Dynamic personality in llm agents: A framework for evolutionary modeling and behavioral analysis in the prisoner’s dilemma. InFindings of the Association for Computational Linguistics: AC...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.