REVIEW 3 major objections 4 minor 14 references
Under a controlled pre-registered test, a non-programmer persona made one frontier model decline to code in 12 of 60 responses while another model ignored the persona entirely.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 17:59 UTC pith:MLWUOLZC
load-bearing objection A careful, pre-registered study with a genuinely striking post-hoc observation about a librarian persona making Claude refuse to code — but the confirmatory evidence for model dependence is fragile, and the paper already half-admits it. the 3 major comments →
The Librarian Who Refused to Code: Model-Dependent Identity Enactment in LLM Code Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that biographical personas are not neutral context or universal quality interventions; they operate as model-dependent behavioral-policy biases. On Claude Opus, a persona with zero behavioral instruction—a research librarian named Linnea—caused the model to generalize from identity to action: it hedged with role-consistent disclaimers in 55 of 60 responses and declined to write any code in 12 of 60, a behavior stated nowhere in the prompt. On GPT-5.5 the same persona produced zero disclaimers and zero genuine refusals. The pre-registered condition-by-model interaction is significant for output length (on both provider-reported tokens and a shared visible-character measur
What carries the argument
The load-bearing objects are the biographically detailed persona texts—especially the 'Linnea' research librarian, which contains no code-, length-, or behavior-related instruction—and the pre-registered mixed-effects analysis with a condition×model interaction term. The identity-enactment phenomenon is quantified through fixed, post-hoc operationalizations: an in-character disclaimer is a first-person librarian/non-programmer assertion near the start of the response, and a genuine no-code response is one with zero fenced code blocks that terminated normally (not via length truncation). These textual definitions, combined with the sharp contrast between the two models, carry the argument tha
Load-bearing premise
The load-bearing premise is that the pre-registered condition-by-model interaction is a robust signal of model dependence—a premise that the paper's own task-clustered sensitivity analysis shows to be fragile, since the provider-token interaction loses significance when one prompt-contaminated task is excluded and the strongest surviving interaction uses a measure elevated to primary only post hoc.
What would settle it
A pre-registered replication that removes the R-TS-2 prompt leak, fixes the visible-character interaction as the primary outcome, and applies a task-clustered test; if the condition×model interaction on that measure fails to meet the pre-registered threshold, the confirmatory basis for model dependence would be falsified. Alternatively, running the same Linnea persona on a third frontier model family and observing disclaimers or refusals at rates comparable to Claude Opus would contradict the claimed model-family specificity.
If this is right
- Personas can change task willingness, not just style: a benign, instruction-free identity can cause a model to hedge or refuse to code, so persona design should be treated as a soft policy override.
- No persona tested improved correctness; on tasks near saturation, persona selection is a matter of fit between activated behaviors and the objective, not a universal quality lever.
- Model families differ in persona susceptibility; the same persona-bearing prompt can be inert on one model and disruptive on another, raising silent-degradation risks in multi-model pipelines.
- The engineer-persona length effects are confounded with an explicit style sentence, so narrative-only causation is established only for the librarian persona and only on one model family.
- The signed pre-registration with hard execution gates offers a reusable methodology for behavioral LLM experiments, independent of the persona results themselves.
Where Pith is reading between the lines
- The identity-enactment result suggests a security-relevant vector: a persona injected through an untrusted channel (e.g., a simulated teammate in a shared agent workspace) might induce silent refusal or hedging without any malicious instruction; the paper flags this as a hypothesis, and I would add that an adversary could craft personas to exploit specific model-family susceptibilities.
- If model dependence generalizes, persona evaluation on a single model family has little predictive value for another family; benchmark designers should report model×condition interactions and cluster-robust tests, not only pooled persona effects.
- The paper's own task-clustered sensitivity analysis shows the confirmatory interaction on the pre-registered provider-token measure is fragile (narrowly above threshold, and not significant when one prompt-contaminated task is removed), implying the model-dependence claim as confirmatory rests largely on a post-hoc visible-character measure; a replication with a pre-registered cluster-robust desig
- Because baseline correctness is saturated, the hypothesis that personas can improve quality on harder tasks remains completely open; a natural extension is to run the same design on tasks with headroom to see whether any persona raises correctness above baseline.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a pre-registered experiment with 480 completions crossing four prompt conditions (no persona, two engineer personas, a research-librarian persona), 12 code-generation tasks, two frontier models, and five runs per cell. Its central claims are: (i) a pre-registered condition×model interaction shows that persona effects are model-dependent; (ii) personas act as behavioral-policy biases rather than universal quality levers; and (iii) an exploratory finding that the librarian persona induces in-character disclaimers and no-code refusals on Claude Opus (55/60 disclaimers, 12/60 no-code) but not on GPT-5.5. The paper transparently marks confirmatory versus exploratory analyses, releases raw logs, derived scores, a classification script, and a pre-registration, but not the executable task harness.
Significance. If the identity-enactment result replicates under a pre-registered operationalization, it is significant: it would show that an instruction-free biographical persona can change not only style but task willingness, and do so differentially across model families. The methodological apparatus—externally signed pre-registration, execution gates, amendment logging, and partial artifact release—is a genuine contribution to behavioral LLM evaluation. However, the confirmatory support for the headline model-dependence claim is fragile: under the paper's own task-level clustered sensitivity analysis, the pre-registered provider-token interaction does not survive exclusion of one contaminated task, and the only interaction that remains strong under clustering uses a measure elevated to primary post hoc. The paper is mostly honest about these limitations, but the abstract and introduction overstate the confirmatory status of the model-dependence contribution.
major comments (3)
- [§4.2, Appendix C.2] The central claim that persona effects are model-dependent rests on H_INT, the condition×model interaction in the pre-registered mixed-effects model (outcome ~ C(condition)×C(model) + (1|task)). Appendix C.2 reports the exploratory task-clustered sensitivity: the correctness interaction does not survive (F(3,11)=2.47, p=0.116); the provider-token interaction passes only narrowly (F(3,11)=6.80, p=0.0074) and rises to p=0.0185 when the R-TS-2 task is excluded—above the pre-registered Bonferroni threshold. Because a random-intercept model with only 12 tasks can substantially understate between-task variability, and because the provider-token result is borderline and vanishes with one excluded task, the pre-registered analysis does not provide robust confirmatory support for model dependence. The paper should either adopt task-clustered inference as a confirmatory test or explicitly downgrad
- [§3.4, §4.2, §6] The pre-registration specified provider-reported completion tokens as E1; the paper elevated visible characters to primary “post hoc, for cross-provider comparability” (§3.4). This creates a circularity risk: the only interaction that survives task-level clustering is on visible characters (p=8.4e-6), a measure selected after seeing the data. The abstract honestly calls it post hoc, but the paper's conclusion and contribution list call the model-dependence result a “pre-registered demonstration.” The robust evidence is post hoc, and the language in the abstract and conclusion should be adjusted so that the confirmatory status of H_INT matches the analysis that actually supports it.
- [§3.2, Appendix B, Data availability] The executable task harness is not released; end-to-end re-execution of test-based scoring from raw prompts is impossible. Since correctness (Q1) is a central outcome—driving the identity-enactment no-code zeros and the disclaimer-correlation analysis—an independent reader cannot verify the reported correctness values. The release of derived scores is useful, but the claim of a “partial-reproducibility package” would be strengthened by releasing the harness or, failing that, by providing a detailed independent evaluation path for the correctness metric.
minor comments (4)
- [§1] Typo: “noimperativeinstructions” should be “no imperative instructions.”
- [§3.4] The notation for E1 visibility could be clarified: the paper uses “provider tokens” and “visible characters” but at times refers to “E1” without specifying which measure; please be consistent.
- [§6] The discussion of the R-TS-2 prompt leak is clear, but it would help to state explicitly that the leak cannot affect between-persona contrasts because the prompt is identical across conditions; currently this is stated only in Appendix C.2.
- [Appendix F.2] Table 3 would benefit from a note that the GPT-5.5 I-TS-1 cell has n=4 due to the excluded truncation artifact; this is stated in F.2 but not in the table caption.
Circularity Check
No significant circularity: the central claims are empirical, pre-registered where confirmatory, and post hoc measures are explicitly labeled as exploratory.
full rationale
The paper makes no derivation-from-first-principles claim that reduces to its inputs. Its central claim—model-dependent persona effects—is tested empirically via a pre-registered condition×model interaction on observed completions, not computed from a fitted parameter or imported from the authors' own prior work. The pre-registered mixed-effects analysis and the exploratory task-clustered sensitivity analysis are both reported; the fragility of the provider-token interaction under clustering and after excluding R-TS-2 is a statistical-inference and robustness concern, not a circular reduction. The identity-enactment result was operationalized post hoc after inspecting outputs, but the paper repeatedly and explicitly labels it exploratory/descriptive (§6, C.3, F), and the no-code outcome is an independently defined behavioral measure rather than a renamed input. There are no load-bearing self-citations, no imported uniqueness theorems, and no fitted parameters renamed as predictions. The post hoc elevation of visible characters to primary is a transparency/validity limitation, not a circular step. Therefore, on the evidence in the manuscript, no circular step meets the required evidentiary bar.
Axiom & Free-Parameter Ledger
free parameters (1)
- Identity-enactment classifier thresholds (proximity ≤40 chars; phrase list including 'wrong Linnea', 'not really my area =
Post hoc; yields 55/60 disclaimers and 12/60 no-code responses on Opus
axioms (5)
- domain assumption Test-based correctness is a valid measure of code quality
- domain assumption The two models' output lengths are comparable via visible character count
- domain assumption The Linnea persona text contains no behavioral instruction
- domain assumption The R-TS-2 user prompt leak does not bias between-persona contrasts
- standard math Mixed-effects model assumptions (random task intercept, independent errors)
read the original abstract
Biographical personas are widely used in system prompts, but their effects on code generation are rarely evaluated under controlled, pre-registered conditions. We tested four prompt conditions (no persona, two engineer personas, and a research-librarian persona), 12 code-generation tasks, two frontier models, and five runs per cell (480 completions). Persona effects differed between the two tested models. Under the pre-registered mixed-effects analysis, the condition-by-model interaction was significant for provider-reported output tokens; a post-hoc visible-character measure showed the same qualitative pattern. Six GPT-5.5 completions were length-capped and are reported separately. On Claude Opus, the minimalist engineer persona reduced visible output by 30% (33% in provider tokens) without improving correctness, while the thorough engineer persona increased output without a correctness gain. In an exploratory post-hoc analysis, the librarian persona elicited in-character disclaimers in 55 of 60 Opus responses and 12 genuine no-code responses, lowering mean correctness from 0.92 to 0.67. GPT-5.5 produced neither behavior in its 59 non-truncated responses. These results are consistent with personas acting as Model-Dependent behavioral-policy biases rather than universal quality interventions. We release raw completions, derived scores, analysis artifacts, a pre-registration document, and an execution gate log; end-to-end test-based rescoring requires an unreleased task harness.
Reference graph
Works this paper leans on
-
[1]
Pedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy, and Benjamin Roth. Principled Per- sonas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 26857–26886, 2025. doi:10.18653/v1/2025.emnlp-main.1364. arXiv:2508.19764
arXiv 2025
-
[2]
Marlene Lutz, Indira Sen, Georg Ahnert, Elisa Rogers, and Markus Strohmaier. The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 23212–23237, 2025. doi:10.18653/v1/2025.findings-emnlp.1261. arXiv:2507.16076
arXiv 2025
-
[3]
Persona Prompting as a Lens on LLM Social Reasoning
Jing Yang, Moritz Hechtbauer, Elisabeth Khalilov, Evelyn Luise Brinkmann, Vera Schmitt, and Nils Feldhus. Persona Prompting as a Lens on LLM Social Reasoning. InProceedings of the 2026 Conference of the European Chapter of the Association for Computational Linguistics (EACL), pages 1152–1170, 2026. doi:10.18653/v1/2026.eacl-long.52
-
[4]
SPeCtrum: A Grounded FrameworkforMultidimensionalIdentityRepresentationinLLM-BasedAgent
Keyeun Lee, Seo Hyeong Kim, Seolhee Lee, Jinsu Eun, Yena Ko, Hayeon Jeon, Esther Hehsun Kim, Seonghye Cho, Soeun Yang, Eun-mee Kim, and Hajin Lim. SPeCtrum: A Grounded FrameworkforMultidimensionalIdentityRepresentationinLLM-BasedAgent. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguist...
-
[5]
Zizhao Hu, Mohammad Rostami, and Jesse Thomason. Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM. arXiv:2603.18507, 2026
arXiv 2026
-
[6]
A Survey on Large Language Models for Code Generation
Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim. A Survey on Large Language Models for Code Generation. arXiv:2406.00515, 2024
Pith/arXiv arXiv 2024
-
[7]
Sarah Fakhoury, Aaditya Naik, Georgios Sakkas, Saikat Chakraborty, and Shuvendu K. Lahiri. LLM-Based Test-Driven Interactive Code Generation: User Study and Empirical Evaluation. IEEE Transactions on Software Engineering, 50:2254–2268, 2024. arXiv:2404.10100
Pith/arXiv arXiv 2024
-
[8]
Giovanni Rosa, David Moreno-Lumbreras, Gregorio Robles, and Jesús M. González-Barahona. Understanding Specification-Driven Code Generation with LLMs: An Empirical Study Design. arXiv:2601.03878, 2026
arXiv 2026
-
[9]
LIFEBENCH: 16 Evaluating Length Instruction Following in Large Language Models
Wei Zhang, Zhenhong Zhou, Kun Wang, Junfeng Fang, Rongwu Xu, Yuanhe Zhang, Rui Wang, Ge Zhang, Xinfeng Li, Li Sun, Lingjuan Lyu, Yang Liu, and Sen Su. LIFEBENCH: 16 Evaluating Length Instruction Following in Large Language Models. InAdvances in Neural Information Processing Systems 38 (NeurIPS 2025), Datasets and Benchmarks Track, 2025. arXiv:2505.16234
Pith/arXiv arXiv 2025
-
[10]
Bowman, and Sara Price
Isha Gupta, Kai Fronsdal, Abhay Sheshadri, Jonathan Michala, Jacqueline Tay, Rowan Wang, Samuel R. Bowman, and Sara Price. Bloom: An Open Source Tool for Automated Behavioral Evaluations. Anthropic, December 2025.https://www.anthropic.com/research/bloom ; code:https://github.com/safety-research/bloom
2025
-
[11]
Cambridge University Press, 2007
Andrew Gelman and Jennifer Hill.Data Analysis Using Regression and Multilevel/Hierarchical Models. Cambridge University Press, 2007
2007
-
[12]
Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta- Analyses.Social Psychological and Personality Science, 8(4):355–362, 2017
Daniël Lakens. Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta- Analyses.Social Psychological and Personality Science, 8(4):355–362, 2017
2017
-
[13]
Nosek and Daniël Lakens
Brian A. Nosek and Daniël Lakens. Registered Reports: A Method to Increase the Credibility of Published Results.Social Psychology, 45(3):137–141, 2014
2014
-
[14]
Simmons, Leif D
Joseph P. Simmons, Leif D. Nelson, and Uri Simonsohn. False-Positive Psychology: Undis- closed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant. Psychological Science, 22(11):1359–1366, 2011. 17
2011
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.