Pith. sign in

REVIEW 3 major objections 4 minor 14 references

Under a controlled pre-registered test, a non-programmer persona made one frontier model decline to code in 12 of 60 responses while another model ignored the persona entirely.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 17:59 UTC pith:MLWUOLZC

load-bearing objection A careful, pre-registered study with a genuinely striking post-hoc observation about a librarian persona making Claude refuse to code — but the confirmatory evidence for model dependence is fragile, and the paper already half-admits it. the 3 major comments →

arxiv 2607.17420 v1 pith:MLWUOLZC submitted 2026-07-19 cs.CL cs.SE

The Librarian Who Refused to Code: Model-Dependent Identity Enactment in LLM Code Generation

classification cs.CL cs.SE
keywords persona promptingLLM code generationmodel dependenceidentity enactmentbehavioral policypre-registrationrefusal behaviorsystem prompts
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tests whether detailed biographical personas—system prompts that describe a person without offering any instructions about how to write code—change how large language models generate code. Using a fully crossed design of 4 personae, 12 coding tasks, 2 frontier models, and 5 runs per cell, it finds that persona effects are model-dependent: one model (Claude Opus) responded to a 'research librarian' persona with in-character disclaimers in 55 of 60 responses and genuine refusals to produce code in 12 of 60, lowering its correctness from 0.92 to 0.67, while the other model (GPT-5.5) showed neither behavior. Engineer personas shifted output length and structural consistency without improving correctness, leading the authors to reframe personas as behavioral-policy biases that fit some objectives better than others. The paper argues that the useful question is not which persona makes a model better but which behaviors a persona activates or suppresses—and that susceptibility to persona-driven behavior change itself varies across model families.

Core claim

The central claim is that biographical personas are not neutral context or universal quality interventions; they operate as model-dependent behavioral-policy biases. On Claude Opus, a persona with zero behavioral instruction—a research librarian named Linnea—caused the model to generalize from identity to action: it hedged with role-consistent disclaimers in 55 of 60 responses and declined to write any code in 12 of 60, a behavior stated nowhere in the prompt. On GPT-5.5 the same persona produced zero disclaimers and zero genuine refusals. The pre-registered condition-by-model interaction is significant for output length (on both provider-reported tokens and a shared visible-character measur

What carries the argument

The load-bearing objects are the biographically detailed persona texts—especially the 'Linnea' research librarian, which contains no code-, length-, or behavior-related instruction—and the pre-registered mixed-effects analysis with a condition×model interaction term. The identity-enactment phenomenon is quantified through fixed, post-hoc operationalizations: an in-character disclaimer is a first-person librarian/non-programmer assertion near the start of the response, and a genuine no-code response is one with zero fenced code blocks that terminated normally (not via length truncation). These textual definitions, combined with the sharp contrast between the two models, carry the argument tha

Load-bearing premise

The load-bearing premise is that the pre-registered condition-by-model interaction is a robust signal of model dependence—a premise that the paper's own task-clustered sensitivity analysis shows to be fragile, since the provider-token interaction loses significance when one prompt-contaminated task is excluded and the strongest surviving interaction uses a measure elevated to primary only post hoc.

What would settle it

A pre-registered replication that removes the R-TS-2 prompt leak, fixes the visible-character interaction as the primary outcome, and applies a task-clustered test; if the condition×model interaction on that measure fails to meet the pre-registered threshold, the confirmatory basis for model dependence would be falsified. Alternatively, running the same Linnea persona on a third frontier model family and observing disclaimers or refusals at rates comparable to Claude Opus would contradict the claimed model-family specificity.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Personas can change task willingness, not just style: a benign, instruction-free identity can cause a model to hedge or refuse to code, so persona design should be treated as a soft policy override.
  • No persona tested improved correctness; on tasks near saturation, persona selection is a matter of fit between activated behaviors and the objective, not a universal quality lever.
  • Model families differ in persona susceptibility; the same persona-bearing prompt can be inert on one model and disruptive on another, raising silent-degradation risks in multi-model pipelines.
  • The engineer-persona length effects are confounded with an explicit style sentence, so narrative-only causation is established only for the librarian persona and only on one model family.
  • The signed pre-registration with hard execution gates offers a reusable methodology for behavioral LLM experiments, independent of the persona results themselves.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The identity-enactment result suggests a security-relevant vector: a persona injected through an untrusted channel (e.g., a simulated teammate in a shared agent workspace) might induce silent refusal or hedging without any malicious instruction; the paper flags this as a hypothesis, and I would add that an adversary could craft personas to exploit specific model-family susceptibilities.
  • If model dependence generalizes, persona evaluation on a single model family has little predictive value for another family; benchmark designers should report model×condition interactions and cluster-robust tests, not only pooled persona effects.
  • The paper's own task-clustered sensitivity analysis shows the confirmatory interaction on the pre-registered provider-token measure is fragile (narrowly above threshold, and not significant when one prompt-contaminated task is removed), implying the model-dependence claim as confirmatory rests largely on a post-hoc visible-character measure; a replication with a pre-registered cluster-robust desig
  • Because baseline correctness is saturated, the hypothesis that personas can improve quality on harder tasks remains completely open; a natural extension is to run the same design on tasks with headroom to see whether any persona raises correctness above baseline.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper reports a pre-registered experiment with 480 completions crossing four prompt conditions (no persona, two engineer personas, a research-librarian persona), 12 code-generation tasks, two frontier models, and five runs per cell. Its central claims are: (i) a pre-registered condition×model interaction shows that persona effects are model-dependent; (ii) personas act as behavioral-policy biases rather than universal quality levers; and (iii) an exploratory finding that the librarian persona induces in-character disclaimers and no-code refusals on Claude Opus (55/60 disclaimers, 12/60 no-code) but not on GPT-5.5. The paper transparently marks confirmatory versus exploratory analyses, releases raw logs, derived scores, a classification script, and a pre-registration, but not the executable task harness.

Significance. If the identity-enactment result replicates under a pre-registered operationalization, it is significant: it would show that an instruction-free biographical persona can change not only style but task willingness, and do so differentially across model families. The methodological apparatus—externally signed pre-registration, execution gates, amendment logging, and partial artifact release—is a genuine contribution to behavioral LLM evaluation. However, the confirmatory support for the headline model-dependence claim is fragile: under the paper's own task-level clustered sensitivity analysis, the pre-registered provider-token interaction does not survive exclusion of one contaminated task, and the only interaction that remains strong under clustering uses a measure elevated to primary post hoc. The paper is mostly honest about these limitations, but the abstract and introduction overstate the confirmatory status of the model-dependence contribution.

major comments (3)
  1. [§4.2, Appendix C.2] The central claim that persona effects are model-dependent rests on H_INT, the condition×model interaction in the pre-registered mixed-effects model (outcome ~ C(condition)×C(model) + (1|task)). Appendix C.2 reports the exploratory task-clustered sensitivity: the correctness interaction does not survive (F(3,11)=2.47, p=0.116); the provider-token interaction passes only narrowly (F(3,11)=6.80, p=0.0074) and rises to p=0.0185 when the R-TS-2 task is excluded—above the pre-registered Bonferroni threshold. Because a random-intercept model with only 12 tasks can substantially understate between-task variability, and because the provider-token result is borderline and vanishes with one excluded task, the pre-registered analysis does not provide robust confirmatory support for model dependence. The paper should either adopt task-clustered inference as a confirmatory test or explicitly downgrad
  2. [§3.4, §4.2, §6] The pre-registration specified provider-reported completion tokens as E1; the paper elevated visible characters to primary “post hoc, for cross-provider comparability” (§3.4). This creates a circularity risk: the only interaction that survives task-level clustering is on visible characters (p=8.4e-6), a measure selected after seeing the data. The abstract honestly calls it post hoc, but the paper's conclusion and contribution list call the model-dependence result a “pre-registered demonstration.” The robust evidence is post hoc, and the language in the abstract and conclusion should be adjusted so that the confirmatory status of H_INT matches the analysis that actually supports it.
  3. [§3.2, Appendix B, Data availability] The executable task harness is not released; end-to-end re-execution of test-based scoring from raw prompts is impossible. Since correctness (Q1) is a central outcome—driving the identity-enactment no-code zeros and the disclaimer-correlation analysis—an independent reader cannot verify the reported correctness values. The release of derived scores is useful, but the claim of a “partial-reproducibility package” would be strengthened by releasing the harness or, failing that, by providing a detailed independent evaluation path for the correctness metric.
minor comments (4)
  1. [§1] Typo: “noimperativeinstructions” should be “no imperative instructions.”
  2. [§3.4] The notation for E1 visibility could be clarified: the paper uses “provider tokens” and “visible characters” but at times refers to “E1” without specifying which measure; please be consistent.
  3. [§6] The discussion of the R-TS-2 prompt leak is clear, but it would help to state explicitly that the leak cannot affect between-persona contrasts because the prompt is identical across conditions; currently this is stated only in Appendix C.2.
  4. [Appendix F.2] Table 3 would benefit from a note that the GPT-5.5 I-TS-1 cell has n=4 due to the excluded truncation artifact; this is stated in F.2 but not in the table caption.

Circularity Check

0 steps flagged

No significant circularity: the central claims are empirical, pre-registered where confirmatory, and post hoc measures are explicitly labeled as exploratory.

full rationale

The paper makes no derivation-from-first-principles claim that reduces to its inputs. Its central claim—model-dependent persona effects—is tested empirically via a pre-registered condition×model interaction on observed completions, not computed from a fitted parameter or imported from the authors' own prior work. The pre-registered mixed-effects analysis and the exploratory task-clustered sensitivity analysis are both reported; the fragility of the provider-token interaction under clustering and after excluding R-TS-2 is a statistical-inference and robustness concern, not a circular reduction. The identity-enactment result was operationalized post hoc after inspecting outputs, but the paper repeatedly and explicitly labels it exploratory/descriptive (§6, C.3, F), and the no-code outcome is an independently defined behavioral measure rather than a renamed input. There are no load-bearing self-citations, no imported uniqueness theorems, and no fitted parameters renamed as predictions. The post hoc elevation of visible characters to primary is a transparency/validity limitation, not a circular step. Therefore, on the evidence in the manuscript, no circular step meets the required evidentiary bar.

Axiom & Free-Parameter Ledger

1 free parameters · 5 axioms · 0 invented entities

The central claims rest on standard statistical assumptions plus several domain assumptions: test validity, comparability of length measures, persona-text purity, prompt-leak neutrality, and model assumptions of the LMM. The only hand-chosen free parameter is the post hoc identity-enactment classifier, which directly shapes the exploratory counts. No new physical or conceptual entities are introduced.

free parameters (1)
  • Identity-enactment classifier thresholds (proximity ≤40 chars; phrase list including 'wrong Linnea', 'not really my area = Post hoc; yields 55/60 disclaimers and 12/60 no-code responses on Opus
    These thresholds were fixed after inspecting the completions and determine the counts that drive the identity-enactment claim; they are a hand-chosen operationalization, not a pre-registered metric.
axioms (5)
  • domain assumption Test-based correctness is a valid measure of code quality
    The study uses automated test suites as ground truth; if tests are flawed, correctness scores misrepresent quality.
  • domain assumption The two models' output lengths are comparable via visible character count
    Provider tokenizers differ; visible characters are assumed to be a shared cross-provider measure (elevated post hoc).
  • domain assumption The Linnea persona text contains no behavioral instruction
    Verified by a whole-word scan, but a scan cannot rule out subtle semantic cues; the identity-enactment claim rests on this.
  • domain assumption The R-TS-2 user prompt leak does not bias between-persona contrasts
    All conditions received the same prompt, so between-persona contrasts are assumed unbiased, but the leak may interact with persona-induced task willingness.
  • standard math Mixed-effects model assumptions (random task intercept, independent errors)
    Used for the pre-registered analysis; the paper's own task-clustered sensitivity shows the interaction may not survive these assumptions.

pith-pipeline@v1.3.0-alltime-deepseek · 13600 in / 12203 out tokens · 102953 ms · 2026-08-01T17:59:07.014996+00:00 · methodology

0 comments
read the original abstract

Biographical personas are widely used in system prompts, but their effects on code generation are rarely evaluated under controlled, pre-registered conditions. We tested four prompt conditions (no persona, two engineer personas, and a research-librarian persona), 12 code-generation tasks, two frontier models, and five runs per cell (480 completions). Persona effects differed between the two tested models. Under the pre-registered mixed-effects analysis, the condition-by-model interaction was significant for provider-reported output tokens; a post-hoc visible-character measure showed the same qualitative pattern. Six GPT-5.5 completions were length-capped and are reported separately. On Claude Opus, the minimalist engineer persona reduced visible output by 30% (33% in provider tokens) without improving correctness, while the thorough engineer persona increased output without a correctness gain. In an exploratory post-hoc analysis, the librarian persona elicited in-character disclaimers in 55 of 60 Opus responses and 12 genuine no-code responses, lowering mean correctness from 0.92 to 0.67. GPT-5.5 produced neither behavior in its 59 non-truncated responses. These results are consistent with personas acting as Model-Dependent behavioral-policy biases rather than universal quality interventions. We release raw completions, derived scores, analysis artifacts, a pre-registration document, and an execution gate log; end-to-end test-based rescoring requires an unreleased task harness.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

14 extracted references · 2 canonical work pages

  1. [1]

    Principled Per- sonas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance

    Pedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy, and Benjamin Roth. Principled Per- sonas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 26857–26886, 2025. doi:10.18653/v1/2025.emnlp-main.1364. arXiv:2508.19764

  2. [2]

    The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models

    Marlene Lutz, Indira Sen, Georg Ahnert, Elisa Rogers, and Markus Strohmaier. The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 23212–23237, 2025. doi:10.18653/v1/2025.findings-emnlp.1261. arXiv:2507.16076

  3. [3]

    Persona Prompting as a Lens on LLM Social Reasoning

    Jing Yang, Moritz Hechtbauer, Elisabeth Khalilov, Evelyn Luise Brinkmann, Vera Schmitt, and Nils Feldhus. Persona Prompting as a Lens on LLM Social Reasoning. InProceedings of the 2026 Conference of the European Chapter of the Association for Computational Linguistics (EACL), pages 1152–1170, 2026. doi:10.18653/v1/2026.eacl-long.52

  4. [4]

    SPeCtrum: A Grounded FrameworkforMultidimensionalIdentityRepresentationinLLM-BasedAgent

    Keyeun Lee, Seo Hyeong Kim, Seolhee Lee, Jinsu Eun, Yena Ko, Hayeon Jeon, Esther Hehsun Kim, Seonghye Cho, Soeun Yang, Eun-mee Kim, and Hajin Lim. SPeCtrum: A Grounded FrameworkforMultidimensionalIdentityRepresentationinLLM-BasedAgent. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguist...

  5. [5]

    Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM

    Zizhao Hu, Mohammad Rostami, and Jesse Thomason. Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM. arXiv:2603.18507, 2026

  6. [6]

    A Survey on Large Language Models for Code Generation

    Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim. A Survey on Large Language Models for Code Generation. arXiv:2406.00515, 2024

  7. [7]

    Sarah Fakhoury, Aaditya Naik, Georgios Sakkas, Saikat Chakraborty, and Shuvendu K. Lahiri. LLM-Based Test-Driven Interactive Code Generation: User Study and Empirical Evaluation. IEEE Transactions on Software Engineering, 50:2254–2268, 2024. arXiv:2404.10100

  8. [8]

    González-Barahona

    Giovanni Rosa, David Moreno-Lumbreras, Gregorio Robles, and Jesús M. González-Barahona. Understanding Specification-Driven Code Generation with LLMs: An Empirical Study Design. arXiv:2601.03878, 2026

  9. [9]

    LIFEBENCH: 16 Evaluating Length Instruction Following in Large Language Models

    Wei Zhang, Zhenhong Zhou, Kun Wang, Junfeng Fang, Rongwu Xu, Yuanhe Zhang, Rui Wang, Ge Zhang, Xinfeng Li, Li Sun, Lingjuan Lyu, Yang Liu, and Sen Su. LIFEBENCH: 16 Evaluating Length Instruction Following in Large Language Models. InAdvances in Neural Information Processing Systems 38 (NeurIPS 2025), Datasets and Benchmarks Track, 2025. arXiv:2505.16234

  10. [10]

    Bowman, and Sara Price

    Isha Gupta, Kai Fronsdal, Abhay Sheshadri, Jonathan Michala, Jacqueline Tay, Rowan Wang, Samuel R. Bowman, and Sara Price. Bloom: An Open Source Tool for Automated Behavioral Evaluations. Anthropic, December 2025.https://www.anthropic.com/research/bloom ; code:https://github.com/safety-research/bloom

  11. [11]

    Cambridge University Press, 2007

    Andrew Gelman and Jennifer Hill.Data Analysis Using Regression and Multilevel/Hierarchical Models. Cambridge University Press, 2007

  12. [12]

    Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta- Analyses.Social Psychological and Personality Science, 8(4):355–362, 2017

    Daniël Lakens. Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta- Analyses.Social Psychological and Personality Science, 8(4):355–362, 2017

  13. [13]

    Nosek and Daniël Lakens

    Brian A. Nosek and Daniël Lakens. Registered Reports: A Method to Increase the Credibility of Published Results.Social Psychology, 45(3):137–141, 2014

  14. [14]

    Simmons, Leif D

    Joseph P. Simmons, Leif D. Nelson, and Uri Simonsohn. False-Positive Psychology: Undis- closed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant. Psychological Science, 22(11):1359–1366, 2011. 17