{"id":"bf493805-eedd-4b8f-9230-3eaad707df4c","arxiv_id":"2607.18566","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Task narrative, not persona, is the dominant driver of LLM agent action profiles in structurally identical investigation games, and transferable personas are those with concrete action words.","lead":"This paper shows that the story framing of a task—medical, IT, or crime—shapes how LLM agents act more strongly than the persona you assign them. It also identifies which persona descriptions transfer across stories and offers a method to pick more portable personas.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Structural isomorphism is asserted, not tested: resource counts and role labels vary across narratives, so the 5–31× narrative effect may be confounded with resource availability.","rationale":"The reader's weakest_assumption identifies exactly this point: the three games are claimed to be functionally equivalent despite differing character counts, document counts, and role labels. My analysis sharpens the concern by connecting it to the specific dependent measures: raw action-type ratios are not normalized for the number of available resources, so resource-count differences can mechanically shift the very features that drive the 5–31× variance claim. The paper's own Appendix A provides the counts, and §6 robustness checks do not include a matched-resource control. This is the most load-bearing concern because every downstream result—narrative priors, persona transfer, behavioral anchors, the CK prediction, and the persona-selection method—depends on the clean isolation of narrative from task structure. If the isomorphism assumption fails, the central claim is at least partially reinterpreted as a resource-availability or role-label effect. This does not require rejection: the paper is otherwise rigorous, with multiple models, robustness checks, and a pre-registered prediction. But the missing control is exactly the kind of condition that a CONDITIONAL verdict should demand. Since the reader already reached CONDITIONAL, my recommendation is unchanged. Agreement with the reader is 'agree' because the same load-bearing assumption was identified, though I have tied it more explicitly to the raw ratio features and the role-label confound.","tokens_in":14808,"tokens_out":4479,"duration_ms":59599,"concrete_test":"Run a matched-resource variant: set all three games to identical counts (7 characters, 15 documents, 6 locations) and an identical role label (e.g., 'investigator') while retaining narrative-specific wording. If the cross-narrative action-profile pattern (SD highest read, CI highest talk, MM highest test) and the narrative η² ratios of §4.1 persist, the structural-isomorphism assumption is supported. If the effect sizes shrink or reorder, resource counts or role labels are confounds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central identification strategy is structural isomorphism (§3.1): the three games are claimed to differ only in task narrative, yet Appendix A shows they differ in character counts (CI 7, SD 5, MM 7), document counts (CI 12, SD 12, MM 15), and role labels (student/engineer/detective). The text asserts these differences are 'narrative needs rather than structural asymmetry' and that resources are 'functionally equivalent,' but no empirical check supports this. This matters because the headline variance decomposition in §4.1 uses raw action ratios (e.g., talk ratio = talk / (read+talk+test+move)). A policy that interacts with available characters or documents will mechanically produce different action ratios when the number of available sources differs, even with no narrative effect. Thus the 5–31× narrative η² could partly reflect resource-count differences, not narrative priors. Role-label changes are also part of the 'narrative' manipulation; the abstract-code experiment in §5 removes action-verb semantics but does not hold role labels or resource counts constant, so it does not resolve this confound. The held-out CK prediction (§5) similarly changes both resource configuration and role, so it cannot validate the isomorphism assumption. This is not an internal contradiction, but an unvalidated identification condition on which the paper's central claim rests.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether the narrative framing of a task can shape LLM agent behavior more strongly than an assigned persona. It introduces three text-based investigation games (disease, IT, murder mystery) claimed to be structurally isomorphic, differing only in narrative surface. Using 1,890 sessions across 3 models and 10 personas, the authors report that narrative explains 5–31× more behavioral variance than persona, and that cross-narrative persona transfer is governed by 'behavioral anchors'—concrete action language in persona descriptions. They validate this with a held-out fourth game, a causal rewriting intervention, and extensive robustness checks, and propose a persona-selection method based on the Behavioral Consistency Index.","tokens_in":15132,"tokens_out":6749,"duration_ms":81755,"significance":"If the identification strategy is valid, the paper makes a substantial contribution: it provides a reusable structural-isomorphism framework, demonstrates large narrative effects on agent behavior, and offers an actionable design principle for persona prompting. The empirical scope is a clear strength: 1,890 sessions, three model families, pre-registered held-out prediction, multiple robustness analyses (Kruskal-Wallis, binomial GLM, leave-one-narrative-out), and an intervention with bootstrap/permutation inference. The practical persona-selection algorithm is useful and the paper is clearly written. The central risk is whether the 'varying only task narrative' claim is actually supported given differences in resource counts and role labels across the games; this concern bears directly on the headline variance decomposition.","major_comments":[{"comment":"The paper asserts structural isomorphism, but the games differ in character counts (CI 7, SD 5, MM 7) and document counts (CI 12, SD 12, MM 15). The central variance decomposition (§4.1) uses action ratios defined in §3.4 as ratio_a = n_a/(n_read+n_talk+n_test+n_move). If agents tend to interact with each available source once, talk actions scale with the number of characters and read actions with the number of documents, producing different ratios even under identical decision policies. The statement in §3.1 that these resources are 'functionally equivalent' is not tested. The §6 robustness with 12 non-ratio features does not remove the confound because features such as 'social breadth' and 'resource coverage' normalize by the same available counts. Please either equalize counts across games or provide an empirical invariance check showing action ratios are unchanged when resource count","section":"§3.1, Appendix A (Table 4)"},{"comment":"The Cooking Kitchen (CK) experiment is used to show narrative priors generalize, but CK changes both the role label (kitchen inspector) and, presumably, resource configuration. Appendix I does not report CK character/document counts. The observed talk-ratio hierarchy (Table 6) is therefore compatible with a resource-availability explanation (e.g., more talk sources in CI/CK than SD). Without CK structural properties and a demonstration that the prediction is not driven by count differences, the CK confirmation does not independently validate structural isomorphism.","section":"§5, Appendix I"},{"comment":"The causal claim for behavioral anchors compares the original 'social collaborative' persona with an abstract rewrite. The rewrite differs not only in the absence of action words but also in register and lexical frequency ('collaborative epistemic exchange' vs. 'conversation'). The observed BCI drop could reflect prompt unnaturalness rather than the missing anchors. A matched control replacing anchor words with equally concrete, equally frequent non-action words (or template variants) is needed. The replication on 'coverage focused' helps but uses a different rewrite style and does not isolate the anchor mechanism.","section":"§5, Tables 2–3, Appendix E"},{"comment":"The claim that narrative priors are not strategic adaptations relies on correlations between the narrative-biased action and task success within narratives. Because the narrative prior is a between-narrative mean shift, a within-narrative correlation can be driven largely by persona variation. To support the claim that the narrative-induced shift itself hurts performance, the paper should either compare success under a counterfactual manipulation of the action bias, compute partial correlations controlling for persona, or compare success across narratives while holding persona fixed.","section":"§4.1, Table 9"}],"minor_comments":[{"comment":"The ANOVA setup states '70 observations per model–narrative cell; N = 630 per model,' but 3 cells × 70 = 210, not 630. Please correct the degrees of freedom or clarify whether trials are included.","section":"§4.1"},{"comment":"'Any behavioral differences across environments must stem not from different action inventories or investigation scaffolds, but from how the same scaffold is interpreted under different narrative frames' is too strong given the admitted resource-count differences. Please soften this and explicitly tie it to the invariance check requested above.","section":"§3.1"},{"comment":"The caption says 'identical decision structure (4 stages, 4 action types, 7 scenarios)' even though Appendix A reports different character and document counts. Specify that action inventories and stage progression are identical, while resource counts are matched but not exactly equal.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is empirically rich and well within the scope of COLM, but the structural-isomorphism assumption is the crux. I recommend requiring either an invariance analysis for resource counts or equalized game configurations before acceptance. The current wording overstates the 'varying only narrative' claim. No concerns about novelty or scholarship; the framing issues are fixable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper deserves a serious referee. It makes a genuinely new contribution: it treats task narrative as an independent variable in interactive LLM agents, constructs three structurally aligned investigation games, and shows narrative explains far more behavioral variance than assigned persona. The behavioral-anchor finding (concrete action words make personas transfer) is concrete and the causal intervention is compelling. The held-out narrative prediction is a nice touch.\n\nThe empirical core is extensive: 1,890 sessions, three models, ten personas, multiple robustness checks (Kruskal-Wallis, binomial GLM, bootstrap, permutation). Credit where it's due: the design is careful and the claims are mostly calibrated.\n\nThe main soft spot is the structural isomorphism assumption. The three games differ in character counts (5–7), document counts (12–15), and role labels. The paper asserts these are 'functionally equivalent' because each source is single-access, but it doesn't test that equivalence. If more characters or documents naturally elicit more talk or read actions, the action-ratio variance decomposition (5–31×) would partly reflect resource availability rather than narrative priors. The abstract-code experiment doesn't solve this, because role labels and resource counts still vary. This doesn't sink the paper—narrative classification still works on non-ratio features, and the held-out prediction is consistent—but it means the headline magnitudes are probably overestimated, and the 'varying only task narrative' framing is not yet established.\n\nMinor issues: no code/data release, and some stats (e.g., r=.96) lack supporting detail. These are fixable.\n\nVerdict: conditionally accept for peer review. The central idea is important, the experiments are extensive, and the main confound can be addressed with a follow-up (e.g., holding resource counts constant across narratives, or varying resource counts within a narrative). Worth reading for anyone working on prompt sensitivity or LLM agent reliability.","headline":"A strong empirical study of narrative framing in LLM agents, with a real confound in the isomorphism design that needs testing.","tokens_in":15557,"tokens_out":2517,"would_cite":true,"duration_ms":28077,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The story framing of a task shapes LLM agent behavior more than the assigned persona; only action-grounded personas transfer across narratives.","keywords":["narrative priors","persona prompting","LLM agents","structural isomorphism","behavioral consistency index","behavioral anchors","cross-narrative transfer","prompt framing effects"],"falsifier":"Swap the resource counts between two narratives (e.g., give the disease game 5 characters and 15 documents, the IT game 7 characters and 12 documents) while keeping each story's wording intact; if the read/talk/test ratios shift with the counts, the isomorphism premise fails and the narrative-prior attribution collapses.","tokens_in":14715,"feed_emoji":"🎭","tokens_out":7178,"duration_ms":79907,"temperature":0.7,"pith_summary":"This paper tries to show that the narrative wrapper of a task—not the persona description a user assigns—is the main driver of what an LLM agent does. It does this by building three text-based investigation games with the same underlying decision structure but different stories (disease, IT, murder mystery) and measuring behavior over 1,890 sessions. Narrative explains 5–31× more variance than persona, the biases are consistent across different model architectures, and in two of the three games the narrative-biased action is tied to worse task success. The paper also identifies why some personas do transfer across stories: concrete action words in the persona description that map directly onto the game's actions. A sympathetic reader would care because it challenges the common practice of persona prompting as a portable control interface for LLM agents.","feed_headline":"Narrative framing beats persona in steering LLM agents","feed_subtitle":"Three identical-in-structure games show the story framing, not the assigned persona, sets the agent's actions.","key_machinery":"The load-bearing object is the structural-isomorphism design: three games with identical action spaces, stage sequences, and resource constraints that differ only in narrative surface. The measurement is the Behavioral Consistency Index (BCI), a pairwise correlation of z-normalized behavioral profiles that tells whether a persona keeps the same signature across narratives. The explanatory mechanism is the behavioral anchor—a concrete action word in a persona description that maps directly onto the game's action types. Together, isomorphism attributes observed differences to narrative, BCI quantifies transfer, and anchor manipulation provides the causal test.","core_discovery":"On the paper's own terms, the central discovery is that task narrative is not a confound but the primary driver of LLM agent behavior. In three text-based investigation games built to be structurally identical—same four action types, same four-stage progression, same single-access resource constraints—the only meaningful variation was the surface story. Narrative explained 5–31× more variance in the read/talk/test action ratios than the assigned persona, and these biases were consistent across three models. The biases are not strategic: in the disease game the talk-heavy bias and in the murder game the test-heavy bias are negatively associated with success. Personas still matter within a fix","pith_inferences":["Editorial extension: if this account generalizes, benchmark designers should treat narrative framing as a controlled variable; two benchmarks measuring the 'same' agent skill but using different story wrappers may be measuring narrative priors rather than skill.","Editorial extension: a direct test of the 'pretraining stereotype' mechanism would compare the observed action biases (talk in medical, read in IT, test in crime) against co-occurrence statistics in a large text corpus; the paper does not quantify corpus evidence.","Editorial extension: the anchor-removal result implies a practical prompt-editing rule—prefer concrete action verbs in personas—but the paper's own cross-model instability of BCI rankings (correlation below 0.24) suggests the rule will need per-model tuning in deployment.","Editorial extension: the structural-isomorphism design could be ported to non-investigation task families, such as planning or tool use; the paper lists this as a limitation, and confirming narrative priors there would strengthen the claim that story framing is a general driver of LLM behavior."],"forward_implications":["Persona effects are largely confined to the narrative they were tested in: cross-narrative transfer classification is 10.4%, barely above the 10% chance level.","Narrative priors are consistent across three models of different architecture, suggesting they are a general property of large-scale pretrained models rather than a quirk of one model.","In two of three narratives, the action tendency the story elicits (talking in the disease game, testing in the murder game) is negatively associated with success, so narrative framing is not a neutral stylistic layer.","Personas described with concrete action words transfer across narratives; personas described with abstract traits do not, and removing the concrete words cuts transfer consistency by 95%.","Selecting personas by their measured cross-narrative consistency (BCI) improves transfer identifiability in all tested conditions without needing target-narrative data."],"fun_headline_variants":["Story beats persona: narrative drives LLM agent actions 31x more","Narrative priors, not personas, steer LLM agents' behavior","Task narrative dominates persona in LLM agent behavior","Why story framing outranks persona for LLM agents","LLM agents follow the story, not the persona, 31x stronger"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper's central premise—stated in §3.1 and Appendix A as 'functionally equivalent'—is that the three games are structurally identical despite character counts of 5–7 and document counts of 12–15; if those resource counts alter the optimal action mix, the behavioral differences attributed to narrative could instead be resource or role-label effects.","fun_headline_variants_meta":{"raw":{"variants":["Story beats persona: narrative drives LLM agent actions 31x more","Narrative priors, not personas, steer LLM agents' behavior","Task narrative dominates persona in LLM agent behavior","Why story framing outranks persona for LLM agents","LLM agents follow the story, not the persona, 31x stronger"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000556,"raw_usage":{"total_tokens":2475,"prompt_tokens":725,"completion_tokens":1750,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":1661}},"tokens_in":469,"tokens_out":1750,"duration_ms":18989,"temperature":1.0,"reasoning_tokens":1661,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T14:59:13.189102+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Swap the resource counts between two narratives (e.g., give the disease game 5 characters and 15 documents, the IT game 7 characters and 12 documents) while keeping each story's wording intact; if the read/talk/test ratios shift with the counts, the isomorphism premise fails and the narrative-prior attribution collapses.","supporting_citations":[],"review_version":1}