{"id":"09290b20-d7b0-4b8c-af30-b90261fdfa76","arxiv_id":"2605.30207","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Persona prefixes reduce brand recommendation Jaccard similarity by 0.12-0.20, with mid-market brands swapping up to 75% of recommendations while category leaders remain ~80% consistent across OpenAI and Anthropic models.","lead":"This paper audits how buyer personas added to prompts change which brands AI models recommend for CRM software queries. A smart generalist should read it to see why generic prompts may not capture real variation in commercial AI outputs and why persona conditioning matters for measurement.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Persona prefixes may introduce prompt-phrasing confounds (length, keywords, structure) independent of buyer-identity interpretation.","rationale":"The reader's weakest_assumption directly isolates the same point. Because the full text is referenced but the abstract supplies the only quantitative claims, the absence of any reported control for prefix surface features remains the single most load-bearing untested assumption for the central causal attribution. This does not invalidate the audit design but conditions the interpretation of the Delta on that assumption holding.","tokens_in":1877,"tokens_out":395,"duration_ms":19859,"concrete_test":"Re-run the 2,000-query design using length-matched and keyword-controlled neutral prefixes (e.g., generic descriptors of equal token count with no buyer-specific nouns) in place of the original personas; recompute all Jaccard deltas and prominence-stratified tables. If the effect size remains within 0.05 of the reported values, the identity interpretation is supported; if it attenuates by >50%, phrasing confounds are material.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline result attributes the observed Jaccard drop (-0.12 to -0.20) to models conditioning on distinct buyer contexts supplied by the 10 persona prefixes. For this attribution to be secure, the prefixes must function solely as identity signals; any systematic differences in token length, lexical content, or syntactic framing could alter retrieval scores or generation priors through mechanisms unrelated to the intended persona. The abstract provides no indication that prefix length, keyword overlap with the query, or surface-form variation were matched or ablated across the 10 personas. The design (10 personas × 8 prompts × 3 configs) therefore leaves open the possibility that the measured Delta and the prominence-stratified pattern partly reflect prompt-engineering artifacts rather than context integration per se. The clustered CIs and the sonnet-cell coverage note do not address this source of variation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper audits how buyer personas affect brand recommendations in retrieval-augmented commercial chat systems. It samples 2000 runs across 10 personas × 8 prompts × 3 model configurations (OpenAI high/low and Anthropic sonnet-4.6/low, with partial coverage in one cell) and reports that persona prefixes reduce recommendation-set Jaccard similarity by Δ = -0.12 to -0.20 relative to same-persona baselines (clustered 95% CIs exclude zero). The effect is prominence-stratified (category leaders ~80% consistent; mid-market brands swap up to 75%), larger in the Anthropic configuration, and attributed to differences in retrieval attribution rates.","tokens_in":2077,"tokens_out":630,"duration_ms":22733,"significance":"If the central empirical deltas hold after controls for prompt artifacts, the result shows that measurements of AI brand perception must condition on buyer persona, as aggregated protocols obscure material variation concentrated at mid-market brands. The prominence stratification and cross-provider comparison (with note on retrieval-unattributed generation) provide a concrete, falsifiable demonstration that context integration strength modulates recommendation stability. The clustered-CI design and explicit coverage limitations are strengths that improve audit transparency.","major_comments":[{"comment":"The design (10 personas × 8 prompts) does not report controls or ablations for systematic differences in prefix length, lexical overlap with the query, or syntactic framing across the 10 personas. Without such matching, the observed Jaccard drops cannot be securely attributed to buyer-identity conditioning rather than prompt-phrasing confounds that could alter retrieval scores or generation priors independently of the intended persona signal.","section":"Abstract (design space description and effect attribution)"},{"comment":"The interpretation that the Anthropic vs. OpenAI asymmetry is consistent with retrieval-unattributed generation (43-52% vs. 8-29%) rests on rates documented in Jack 2026. While the empirical deltas are independent measurements, the explanatory claim for why the effect is larger on one route reduces in part to that prior result; direct within-study attribution measurements or clearer separation of the descriptive claim from the causal interpretation would strengthen the argument.","section":"Abstract (asymmetry paragraph)"}],"minor_comments":[{"comment":"The sonnet cell's CI is noted as resting on only 4 prompt clusters; a table or appendix explicitly listing per-cell prompt coverage and cluster counts would improve reproducibility.","section":"Abstract"},{"comment":"The abstract states N=10 reps but does not specify whether the clustered CIs account for prompt-level or persona-level clustering; a brief methods note on the clustering structure would clarify the statistical procedure.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The explanatory reliance on the Jack 2026 self-citation for the model-asymmetry interpretation raises a minor novelty/citation-pattern concern; the core empirical contribution stands independently but the interpretive framing could be tightened."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments on attribution and design controls. We address each point below and indicate planned revisions.","responses":[{"response":"We agree the manuscript does not report quantitative controls or ablations for prefix length, lexical overlap, or syntactic framing. Personas were constructed to vary primarily on buyer identity with fixed core queries, but without explicit matching this leaves room for prompt artifacts. In the revised version we will add (i) summary statistics on prefix lengths and token overlap across the 10 personas and (ii) a sensitivity check re-running a subset of prompts with length-normalized prefixes to test robustness of the reported Jaccard deltas.","revision_made":"yes","referee_comment":"[Abstract (design space description and effect attribution)] The design (10 personas × 8 prompts) does not report controls or ablations for systematic differences in prefix length, lexical overlap with the query, or syntactic framing across the 10 personas. Without such matching, the observed Jaccard drops cannot be securely attributed to buyer-identity conditioning rather than prompt-phrasing confounds that could alter retrieval scores or generation priors independently of the intended persona signal."},{"response":"The Jaccard deltas and the within-study retrieval-attribution percentages we report are measured directly in our runs and do not depend on Jack 2026. The asymmetry is described as consistent with rather than proven by the cited rates. We will revise the abstract and discussion to separate the descriptive finding (larger point estimate on the Anthropic route) from the interpretive discussion, and will explicitly note that a stronger causal link would require additional within-study experiments not present in the current audit.","revision_made":"partial","referee_comment":"[Abstract (asymmetry paragraph)] The interpretation that the Anthropic vs. OpenAI asymmetry is consistent with retrieval-unattributed generation (43-52% vs. 8-29%) rests on rates documented in Jack 2026. While the empirical deltas are independent measurements, the explanatory claim for why the effect is larger on one route reduces in part to that prior result; direct within-study attribution measurements or clearer separation of the descriptive claim from the causal interpretation would strengthen the argument."}],"tokens_in":1636,"tokens_out":472,"duration_ms":22993,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The key things to know are that adding a persona prefix drops recommendation-set Jaccard by 0.12 to 0.20 relative to baseline and that the shift concentrates on mid-market brands while category leaders stay stable around 80 percent. The Anthropic setup shows a larger point estimate than the OpenAI ones, tied to higher rates of retrieval-unattributed generation.\n\nThe paper applies a clean audit across 10 personas and 8 prompts, then stratifies the results by brand prominence. That breakdown is the useful part: it turns a generic context-sensitivity claim into a more precise statement about where the variation actually occurs. The cross-provider comparison and the link to retrieval attribution rates give the result a bit more grounding than a single-model study would have.\n\nThe soft spot is the absence of any check that the 10 persona prefixes were matched on length, lexical overlap with the query, or syntactic structure. Without that control, some of the measured Delta could trace to prompt surface differences rather than the buyer-context signal the prefixes are supposed to carry. The abstract gives no sign of an ablation or balancing step on this, so the attribution to persona interpretation is not fully locked down. The Anthropic cell also rests on only four prompts, which makes that contrast noisier.\n\nThis is aimed at researchers who audit commercial AI systems or study context effects in retrieval-augmented generation. It offers a replicable protocol and effect sizes that others could test or extend.\n\nI would send it to peer review. The core pattern is straightforward enough that referees can examine the methods and data once the full manuscript is available.","headline":"Persona prefixes cut brand rec overlap by 0.12-0.20 Jaccard with mid-market brands shifting most, but the design leaves prompt-form confounds unaddressed.","tokens_in":2563,"tokens_out":409,"would_cite":false,"duration_ms":27624,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Prefixing the same query with different buyer personas drops AI recommendation-set overlap by 0.12-0.20 in Jaccard index.","keywords":["persona conditioning","brand recommendations","retrieval-augmented generation","AI chatbots","Jaccard similarity","prominence stratification","commercial queries"],"falsifier":"Repeating the audit with a fresh set of personas and finding all clustered confidence intervals for the Jaccard delta include zero.","tokens_in":2775,"feed_emoji":"🤖","tokens_out":355,"duration_ms":13547,"temperature":0.7,"pith_summary":"The paper tests whether buyer context, signaled by a persona prefix, changes which brands commercial AI chatbots recommend for identical prompts. Across 2000 runs on OpenAI and Anthropic models, it measures recommendation-set similarity and finds consistent drops when personas differ. The effect concentrates on mid-market brands while category leaders remain stable. The audit shows that aggregating recommendations without conditioning on persona masks real variation in model output.","feed_headline":"Persona prefixes shift AI brand recommendations by 12-20 percent","feed_subtitle":"Mid-market brands swap up to 75 percent of suggestions when the model infers different buyer contexts.","key_machinery":"Jaccard similarity computed on persona-conditioned recommendation sets, stratified by brand prominence category.","core_discovery":"The same prompt produces materially different recommendation sets depending on who the model thinks is asking, with the effect sharply stratified by brand prominence and largest on the most priors-reliant generation route.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Persona prefixes reduce brand rec similarity by 12-20%","Mid-market brands swap 75% of AI recs when persona changes","AI brand recommendations vary by 12-20% with buyer persona","Persona effects on recs larger for Anthropic than OpenAI"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The ten chosen personas accurately represent distinct buyer contexts and models treat the prefixes only as identity signals.","fun_headline_variants_meta":{"raw":{"variants":["Persona prefixes reduce brand rec similarity by 12-20%","Mid-market brands swap 75% of AI recs when persona changes","AI brand recommendations vary by 12-20% with buyer persona","Persona effects on recs larger for Anthropic than OpenAI"]},"model":"grok-4.3","cost_usd":0.009061,"raw_usage":{"total_tokens":4121,"prompt_tokens":778,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":90612000,"prompt_tokens_details":{"text_tokens":778,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3272,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":778,"tokens_out":71,"duration_ms":28141,"temperature":1.0,"reasoning_tokens":3272,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T07:18:00.886301+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Repeating the audit with a fresh set of personas and finding all clustered confidence intervals for the Jaccard delta include zero.","supporting_citations":[],"review_version":1}