{"id":"a6c5c8a9-ca21-4b6d-aa14-da61de7a60d6","arxiv_id":"2607.07480","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":4,"one_line_summary":"Changing only the prompter's age and gender in AI coding prompts produces statistically significant differences in generated website interface design, template content, and code structure across 800 generated websites and a 20-person user study.","lead":"This paper shows that when AI coding assistants are given a user's age and gender, they generate measurably different software — different colors, different template content, different code structure — even when the programming task is identical. A smart generalist should read it because it reveals that AI-generated code is not task-neutral: invisible demographic signals can shape the structure and design of software in ways users don't notice.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The controlled experiment demonstrates that explicit demographic cues (name+age) influence outputs, but the leap to 'inferred developer attributes' is untested: the user study's 19–36 age range cannot validate age-related inference, leaving the central claim's generalization to real-world personaliz","rationale":"The reader correctly identifies the most load-bearing concern: the gap between explicit demographic cues (name+age in prompts) and 'inferred developer attributes' that the paper's framing emphasizes. I agree this is the central weakness. The controlled experiment is well-designed for what it tests — the effect of explicit demographic signals on AI-generated software — and the statistical findings are robust for interface design and template content (consistent across models, large effect sizes, appropriate corrections). The code structure findings are weaker (model- and task-specific with reversed directions), but the paper acknowledges this. The user study provides valuable qualitative insights about perceptions of personalization but cannot bridge the explicit-to-inferred gap for age due to its 19–36 age range. The CONDITIONAL verdict is appropriate: the core finding that demographic signals influence generated software is well-supported, but the generalization to real-world inference scenarios remains untested. The reader's three secondary concerns (120-website subsample for key analyses, model/task-specific code structure results, and lack of shared artifacts) are valid but do not independently undermine the central claim. I recommend UNCHANGED because the reader's verdict already captures these limitations appropriately.","tokens_in":21853,"tokens_out":8200,"duration_ms":332654,"concrete_test":"Re-run the controlled experiment with an additional condition where demographic information is conveyed implicitly rather than explicitly — e.g., replacing 'Hi, my name is Susan, and I'm 66' with contextual cues like 'I recently retired after a long career in accounting' or 'I'm a freshman excited to build my first website' (age) and using gender-neutral names with hobby/context cues instead of gendered names. Compare effect sizes for interface design, template content, and code structure between the explicit-cue and implicit-cue conditions. If effects are substantially weaker or absent in the implicit condition, the generalization from 'explicit demographic cues influence outputs' to 'inferred developer attributes influence outputs' is not supported by the current evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (§10) states that 'the age and gender of the prompter can significantly and substantially influence AI-generated software.' The controlled experiment provides causal evidence that explicit demographic signals — a persona name and age stated directly in the prompt — produce significant differences across interface design, template content, and code structure. This part of the argument is well-supported: the design holds all non-demographic prompt content constant, uses two models, two tasks, and applies BH correction. However, the paper's framing throughout (title, abstract, §1, §3) emphasizes 'inferred developer attributes' and 'personal information' broadly, while the experiment varies only explicit, unambiguous demographic cues. The mechanism by which real-world AI systems would obtain demographic information — inferring age and gender from conversational patterns, writing style, or stored memory — is fundamentally different from receiving 'Hi, my name is Susan, and I'm 66.' Explicit cues are strong, unambiguous signals; inferred characteristics are weaker, noisier, and may not even be present in many interactions. The paper acknowledges this gap (§7: 'real-world coding assistants may infer similar information from conversational history or writing style'), but the headline claim and conclusion do not adequately distinguish 'demographic cues in prompts' from 'inferred developer attributes.' The user study is designed to bridge this gap by using real ChatGPT accounts with memory, but it cannot adequately test the age dimension: participants range from 19–36 (Table 1), providing almost no variation to assess whether age-related inference produces effects comparable to the controlled experiment. With only 4 participants over 30 and none over 36, the user study offers no meaningful evidence about age-related personalization through inference. For gender, the user study is observational (no control over what demographic information ChatGPT has","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper investigates how demographic information (age and gender) in prompts to LLM-based coding assistants influences generated software artifacts. The authors conduct controlled experiments on 800 AI-generated websites across two tasks (personal website, online shop) and two models (GPT-4.1, DeepSeek-V3.2), varying only persona name and age while holding all other prompt content constant. They find statistically significant differences across three dimensions: interface design (e.g., color palettes, section presence), template content (e.g., generated skills, product categories), and code structure (e.g., lines of code, file organization). A complementary user study with 20 participants examines how developers perceive the boundary between personalization and bias in practice. The controlled experiment is well-designed with appropriate statistical corrections (BH), and the mixed-methods approach strengthens the contribution.","tokens_in":21967,"tokens_out":3931,"duration_ms":138049,"significance":"The paper addresses a timely and important question at the intersection of AI-assisted software engineering and algorithmic fairness. The finding that demographic cues irrelevant to the programming task systematically alter generated artifacts—including non-obvious dimensions like code structure—is novel and has practical implications for AI coding tool design. The three-dimensional framework (interface design, template content, code structure) provides a useful analytical lens. The controlled experiment design—varying only persona name/age across 800 generations with two models and two tasks, applying BH correction—is methodologically sound. The user study adds valuable qualitative depth on developer perceptions of personalization versus bias. The paper ships falsifiable, statistically tested predictions with reported effect sizes (Cramér's V, rank-biserial correlation).","major_comments":[{"comment":"§1, Abstract: The paper frames its contribution around 'inferred developer attributes' and 'inferred user characteristics,' but the controlled experiment provides demographic information explicitly via persona name and age stated directly in the prompt (§3.1.2, Figure 2). The mechanism of explicit demographic cues is fundamentally different from real-world inference of age/gender from conversational patterns or writing style. The paper acknowledges this gap in §7 ('real-world coding assistants may infer similar information from conversational history or writing style'), but the abstract, introduction, and research questions do not adequately distinguish 'demographic cues in prompts' from 'inferred developer attributes.' The conclusion (§10) is more carefully worded ('the age and gender of the prompter'), but the surrounding framing overstates the generalization. This is load-bearing for,","section":null},{"comment":"§3.2, Table 1: The user study is designed to bridge the gap between explicit and implicit demographic signals by using real ChatGPT accounts with memory. However, participants range from 19–36 years old (Table 1), providing almost no age variation to test whether age-related inference occurs in practice. The paper does not discuss this limitation in §8 (Limitations). Since age is one of two demographic variables studied and a central part of the paper's claim, the user study's inability to speak to age-related personalization in practice should be explicitly acknowledged as a limitation, and claims about the user study validating real-world age effects should be tempered accordingly.","section":null},{"comment":"§5.2, Table 4: The skills analysis is conducted only on the 120-website qualitative subsample (15% of 800 generated websites). While the paper is transparent about this (Table 2, §4.1), some significant associations are based on very small counts (e.g., 'Knitting, Crocheting' for GPT: 5 occurrences in one cell, 0 in all others; 'Home Repair' for GPT: 4 occurrences in one cell). The Fisher-Freeman-Halton test is appropriate for small samples, but the paper should discuss the statistical power limitations of this subsample more explicitly, particularly for skills that appear in fewer than 10 instances across the full subsample. The color analysis is partially validated on the full Task 2 sample (Figure 4), but no such validation exists for the skills analysis.","section":null}],"minor_comments":[{"comment":"§3.1.3: The paper states websites were generated through 'each model's chat interface instead of its API' following a 'vibe coding architecture.' It would help to specify the exact dates of data collection and model versions (e.g., GPT-4.1 access dates), as model behavior may change over time.","section":null},{"comment":"§5.3, Table 6: The description of Task 2 code metrics notes that 'the strongest effects occur for DeepSeek; GPT had age-related differences that did not survive comparison correction.' However, several GPT cells in Table 6 show p-values below 0.05 (e.g., CSS files p=0.045, Python files p=0.025, JS files p=0.021). The text should clarify which specific GPT results did not survive BH correction, as the table formatting (darker cells with bold text) is not fully self-explanatory.","section":null},{"comment":"§6.1.2: The observation that '16/20 participants explicitly specified a color scheme' is used to argue that default color choices may go unnoticed. However, this also means the controlled experiment's color findings (where no color was specified) may not directly generalize to real-world usage where users often specify colors. This tension should be acknowledged.","section":null},{"comment":"§3.2.1: The planning vs. no-planning manipulation is introduced but its analysis is limited to brief observations in §6.1.2 (e.g., '7/10 no-planning participants kept the layout'). A more systematic comparison of outcomes between the two conditions would strengthen the user study contribution, or the manipulation should be explained as exploratory.","section":null},{"comment":"Table 1: Gender labels are inconsistent across participants (e.g., 'W' for P3–P8, 'N' for P10, 'F' for P14–P20, 'M' for P1–P2, P11–P13). Standardizing to a single convention would improve readability.","section":null},{"comment":"§5.1, Table 3: The 'Pink' row for GPT-4.1 shows p=0.210 but is included in a table of significant results. The caption states 'We only include colors for which there were at least three examples for a given model,' but the table also includes non-significant results. Clarifying the inclusion criteria would help.","section":null},{"comment":"§9 (Related Work): The paper cites Tonneau et al. [66] ('Different demographic cues yield inconsistent conclusions about LLM personalization and bias'), which appears directly relevant to the explicit-vs-inferred concern. This work should be discussed in §7 (Discussion) rather than only listed in related work, as it bears on the generalizability of the paper's findings.","section":null}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid empirical contribution with a well-designed controlled experiment. The main issue is a framing mismatch: the abstract and introduction promise findings about 'inferred' demographic attributes, but the experiment tests explicit cues. This is fixable through careful revision of framing language and more thorough discussion of the explicit-vs-inferred distinction, possibly foregrounding Tonneau et al. [66] which seems directly on point. The user study, while valuable for perceptions, is too narrow in age range to bridge the explicit-inferred gap for age effects. I would encourage the authors to either reframe the contribution as being about explicit demographic cues (which is still novel and important) or to add a clearer causal chain explaining how explicit cue findings inform expectations about inferred attributes. The core experiment is sound and the central finding is defensible; the revision should primarily address framing and transparency."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a thorough and constructive review. The referee identifies three major concerns: (1) a framing gap between 'inferred developer attributes' in the abstract/introduction and the explicit demographic cues used in the controlled experiment, (2) the user study's limited age range (19–36) and the absence of this limitation from §8, and (3) statistical power limitations in the skills analysis based on the 120-website qualitative subsample. We address each point below and commit to revisions for all three.","responses":[{"response":"The referee is correct that there is a meaningful distinction between explicit demographic cues in prompts and implicit inference of demographic attributes from conversational patterns or writing style, and that our framing in the abstract and introduction does not adequately distinguish the two. We will revise the abstract, introduction, and research questions to accurately characterize the controlled experiment as studying explicit demographic cues in prompts, while reserving the language of 'inferred' attributes for the broader motivating context and the user study (which does involve ChatGPT's memory and cross-conversational history). Specifically, we will: (1) revise the abstract to say that we study how 'age- and gender-related signals in prompts' produce significant differences, rather than 'inferred developer attributes'; (2) adjust the introduction's framing to distinguish our controlled experiment (explicit cues) from the broader phenomenon of inference (implicit cues), positioning the user study as a complementary investigation of the latter; (3) ensure the research questions in §1 and §5 reflect this distinction. We agree that §10's wording ('the age and gender of the prompter') is the appropriate level of generality and will align the rest of the paper accordingly. We will also expand the §7 discussion of this gap, which currently receives only a single sentence, to more thoroughly address the relationship between explicit cues and implicit inference.","revision_made":"yes","referee_comment":"§1, Abstract: The paper frames its contribution around 'inferred developer attributes' and 'inferred user characteristics,' but the controlled experiment provides demographic information explicitly via persona name and age stated directly in the prompt (§3.1.2, Figure 2). The mechanism of explicit demographic cues is fundamentally different from real-world inference of age/gender from conversational patterns or writing style. The paper acknowledges this gap in §7 but the abstract, introduction, and research questions do not adequately distinguish 'demographic cues in prompts' from 'inferred developer attributes.' The conclusion (§10) is more carefully worded ('the age and gender of the prompter'), but the surrounding framing overstates the generalization."},{"response":"The referee correctly identifies that our user study participants (ages 19–36) do not provide sufficient age variation to test age-related personalization in practice. This is a genuine limitation that we failed to acknowledge in §8. We will add an explicit limitation in §8 noting that the user study's age range (19–36) is too narrow to draw conclusions about age-related personalization in real-world settings, and that the user study's findings regarding personalization in practice should be understood as primarily speaking to gender-related effects and general personalization perceptions, not age effects. We will also review §6 to ensure that no claims about the user study validating real-world age effects are made or implied. Where the user study results are discussed in relation to the controlled experiment's age findings, we will clarify that the user study cannot independently confirm age-related effects due to the restricted age range of participants.","revision_made":"yes","referee_comment":"§3.2, Table 1: The user study is designed to bridge the gap between explicit and implicit demographic signals by using real ChatGPT accounts with memory. However, participants range from 19–36 years old (Table 1), providing almost no age variation to test whether age-related inference occurs in practice. The paper does not discuss this limitation in §8 (Limitations). Since age is one of two demographic variables studied and a central part of the paper's claim, the user study's inability to speak to age-related personalization in practice should be explicitly acknowledged as a limitation, and claims about the user study validating real-world age effects should be tempered accordingly."},{"response":"The referee raises a valid concern about statistical power for low-frequency skills in our 120-website qualitative subsample. We agree that this limitation should be discussed more explicitly. We will add a paragraph in §5.2 (or §4.1) acknowledging that skills appearing in fewer than 10 instances across the subsample (e.g., 'Knitting, Crocheting' with 5 occurrences for GPT, 'Home Repair' with 4 occurrences) are based on very small counts, and that while the Fisher-Freeman-Halton test is appropriate for small samples, the statistical power to detect associations for these rare skills is limited. We will note that these results should be interpreted as suggestive rather than definitive, and that the broader patterns (e.g., technical skills for younger personas, physical/creative skills for older personas) are more robustly supported because they appear across multiple skills and both models. We will also acknowledge the asymmetry the referee notes: the color analysis was partially validated on the full Task 2 sample (Figure 4), but no analogous validation was conducted for the skills analysis. We will discuss this as a limitation and note that future work could validate skills findings on larger samples, though we note that the structured nature of skills sections (present in all 120 websites) provides a more systematic basis for comparison than unstructured content.","revision_made":"yes","referee_comment":"§5.2, Table 4: The skills analysis is conducted only on the 120-website qualitative subsample (15% of 800 generated websites). While the paper is transparent about this, some significant associations are based on very small counts (e.g., 'Knitting, Crocheting' for GPT: 5 occurrences in one cell, 0 in all others; 'Home Repair' for GPT: 4 occurrences in one cell). The Fisher-Freeman-Halton test is appropriate for small samples, but the paper should discuss the statistical power limitations of this subsample more explicitly, particularly for skills that appear in fewer than 10 instances across the full subsample. The color analysis is partially validated on the full Task 2 sample (Figure 4), but no such validation exists for the skills analysis."}],"tokens_in":21715,"tokens_out":1340,"duration_ms":162924,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Here's my read on Entezari and Endres. The core finding is solid: when you vary only the name and age in a prompt to an LLM, the generated websites differ across interface design, template content, and code structure. This is a genuine extension of prior bias work, which has focused on runtime code behavior or single-function generation. Showing that demographic signals propagate into file organization, CSS volume, and color palette — not just text output — is new and worth attention. The experimental design is clean: 800 websites, two models, two tasks, ten generations per cell, BH correction applied. The statistical work is appropriate, and the effect sizes (Cramér's V 0.36–0.53 for color) are not trivial. The user study is a nice complement, particularly the finding that participants noticed personalized content but not structural differences in their generated code. That asymmetry is the most interesting qualitative result. The stress-test concern about 'inferred' vs. explicit demographic cues is real but somewhat overstated. The paper does acknowledge this gap in Section 7, and the controlled experiment is upfront about using explicit name+age signals. The title says 'personal information,' not 'inferred attributes,' so the framing is less misleading than the stress-test suggests. That said, the abstract does lean on 'inferred user characteristics,' and the conclusion's claim that 'the age and gender of the prompter can significantly and substantially influence AI-generated software' would be more precise if it said 'explicitly provided age and gender.' The user study can't really bridge this gap — 20 participants aged 19–36 with no control over what ChatGPT knows about them is observational and underpowered for age effects. But it was never going to settle the inference question; it adds texture, not causal evidence. The more grounded concern is the 120-website qualitative subsample for color and skills analysis. That's 15% of the data, and some findings (photo galleries, 10/120 websites) rest on small counts. The paper is transparent about this, and the automated color extraction on all 400 shops for Task 2 helps validate the pattern, but a few specific claims are thin. No code, data, or artifacts are shared, which is a real limitation for a paper making empirical claims about generated software. The qualitative codebook isn't provided either. This matters for a paper where the analysis pipeline (HSB color grouping, skill categorization) involves judgment calls. Overall: the central experimental finding holds up. The paper is for SE researchers working on fairness in AI-assisted development, and for anyone studying how LLM personalization affects software artifacts. It deserves a serious referee who can push on the subsample sizes, the missing artifacts, and the framing gap between explicit cues and inferred attributes.","headline":"Demographic cues in prompts measurably change AI-generated website structure, content, and design — but the leap from explicit cues to 'inferred' attributes is untested","tokens_in":22922,"tokens_out":647,"would_cite":true,"duration_ms":101320,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"AI coding tools change your software based on your name and age","keywords":[],"falsifier":"If a replication held the persona name and age constant but varied only writing style or conversational tone, and found no statistically significant differences in interface design, template content, or code structure, the paper's claim that 'inferred developer attributes' meaningfully influence generated software would be substantially weakened.","tokens_in":22009,"feed_emoji":"🎨","tokens_out":1105,"duration_ms":198052,"temperature":0.7,"pith_summary":"This paper sets out to prove that when an AI coding assistant receives demographic signals about the user — specifically their age and gender — it changes the software it generates in ways that go well beyond what the user asked for. The authors ran controlled experiments on 800 AI-generated websites, holding the actual programming task constant while varying only the persona name and age in the prompt. They found statistically significant differences across three dimensions: interface design (colors and layout sections), template content (skills and product categories assigned to placeholder text), and code structure (file counts, language distribution, and code volume). For instance, older personas were more likely to receive photo galleries, women's online shops had fewer JavaScript files, and color palettes tracked gender stereotypes — more blue for men, more pink and purple for women. A complementary user study with 20 participants showed that developers tend to notice when AI personalizes content but largely fail to notice when it changes interface design or code structure, meaning the less visible forms of demographic influence can persist undetected in deployed software.","feed_headline":"AI coding tools change your software based on your name and age","feed_subtitle":"800 generated websites show demographic cues shift colors, content, and code structure — and developers don't notice the deepest changes","key_machinery":"The core mechanism is the persona-based controlled experiment: identical programming prompts differing only in the prompter's name and age, generating 800 websites across two models and two tasks, then measuring differences in interface design, template content, and code structure with statistical tests (chi-square, Fisher-Freeman-Halton, Mann-Whitney U) and Benjamini-Hochberg correction for multiple comparisons.","core_discovery":"The central finding is that demographic information irrelevant to the programming task — a user's name and age — systematically and significantly alters AI-generated software across all three measured dimensions: visual interface design, template content, and code structure. The mechanism is that the AI model treats demographic cues as implicit design preferences, mapping them to stereotyped assumptions about what colors, skills, products, and code organizations suit different age and gender groups. This happens even though neither task requested personalized output; the models could have produced neutral placeholders. The effects are model- and task-dependent in direction (e.g., one model缩短","pith_inferences":["If the mechanism of demographic influence differs between explicit name-plus-age signals and implicit inference from conversational patterns, then the controlled experiment's effect sizes represent an upper bound rather than a baseline; real-world inference-based effects could be weaker, differently directed, or interact with additional attributes the study did not test.","The user study's age range of 19–36 means the paper cannot test whether older developers would recognize or resist demographic personalization differently than the predominantly young sample did — an empirical gap that matters if older users are both more affected by age-related bias and less likely to detect it.","A natural extension would test whether providing users with a transparency mechanism — showing them what demographic attributes the model inferred and which design decisions those attributes influenced — reduces the acceptance of biased defaults without eliminating genuinely helpful personalization.","The dichotomy between 'web development' for young men and 'web design' for young women in generated skills suggests the models encode a hierarchy that maps technical depth to gender; if this propagates into educational or portfolio contexts, it could reinforce the very pipeline disparities the field has been trying to correct."],"forward_implications":["If demographic cues in prompts can shift code structure — file organization, language distribution, code volume — then two developers of different demographics given the same task could receive software of measurably different maintainability, readability, or complexity, with neither party aware of the disparity.","The finding that users notice content personalization but miss design and code-structure changes suggests that the most consequential forms of demographic bias in AI-generated software are also the least likely to be caught by human review.","As AI coding assistants increasingly retain cross-conversational memory, the range of demographic attributes they can infer — from writing style, interaction patterns, or stored personal data — may expand beyond the two attributes tested here, potentially amplifying the scope of unintended personalization.","The model- and task-dependent direction of code-structure effects means that bias mitigation strategies cannot assume a single consistent direction of harm; the same demographic signal may produce different structural outcomes depending on the model and programming context."],"fun_headline_variants":["Your name changes the code AI writes for you","AI code generators silently tailor output to your age and gender","Developer demographics reshape AI-generated websites","AI coding assistants infer your identity and adjust code accordingly","Age and gender cues alter AI-generated software across three dimensions"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The paper frames its findings around demographic attributes that AI systems 'infer,' but the controlled experiment provides age and gender explicitly via a name and age in the prompt. Real-world AI systems would need to infer these attributes from conversational patterns, writing style, or stored memory, which may produce different — potentially weaker or differently directed — effects than explicit name-plus-age signals.","fun_headline_variants_meta":{"raw":{"variants":["Your name changes the code AI writes for you","AI code generators silently tailor output to your age and gender","Developer demographics reshape AI-generated websites","AI coding assistants infer your identity and adjust code accordingly","Age and gender cues alter AI-generated software across three dimensions","Your demographics change what AI coding tools build","AI-generated code shifts based on developer name and age","Demographic signals rewrite AI-generated software without developers noticing"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":836,"prompt_tokens":521,"completion_tokens":315,"prompt_tokens_details":null},"tokens_in":521,"tokens_out":315,"duration_ms":26560,"temperature":1.0,"reasoning_tokens":289,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T09:29:53.366947+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If a replication held the persona name and age constant but varied only writing style or conversational tone, and found no statistically significant differences in interface design, template content, or code structure, the paper's claim that 'inferred developer attributes' meaningfully influence generated software would be substantially weakened.","supporting_citations":[],"review_version":1}