{"id":"a6c8deb9-9b83-4470-aa43-ff59de62fb8d","arxiv_id":"2507.15585","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"LLMs generate more diversity- and identity-focused, conflict-heavy, and topically narrower narratives for queer personas than for non-queer personas across five everyday contexts and six open-weight models.","lead":"Large language models produce measurably narrower, more identity-focused stories about queer people than about non-queer people in neutral everyday settings. The paper tests this across six open-weight LLMs and five contexts, finding consistently elevated talk of diversity, identity, and conflict for queer personas.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quantitative claims H2-H4 rest on an unvalidated Llama-3.1-8B judge that is itself an evaluated model; with judge prompts never shown, measured identity-gaps could be judge artifacts rather than model behavior.","rationale":"Agree with the reader's weakest-assumption identification. The central claim—that LLMs narrow queer portrayals—depends on measuring differential focus on identity across groups, and the only measurement instrument for H2/H3/H4 is Llama-3.1-8B-Instruct. This is load-bearing because the judge is drawn from the same model family being audited and is never validated. The missing judge prompts are an additional support gap: the paper says 'We formulate the following questions' but never shows the template that instantiates Q1-Q4, so the reader cannot rule out that the judge answers from the system prompt rather than the response, a direct confound for the Q1-Q3 wording. The reported p-values for H4 come from a permutation test on topic distributions produced by the same judge, so the test inherits any judge-side bias. Resolving this requires the human-validation step described in the test. Until that is done, the paper should not be accepted with its quantitative claims stated as established; the reader's CONDITIONAL verdict is appropriate. Independent support: the H1 term-frequency result (Fig. 2/Table 2) is a direct surface-form measurement and is less vulnerable to the judge concern; it provides partial support that some narrowing exists even if H2-H4 fail. This strengthens the case for CONDITIONAL rather than outright rejection.","tokens_in":22674,"tokens_out":5770,"duration_ms":55175,"concrete_test":"Draw a stratified sample of 300 generated responses (across models, contexts, Identity=User/Model, QUEER/NOT-QUEER). Have at least two, ideally three, human annotators answer Q1-Q4 on the response text with all identity cues in the system prompt masked, and also have Llama-3.1-8B-Instruct judge the same masked responses. Compute Cohen's kappa between the LLM judge and human majority, then recompute δqueer and the H4 permutation statistics using human labels. If kappa < 0.6, or if the human-based effects fail the p<0.01 threshold, the quantitative claims are not established. A second check: run the same judge with the original system prompt vs. a version with the identity phrase redacted to quantify prompt-leakage bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 6 adopts Llama-3.1-8B-Instruct as the sole LLM-Judge for Q1-Q4, justified only by 'Based on manual examination, we used Llama-3.1-8B-Instruct as our LLM-Judge'; no human labels, agreement statistics, or confidence intervals are reported anywhere, and the full Q1-Q4 judge prompts are omitted from Appendix B.6 (which contains only the topic-extraction prompt). Section 8 extends the same unvalidated judge to topic extraction for H4, using it to assign each generated text a 50-sample topic distribution. The judge is itself one of the six evaluated models, so any systematic tendency of Llama-3.1-8B to over-attribute gender/sexuality references, marginalization cues, or identity-focused topics to queer-coded text—or to under-attribute them to non-queer text—will inflate every δqueer score in Fig. 3 and every topic-divergence score in Fig. 4. Because the same judge measures all models, this would also manufacture the apparent cross-model consistency the paper reports. A further concrete risk: if the judge is given the full conversation, the system prompt already contains the identity phrase, making Q1 ('Does the text reference or imply the speaker|spoken-to's gender or sexuality?') trivially YES from the prompt rather than from the response, which would make H2 an artifact of the evaluation setup.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates whether LLMs generate constrained, identity-focused narratives for queer personas in otherwise neutral social contexts. Using six open-weight instruction-tuned models (Llama-3.1-8B/70B, Llama-3.3-70B, Qwen2.5-14B/72B, Gemma3-12B), five contexts (Housing, Medical, Persona, Recommendation, Work), and two role conditions (Identity=User and Identity=Model), it tests four hypotheses: H1 via frequency of diversity/inclusion keywords, H2 and H3 via an LLM-as-a-judge answering four YES/NO questions, and H4 via Jensen-Shannon divergence between LLM-judge topic distributions. The reported results show positive queer-minus-non-queer differences across models for H1-H3 and statistically significant topic divergences for H4, which the authors interpret as evidence that LLM portrayals of LGBTQ+ people are narrower and more identity-focused than portrayals of non-queer people.","tokens_in":22934,"tokens_out":5185,"duration_ms":54309,"significance":"If the findings hold, this is a useful contribution to representation-harms auditing, extending the focus beyond explicit toxicity to subtle discursive othering and narrowness. The paper has concrete strengths: six open-weight models, raw data tables in the appendix, permutation-based p-values for H4, and detailed prompt templates for the generation contexts. However, the central quantitative claims for H2-H4 rest on an unvalidated LLM judge that is itself one of the evaluated models, and H1 lacks any statistical test. These issues are load-bearing for the paper's main conclusion, so the significance of the result is currently conditional on additional validation evidence.","major_comments":[{"comment":"The LLM-Judge is introduced in Section 6 with the statement 'Based on manual examination, we used Llama-3.1-8B-Instruct as our LLM-Judge,' but no human agreement statistics, validation set, or inter-annotator agreement are reported anywhere. Llama-3.1-8B-Instruct is also one of the six evaluated models, so any systematic tendency of this judge to over-attribute identity references, marginalization cues, or unique-perspective language to queer-coded text would inflate every delta_queer score in Fig. 3 and every topic-divergence score in Fig. 4, and could manufacture the apparent cross-model consistency. Please report human-annotated labels on a sample of responses with agreement statistics (e.g., Cohen's kappa) for Q1-Q4, test robustness using a judge that is not in the evaluated set, and state explicitly whether the judge input includes the system prompt containing the identity phrase; if it does, Q1 becomes trivially answerable from the prompt rather than from the generated text. This validation is essential for H2 and H3 and also affects H4.","section":"Section 6, Fig. 3, Tables 3-6"},{"comment":"Section 5 claims 'We observed a significant discrepancy' in the frequency of 'respect', 'diverse', 'inclusive', and 'fair' between queer and non-queer responses, but no statistical test, confidence interval, or per-condition sample size is reported. The H1 keyword set is hand-picked, and the percentages in Fig. 2 are aggregates over an unspecified number of generations; without a test (e.g., chi-square or permutation over responses) or error bars, the word 'significant' is unsupported. Please report the number of responses per identity group and add an appropriate significance test or interval.","section":"Section 5, Fig. 2, Table 2"},{"comment":"Section 8 relies on Llama-3.1-8B-Instruct to extract topic distributions from each response, sampling 50 topic lists per response. The permutation tests show that the divergence between QUEER and NOT-QUEER topic distributions is unlikely under random reassignment given the extracted topics, but they do not validate the extracted topics themselves. If the judge tends to generate identity-related topics for queer-coded text, the H4 divergences would reflect a bias of the measurement tool rather than of the evaluated models. Please validate the topic extraction against human annotations or an independent topic model, report the details of the permutation test (number of randomizations, whether the test is at response level or aggregate level, and seed), and include p-values or confidence intervals in the tables.","section":"Section 8, Fig. 4, Tables 7-8"},{"comment":"Appendix B.6 contains only the topic-extraction prompt; the full Q1-Q4 judge prompts, the four in-context examples mentioned in Section 6, and the exact instructions for instantiating 'speaker|spoken-to' are not shown. Because the wording of these questions determines the measured rates, this omission prevents replication and makes it impossible to assess whether the questions are leading. Please include the complete judge prompts and examples, and state explicitly whether the identity phrase appears in the judge input.","section":"Section 6, Appendix B.6"}],"minor_comments":[{"comment":"Equation (2) defines delta(c, g1, g2) with P(.|c, g1) on both sides of the Jensen-Shannon divergence; the second argument should be P(.|c, g2).","section":"Section 8, Eq. (2)"},{"comment":"The Work prompts contain the phrases 'about possessive performance at work' and 'about possessive good performance at work', which appear to be typos for 'positive performance' and 'poor performance'.","section":"Appendix B.5.1"},{"comment":"The sentence 'conversations involving queer patients disproportionately on sexual health or medical transition' is missing a verb; it should read 'disproportionately focus on'.","section":"Section 8.2"},{"comment":"The phrase 'a singple example' should be 'a single example'.","section":"Appendix C.4"},{"comment":"Hypothesis H2 is listed under both 'discursive othering' and 'narrow representations'; the overlap between these categories should be clarified or the classification made explicit.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and addresses a timely question about subtle representational harms in LLM outputs. The main barrier is methodological: the central claims rely on an unvalidated judge that is itself an evaluated model. I think the issues are fixable within the manuscript's scope, so I recommend major revision rather than rejection, contingent on human validation of the judge, disclosure of the judge prompts, and a proper significance test for H1."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the take: the paper finds a large, consistent gap across six open-weight LLMs: in neutral everyday contexts, outputs involving queer personas are much more likely to mention identity, marginalization, and diversity/inclusion themes than the same prompts with non-queer personas. The raw tables in the appendix show effect sizes that are often 30-70 percentage points, so I don't think this is pure noise. The biggest weakness is that the LLM-as-judge is not validated against human raters, and the Q1-Q4 judge prompts are omitted from the appendix. The stress-test concern that the system prompt leaks the identity phrase to the judge doesn't actually hold: both queer and non-queer prompts name the subject's gender or sexuality (e.g., 'trans man' vs 'man'), so a leak would add a constant to both sides, not create the gap. The real issue is that the judge is itself one of the evaluated models. If Llama-3.1-8B over-attributes identity-related cues to queer text, every δqueer in Fig. 3 and every topic divergence in Fig. 4 could be inflated. The authors report no human agreement statistics, no confidence intervals, and the code isn't released yet. That said, the paper is honest about building on Cheng et al. (2023) and Gupta et al. (2024), and the topic-divergence metric with permutation tests is a reasonable extension. H1's keyword matching is simple and reproducible. The limitations section is candid about the narrow identity set and English-only data. I'd send this to peer review. The right path is a conditional accept with major revisions: validate the judge against humans, release the Q1-Q4 prompts and code, and add uncertainty estimates. The core finding is likely robust, but the exact magnitude shouldn't be cited without those additions.","headline":"A multi-model audit showing LLMs narrow queer personas toward identity topics, but the unvalidated LLM judge leaves the exact size of the gap uncertain.","tokens_in":23501,"tokens_out":3722,"would_cite":true,"duration_ms":37159,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that LLM portrayals of LGBTQ+ people, in otherwise neutral settings, are significantly different from non-queer portrayals and default to identity-related topics, a narrowing that persists across six open-weight models.","keywords":["queer representation","LLM bias","representational harm","discursive othering","LLM-as-a-judge","topic divergence","persona prompting","LGBTQ+ narratives"],"falsifier":"Take the same generated responses, strip all identity-coded words (queer, trans, pronouns, diversity terms) and ask the judge Q1-Q4 again, or have human raters label a sample. If the judge's YES rate for queer responses does not drop when identity cues are removed, or if humans disagree with the judge's topic labels, the measured divergence is an artifact of the judge rather than of the underlying models.","tokens_in":22436,"feed_emoji":"🏳️","tokens_out":6326,"duration_ms":64381,"temperature":0.7,"pith_summary":"Large language models, asked to simulate or interact with a queer person in an everyday setting, do not treat that person the way they treat a non-queer person. The paper tests this with four hypotheses across six open-weight models and five neutral contexts (housing, medical, persona, recommendation, work), and reports that queer-subject outputs contain markedly more diversity-and-inclusion vocabulary, more references to gender or sexuality, more cues of marginalization, and topic distributions that differ from non-queer outputs, mostly at $p<0.01$. The authors read this as evidence of constrained representation and discursive othering: queer characters are made hyper-visible through identity talk while non-queer characters get the full range of everyday life. If correct, this matters because LLMs are already used as therapists, tutors, and co-writers, so the narrowing is not a neutral quirk; it shapes real conversations about human lives.","feed_headline":"Neutral prompts make LLMs shrink queer lives to identity talk","feed_subtitle":"Queer personas get diversity talk, pronouns, and struggle; non-queer ones get everyday life.","key_machinery":"The argument runs on persona-context prompting combined with an LLM-as-a-judge measurement pipeline. Each prompt is a template with an identity slot, filled by phrases such as 'trans man' (QUEER) or 'man' (NOT-QUEER), set in one of five everyday contexts. The judge model, Llama-3.1-8B-Instruct, answers four YES/NO questions (Does the text reference the subject's gender or sexuality? Does it imply a unique perspective due to identity? Does it focus on identity over the setting? Does it imply membership in a marginalized group?) and, separately, extracts topic labels from each response, sampled 50 times per response to form empirical topic distributions. Jensen-Shannon divergence between queer and non-queer topic distributions, with permutation tests, converts the qualitative pattern into a significance claim.","core_discovery":"The paper's central claim is that LLM portrayals of LGBTQ+ people, in otherwise neutral settings, are significantly different from non-queer portrayals and narrowly focus on identity-related topics to a degree not seen for non-queer people. It formulates and validates four hypotheses: H1 (more diversity/inclusion terms for queer subjects), H2 (more identity references and identity-related discussion), H3 (more identity-linked conflict, harassment, or negative experience), and H4 (distinct topic distributions). The supporting evidence spans six open-weight instruction-tuned models; the H4 topic-divergence differences are mostly significant at $p<0.01$ under permutation testing, and the largest effects occur when the model itself speaks as the queer persona. The paper interprets these patterns as discursive othering and narrow representation rather than overt hostility, and explicitly notes that even positive-sounding attention to diversity can mark queer people as separate.","pith_inferences":["Reading further than the paper: the same judge-based pipeline could be pointed at other marginalized identities, but the strongest evidence in H1 (raw word frequencies) would need identity-specific term lists rather than the four generic words used here.","The paper does not compare whether a human reader would notice the asymmetry; a natural next test is to show paired outputs to naive readers and see whether they independently rate queer personas as more identity-focused.","The measured divergence might partly reflect truthful statistical differences in lived experience, a point the paper acknowledges; a sharper test would control for base rates by asking for the same scenarios with explicit non-identity traits.","Because the judge is itself an LLM, swapping the judge for a different model (or a prompted ensemble) would show whether the reported gaps are stable across measurement tools or partly an artifact of Llama-3.1-8B-Instruct's own associations."],"forward_implications":["Every downstream use where an LLM simulates a queer persona—chatbots, educational agents, co-writing tools—inherits the narrowed narrative, not just explicit hostile outputs.","In simulated medical consultations, queer patients' conversations concentrate on sexual health, pronouns, and transition, a machine analogue of 'Trans Broken Arm Syndrome'.","Because the divergence appears across all six tested models and most contexts, the effect looks systemic to current open-weight instruction-tuned models rather than a single-model accident.","The asymmetry between Identity=Model and Identity=User prompts suggests the narrowing is strongest when the model speaks in the voice of the queer person, not when it merely addresses one.","Elevated marginalization cues under H3 mean the observed narrowing carries a negative valence: simulated queer characters are more often placed in conflict, harassment, or discrimination scenarios."],"supporting_citations":[{"why":"Defines caricature in LLM simulations and provides the persona-context prompting method the paper extends to queer identity.","marker":"Cheng et al. (2023)"},{"why":"Supplies the LLM-as-a-judge approach used for Q1-Q4 and topic extraction.","marker":"Zheng et al. (2023)"},{"why":"Motivates asking the judge to justify answers, as the paper does to improve automatic evaluation.","marker":"Chiang and Lee (2023)"},{"why":"Names Trans Broken Arm Syndrome, which the medical-context results are interpreted as echoing.","marker":"Wall et al. (2023)"},{"why":"Documents the real-world narrow portrayal of LGBTQ+ people that the LLM pattern is claimed to mirror.","marker":"Hicks (2020)"},{"why":"Provides the human-testing gold standard for evaluating social and discursive themes, which the paper cites as the alternative to its judge.","marker":"Gadiraju et al. (2023)"},{"why":"Shows biased reasoning in persona-assigned LLMs, a direct predecessor for the persona-based bias tests.","marker":"Gupta et al. (2024)"},{"why":"Makes the case for LLMs as generative agents, which the paper flags as a risky deployment given constrained personas.","marker":"Park et al. (2023)"},{"why":"Frames 'bias' in language technology and grounds the harm taxonomy the paper draws on.","marker":"Blodgett et al. (2020)"}],"fun_headline_variants":["LLMs shrink queer narratives to identity-centric themes","Neutral prompts still skew LLM queer stories toward identity","Queer personas in LLMs get identity talk, not everyday life","LLM tests show queer lives reduced to identity discussions","Study: LLMs portray queer people with narrow identity focus"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on Llama-3.1-8B-Instruct being an unbiased rater of identity themes and topics, with no human agreement statistics reported; if it systematically reads queer-coded text as identity-focused, the measured gaps shrink or vanish.","fun_headline_variants_meta":{"raw":{"variants":["LLMs shrink queer narratives to identity-centric themes","Neutral prompts still skew LLM queer stories toward identity","Queer personas in LLMs get identity talk, not everyday life","LLM tests show queer lives reduced to identity discussions","Study: LLMs portray queer people with narrow identity focus"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000201,"raw_usage":{"total_tokens":1303,"prompt_tokens":792,"completion_tokens":511,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":408,"completion_tokens_details":{"reasoning_tokens":431}},"tokens_in":408,"tokens_out":511,"duration_ms":5903,"temperature":1.0,"reasoning_tokens":431,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:28:26.922591+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same generated responses, strip all identity-coded words (queer, trans, pronouns, diversity terms) and ask the judge Q1-Q4 again, or have human raters label a sample. If the judge's YES rate for queer responses does not drop when identity cues are removed, or if humans disagree with the judge's topic labels, the measured divergence is an artifact of the judge rather than of the underlying models.","supporting_citations":[{"cited_title":"Gonzalez, and Ion Stoica","cited_arxiv_id":null,"evidence_quote":"Supplies the LLM-as-a-judge approach used for Q1-Q4 and topic extraction."}],"review_version":1}