{"id":"90b59fe9-75ed-4159-b592-2151f172d618","arxiv_id":"2602.16151","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 86 ACL papers finds queer NLP research is reactive, English-centric, and under-involves queer communities.","lead":"This survey systematically reviews 86 papers from the ACL Anthology on LGBTQIA+ topics in natural language processing, mapping trends and gaps. It finds that most queer NLP research is reactive, English-centric, and rarely involves queer stakeholders, offering a roadmap for future work.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Selection bias in the corpus is the load-bearing risk: the semi-manual, under-specified search and the unreconciled 122 vs 86 paper count make the headline percentages (76.7% English, reactive/proactive, stakeholder gaps) conditional on corpus construction.","rationale":"The reader identified the weakest assumption as corpus representativeness, specifically the keyword-based search plus manual expansion. I agree: all of the paper's headline statistics are proportions over this corpus, so selection bias is the most load-bearing threat. The unreconciled n=122 versus n=86 is an objective, checkable red flag that makes the denominator unreliable. The annotation reliability is a secondary concern, not the primary one. The paper has genuine strengths: it is transparent about its ACL focus, acknowledges the keyword limitation, and the qualitative synthesis (pronouns, hate speech, stakeholder gaps) is plausible and supported by the cited papers. However, unconditional acceptance is not appropriate until the corpus and selection procedure are released and the counts are reconciled. The reader's CONDITIONAL verdict is therefore the right call, and my concern does not change it.","tokens_in":23771,"tokens_out":5801,"duration_ms":58502,"concrete_test":"Release the full list of 86 (or 122) papers, together with the search log and PRISMA screening decisions. Then independently re-run the ACL Anthology search with an expanded, preregistered keyword set adding at least 'non-binary', 'nonbinary', 'homophobia', 'transphobia', 'gender-neutral', 'sexual orientation', 'same-sex', and non-English equivalents such as 'homofobia' and 'transfobia', and compare the retrieved set with the original list. If the expanded search adds papers not in the original set, or if the English-only proportion or the reactive/proactive split changes by more than 5 percentage points, the central survey claims are materially affected by selection bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims—most papers are reactive, 76.7% are English-only, stakeholders are rarely involved—are computed over a corpus whose construction is not fully reproducible. Section 3.1 uses a keyword list ('queer, aromantic, gender non*, lgbt*, agender, glbt, lesbian, gay, bisexual, transgender') that omits common alternative terms such as 'nonbinary', 'homophobia', 'transphobia', 'gender-neutral', and 'sexual orientation'. After the automated search returned 3,864 entries and manual filtering left 55 papers, the authors 'expanded' the list based on personal knowledge and Semantic Scholar, and added 19 ACL 2025 papers. No PRISMA flow diagram, inclusion/exclusion log, or full paper list is provided in the manuscript, despite the PRISMA claim. The denominator is also unstable: the Abstract says n=122, while Section 3.1 says the final survey covers 86 ACL papers; the headline 76.7% English corresponds to 66/86, indicating the statistics use 86 without reconciling the abstract's 122. If papers using different terminology (or in less-central venues) were missed because no author happened to know them, the headline findings about English-centrism, reactive research, and stakeholder gaps could be artifacts of the selection procedure rather than properties of queer NLP as a field. The Limitations section explicitly acknowledges this risk. The moderate inter-rater reliability for intersectionality (Cohen's kappa = .62, based on only 11 double-coded papers) adds noise, but the selection process is the primary threat to the central claims.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a systematic, PRISMA-style survey of NLP research addressing LGBTQIA+ topics in the ACL Anthology. The authors describe a keyword search, manual filtering, and community-based expansion resulting in a corpus of 86 ACL papers (the abstract states n=122). They annotate papers along axes including task, method, queer groups, harms, language, intersectionality, and stakeholder involvement; group them into seven categories; and find that most work is reactive (bias discovery), English-centric, and lacking stakeholder/intersectional engagement. The paper also offers qualitative discussion and future directions from queer theory.","tokens_in":24138,"tokens_out":3757,"duration_ms":37989,"significance":"If the corpus is accepted as representative, the survey is a valuable and timely map: it is, to my knowledge, the first systematic survey of queer NLP across the ACL Anthology, and its descriptive statistics give a concrete basis for claims about Anglocentrism and reactive research. Strengths include the community-led author team, explicit positionality, the open-source living repository, and the careful distinction between ACL and non-ACL works. The qualitative sections on refusal and non-English venues are thought-provoking. However, the quantitative findings inherit the reproducibility limitations described below.","major_comments":[{"comment":"The sample size is inconsistent: the Abstract states n=122 and 'all such papers published in the ACL Anthology', while §3.1 says the final survey covers 86 ACL papers. Figure 2/Table 1 report 66 English papers, and 66/86 = 76.7%, so the headline percentages are computed on 86, not 122. The discrepancy is not explained (e.g., the 122 could include non-ACL papers, duplicates, or pre-2025 search results). Since the paper's central quantitative claims depend on the denominator, the authors must reconcile these numbers and state which n underlies each statistic.","section":"Abstract/§3.1"},{"comment":"The paper selection is not fully reproducible despite the PRISMA label. The keyword list (queer, aromantic, gender non*, lgbt*, agender, glbt, lesbian, gay, bisexual, transgender) omits common alternative terms such as nonbinary, homophobia, transphobia, gender-neutral, and sexual orientation. From 3,864 search hits, manual filtering yielded 55 papers; the authors then 'expanded' the list based on personal knowledge and Semantic Scholar and added 19 ACL 2025 papers. No PRISMA flow diagram, inclusion/exclusion log, or complete list of the 86 papers is provided in the manuscript (the repository is cited, but the manuscript should be self-contained). Because the headline claims about English-centrism, reactive research, and stakeholder gaps are proportions over this corpus, a non-reproducible or biased selection could change the conclusions. The Limitations section acknowledges this risk, b","section":"§3.1 Paper Selection"},{"comment":"Inter-rater reliability is computed on only 11 of 86 papers, and the reported agreement is 90.09% with Cohen's kappa = .79 for stakeholder involvement and .62 for intersectionality. Kappa = .62 is moderate, not high. Intersectionality is one of the three headline gap claims (Figure 1 and §5.3). With single annotation for the other 75 papers and inductive coding used to form the seven categories, the reliability evidence is too thin to support the strength of the intersectionality and stakeholder-gap claims. Please report per-axis kappas with confidence intervals, the full contingency tables, and a larger reliability sample, or soften the corresponding quantitative claims.","section":"§3.2 Annotation Process"},{"comment":"The classification of papers as 'reactive' versus 'proactive' appears to be interpreted by the annotators, but the coding scheme for this distinction is not specified in §3.2 or §5.2. Since the Abstract's central claim is that 'most papers take a reactive rather than a proactive approach,' the operational definition of this binary (e.g., which categories count as proactive, how mitigation-only vs. evaluation-only papers are coded) should be stated explicitly and supported by reliability evidence.","section":"§5.1/§5.2"}],"minor_comments":[{"comment":"'queer ALC papers' should be 'queer ACL papers'.","section":"§5.1"},{"comment":"'Retrivied' is misspelled (Human Rights Campaign, Nonbinary Wiki entries).","section":"References"},{"comment":"The caption should state explicitly that percentages refer to the share of papers, not languages, and that the total exceeds 100% because a paper can cover multiple languages.","section":"Figure 2"},{"comment":"The header 'Language # Singletons (1 each)' is confusing; consider separating singleton count from the language list.","section":"Table 1"},{"comment":"Please clarify whether 'language diversity' reliability is perfect because all 11 papers were coded identically, and report the number of categories considered.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a community-led survey, and self-citation is expected; nevertheless, the manual expansion step and the 'reactive vs. proactive' coding could benefit from an explicit reflexivity note about potential confirmation bias. The main scientific contribution is the quantitative gap analysis, so the corpus must be made fully auditable. The inconsistency between n=122 and n=86 and the thin reliability sample are fixable within the manuscript's scope, hence major_revision rather than reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first systematic map of queer NLP in the ACL Anthology, and it is genuinely useful. The headline findings — English-centric, reactive rather than proactive, little stakeholder involvement — are plausible, but the corpus construction is the load-bearing part and it is not fully reproducible as written. Worth sending to peer review; it will come back with requests for transparency fixes, not because the core survey is wrong.\n\nWhat's new: prior surveys cover gender bias or sexuality bias; this one covers all queer NLP papers across tasks in the ACL Anthology, with an explicit coding scheme, a repository, and quantitative gap analysis. The section on non-English venues (IberLEF, EVALITA, Brazilian workshops) is a nice corrective to the usual Anglocentric view, and the discussion of 'critical refusal' and queer theory points to concrete gaps in how the field frames harm. The authors are transparent about positionality and limitations, which is more than most surveys do.\n\nWhere it gets soft: the paper claims PRISMA but does not show a flow diagram, inclusion/exclusion log, or full paper list in the manuscript. The search keyword list omits terms like 'nonbinary', 'transphobia', 'homophobia', 'sexual orientation', and 'gender-neutral', and the expansion step relies on the authors' collective knowledge. That is not a fatal flaw for a community survey — but then the quantitative claims (76.7% English, reactive/proactive split, stakeholder rates) are conditional on that corpus, and the paper should say so more loudly. The abstract says n=122 while Section 3.1 says 86 ACL papers; the percentages appear to use 86. This is a concrete inconsistency that a referee should catch. Inter-rater reliability on 11 papers is thin, but this is a survey, not a psychometric study; the .62 kappa on intersectionality is worth noting but not disqualifying.\n\nThe central claims are not fitted from data; they are descriptive. The claim that most papers are reactive rather than proactive is a judgment call, and the authors support it with examples and category counts. It holds up reasonably.\n\nWho is this for? Anyone working on bias, fairness, queer NLP, or social NLP. It would be a good reading-group paper. I would cite it. It deserves a serious referee — conditional accept, with the caveat that the authors must reconcile the counts and publish the full paper list and annotations.","headline":"First systematic map of queer NLP in the ACL Anthology; the corpus-building is the weak spot, not the core findings.","tokens_in":24692,"tokens_out":1582,"would_cite":true,"duration_ms":15401,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Queer NLP research is growing but reactive, English-centric, and rarely community-led, a systematic survey finds.","keywords":["queer NLP","LGBTQIA+","literature survey","bias in natural language processing","intersectionality","stakeholder involvement","language diversity","NLP ethics"],"falsifier":"A concrete check is to rerun the annotation on an expanded corpus: search beyond the main archive with additional identity keywords (including community terms in languages such as Arabic, Hindi, Portuguese, and Japanese) and include regional conference proceedings and preprint servers. If the expanded sample shows proactive mitigation papers are as common as reactive bias papers, or that non-English papers form a large share of the work, then the survey's central claims about a predominantly reactive, English-centric field would not survive.","tokens_in":23672,"feed_emoji":"🏳️‍🌈","tokens_out":5755,"duration_ms":58061,"temperature":0.7,"pith_summary":"This survey aims to establish what natural language processing (NLP) research about LGBTQIA+ communities actually looks like, by systematically collecting and annotating papers from the field's main publication archive through early 2026. The authors argue that queer NLP is growing but lopsided: most papers are reactive—they expose bias in existing models—rather than proactive, and they lean heavily on template-based evaluations and data augmentation. The paper also finds that the research is dominated by English and Western European languages, rarely takes an intersectional view, and seldom involves queer stakeholders despite citing harm to queer communities as motivation. The survey matters because it turns scattered findings into a map of gaps and a call for future work that is multilingual, community-centered, and not limited to post-hoc fixes.","feed_headline":"Most queer NLP research finds bias, not fixes","feed_subtitle":"A systematic review of 100+ papers shows English dominates and queer stakeholders are rarely involved.","key_machinery":"The load-bearing apparatus is the survey's annotation framework: each paper is coded for NLP task, method, explicit queer groups, harms addressed, language and geographic region of the data, whether intersectionality is addressed, and whether stakeholders are involved. The central distinction is 'reactive' versus 'proactive' research—detecting bias in systems that exist versus building new systems or solutions—and the paper uses this contrast to organize its findings and future agenda.","core_discovery":"The central claim is that queer NLP, as represented in the main NLP archive, has three structural features: a reactive orientation in which identifying bias far outpaces mitigating it; a language distribution heavily weighted toward English; and a systematic absence of stakeholder involvement and intersectionality, even though most papers name queer communities as their motivation. The authors reach this by annotating papers along task, method, named queer groups, harms, language and region of data, treatment of intersectionality, and stakeholder engagement. They conclude that these gaps are not random but reflect methods—template prompting, augmentation, post-hoc debiasing—that are convenie","pith_inferences":["Our inference: the reactive pattern may be reinforced by publishing incentives, since bias findings are comparatively easy to benchmark, and the survey does not examine this incentive structure.","Our inference: because the survey is tied to one publication archive, it may undercount proactive and non-English work published in regional venues or as preprints, so the true gap could be smaller than reported.","Our inference: the paper's discussion of 'critical refusal' points to a limit of inclusion—some users may want systems that decline to classify them at all, a design principle that goes beyond better benchmarks.","Our inference: a direct extension would apply the same annotation protocol to non-archive venues and preprint repositories; a finding that proactive work is common there would qualify the survey's conclusions as archive-specific."],"forward_implications":["If the field is as reactive as the survey shows, existing bias benchmarks are an incomplete evidence base for understanding queer harm in NLP.","Community-in-the-loop evaluation, such as building benchmarks with queer annotators, would have to shift from a rare exception to a standard practice.","The survey's examples of non-English shared tasks—like hope-speech detection and slur-reclamation—provide concrete templates for multilingual queer NLP that could be adapted beyond those languages.","A shift toward proactive mitigation would reallocate research effort from reporting model failures to designing systems that support queer users from the start.","The paper's structural suggestions—allowing users to opt out of classification and building dynamic evaluations—imply that future systems should accommodate fluid identities and shifting context over time."],"fun_headline_variants":["Queer NLP survey: 100+ papers, few fixes","Reactive bias hunting dominates queer NLP","Queer NLP: English-centric, stakeholder-free","Queer NLP: Most papers detect, few repair"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the keyword-based search of the main publication archive, expanded by the authors' familiarity with the field, yields a representative and complete set of queer NLP papers; if many relevant papers use different terminology, appear in non-English venues, or exist only as preprints, then the quantitative claims about English dominance and the reactive/proactive split would be skewed.","fun_headline_variants_meta":{"raw":{"variants":["Queer NLP survey: 100+ papers, few fixes","Reactive bias hunting dominates queer NLP","Queer NLP: English-centric, stakeholder-free","Queer NLP: Most papers detect, few repair"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000167,"raw_usage":{"total_tokens":1093,"prompt_tokens":744,"completion_tokens":349,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":300}},"tokens_in":488,"tokens_out":349,"duration_ms":3986,"temperature":1.0,"reasoning_tokens":300,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T22:36:44.503004+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check is to rerun the annotation on an expanded corpus: search beyond the main archive with additional identity keywords (including community terms in languages such as Arabic, Hindi, Portuguese, and Japanese) and include regional conference proceedings and preprint servers. If the expanded sample shows proactive mitigation papers are as common as reactive bias papers, or that non-English papers form a large share of the work, then the survey's central claims about a predominantly reactive, English-centric field would not survive.","supporting_citations":[],"review_version":1}