{"id":"a89e7e2a-dbf6-45f4-a1b2-d1cf73dc8ad3","arxiv_id":"2607.07937","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Debiasing language-model training data for a target group frequently increases stereotyping or counter-stereotyping for non-target groups across categories, models, and scales.","lead":"Preprocessing methods that scrub stereotypes from training data for one demographic often raise stereotyping or reverse-stereotyping for other groups, even unrelated ones. This matters because data-level debiasing is widely used yet can silently redistribute harm that standard benchmarks miss.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged single-seed and proxy-score limitations.","rationale":"The paper's strongest claim is an empirical regularity documented across a large, publicly released experimental grid rather than a causal or mechanistic assertion. The single most load-bearing condition for that claim is that the observed directional SS movements are not pure artifacts of seed, detector noise, or benchmark composition. The authors already flag exactly these issues in the Limitations section and calibrate language to directional trends. Because the grid is large, the pattern appears under multiple interventions (including pure removal and pure swapping), and code is released, the claim holds within the stated scope. The reader's CONDITIONAL verdict already conditions on multi-seed confirmation and broader model coverage; no stronger objection is required. The proposed concrete_test is simply the natural next verification step already implied by the reader's weakest_assumption.","tokens_in":25038,"tokens_out":496,"duration_ms":6549,"concrete_test":"Re-run the full TinyBERT DG-male and RG-female pre-training cells (full Wikipedia) with three additional seeds (e.g., 0, 1, 123) and recompute per-group SS deltas relative to the matched base model; if the sign of the largest non-target shifts (female SS under DG-male; Christian SS under RG-female) flips in a majority of seeds, the side-effect claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that preprocessing (DG/RG/SR) that moves target-group StereoSet SS / CrowS-Pairs scores toward 50 frequently moves non-target scores away from 50, including across categories, and that this pattern is robust across the experimental grid. The manuscript supplies the full grid (TinyBERT + GPT-2, pre-/post-training, 100 % / 5 % Wikipedia, three interventions, six groups) plus one LLaMA-2-7B slice, public code, and explicit Limitations that already list the single fixed seed (42) and the possibility that DG detector false positives or SR distributional artifacts contribute. Attention-rollout distances remain small even when SS shifts, so the paper does not over-claim a mechanism. No internal inconsistency or hidden assumption that would overturn the directional pattern is present; the reader's weakest_assumption already captures the main soundness limit.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that preprocessing-based stereotype mitigation (DG: remove stereotypical sentences; RG: remove group mentions; SR: swap group references) on Wikipedia, used for pre- or post-training of TinyBERT and GPT-2 (plus one LLaMA-2-7B RG-female slice), reliably lowers StereoSet SS / CrowS-Pairs scores for the target demographic while frequently moving non-target scores away from the neutral baseline of 50, including across gender/race/religion categories. These side effects appear under both full and 5% data scales, are not explained by simple stereotype/anti-stereotype content redistribution, and are not accompanied by large attention-rollout shifts (PD/SD/JSD/L2). The authors supply the full experimental grid, extensive appendix tables (A1–A42), public code, and an explicit Limitations section, and recommend side-effect-aware diagnostics and evaluation.","tokens_in":25283,"tokens_out":1146,"duration_ms":13104,"significance":"If the directional pattern holds, the result is practically important: data-level debiasing is widely used precisely because it is simple and inference-cost-free, yet the paper shows that target-group gains can be purchased by redistributing harm to other groups, including groups outside the intervention category and outside the guiding benchmarks. Strengths that raise the contribution include the systematic grid (two architectures, three interventions, pre-/post-training, two scales), the operational definition of side effects, the public code release, the attention-rollout negative result that avoids over-claiming mechanism, and the open Limitations (single seed, compact models + one large slice, English Wikipedia only, possible DG false positives and SR artifacts). The work therefore supplies both a cautionary empirical finding and concrete evaluation recommendations for the fairness community.","major_comments":[{"comment":"§3 Evaluation protocol and Limitations: all SS/CrowS-Pairs directional claims rest on a single fixed seed (42) with no multi-seed uncertainty or significance tests. Because the central claim is the frequency and robustness of non-target shifts (not merely that a target SS can move toward 50), the manuscript should either report multi-seed runs for a representative subset of the grid or supply bootstrap/permutation intervals on the directional changes so that readers can judge how often the side-effect pattern survives sampling variation.","section":null},{"comment":"§5.1–5.2 and Tables A1–A21: the paper correctly notes that side effects are not fully explained by altered stereotype/anti-stereotype content ratios, yet it never quantifies those ratios (or co-occurrence statistics) before vs. after each DG/RG/SR intervention. Without that measurement, the claim that distributional explanations are insufficient remains qualitative; a short appendix table of pre-/post-intervention stereotype-sentence counts or PMI shifts for the six groups would make the argument load-bearing rather than suggestive.","section":null},{"comment":"§6 and Figure 5: attention-rollout distances are reported as small (max PD 0.0061, JSD 0.0925, etc.) even when SS moves substantially, which is a useful negative result. However, the paper treats rollout as the primary mechanistic probe without a positive control (e.g., a known semantic intervention that does produce large rollout shifts on the same StereoSet-derived bench). Adding one such control, or explicitly bounding what rollout can and cannot detect, would strengthen the conclusion that “semantic routing alone does not explain the effect.”","section":null}],"minor_comments":[{"comment":"Figures 1–2 and the corresponding appendix tables mix “closer to 50” (underlined) and “farther from 50” (bold) conventions; a single consistent legend and color scale across all SS plots would improve readability.","section":null},{"comment":"§3: the DG detector is cited with F1=0.98 on CrowS-Pairs; a brief note on its false-positive rate on Wikipedia-style prose (or a small manual audit) would help readers assess contamination risk for the DG condition.","section":null},{"comment":"§7.2 Table 1: LLaMA-2 results are reported for only the RG-female setting; even a one-sentence statement of why other interventions were computationally infeasible would clarify the scope of the “massive models” claim.","section":null},{"comment":"Typos and notation: “canincrease” (abstract), occasional inconsistent hyphenation of “pre-/post-training,” and the use of both “SG-” and “SR-” labels in Figure 5 should be cleaned for camera-ready.","section":null},{"comment":"Ethics / Limitations: the recommendation to document non-targeted groups in model cards is valuable; a short checklist or example diagnostic (e.g., “report ΔSS for all six groups after any single-group intervention”) would make the actionable advice more concrete.","section":null}],"recommendation":"major_revision","confidential_remarks":"The central empirical pattern is well-documented and the Limitations section already flags the main soundness constraints (single seed, proxy benchmarks, compact models). The requested revisions are fixable within the existing experimental scope and do not require a new research program. Fit for a fairness/NLP venue is strong; I would not reject on novelty or scope grounds."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: when you clean Wikipedia for one demographic (remove stereotypical sentences, strip group mentions, or swap references) and then pre- or post-train TinyBERT or GPT-2, StereoSet and CrowS-Pairs scores for the target group usually move toward 50, but non-target groups—including across gender/race/religion—frequently move away from 50. The pattern holds at full and 5 % data scale and shows up in one LLaMA-2-7B post-training slice. That is the new result.\n\nWhat the paper does well is the experimental grid itself. Three distinct preprocessing operators, two architectures, pre- and post-training, two data scales, six groups, full tables A1–A42, public code, and an attention-rollout check that shows the SS shifts are not accompanied by large changes in attention flow. They do not invent a mechanism they cannot support; they just document that the side effects are common, asymmetric, and hard to predict from co-occurrence counts alone. LMS stays stable so iCAT often looks better, which is exactly why standard aggregate reporting misses the problem. Limitations are stated cleanly: single seed 42, compact models plus one larger slice, English Wikipedia only, possible DG detector false positives and SR artifacts.\n\nThe soft spots are real but already named. Directional SS movement under one seed is a proxy, not a multi-seed significance claim, and StereoSet/CrowS-Pairs composition can amplify or hide certain associations. That does not overturn the directional pattern across dozens of matched conditions, but it does mean the result is still conditional on broader confirmation. Citation pattern is ordinary survey-plus-benchmark; no circularity.\n\nThis is for people who actually ship or audit data-level fairness interventions and for anyone writing model cards. It is not a theory paper and does not reorganize the field, but it is a clear, reproducible warning that “debias the target group and check the aggregate score” is insufficient. I would bring it to reading group, cite the side-effect observation when discussing preprocessing, and send it to peer review. The evidence is extensive enough and the claim important enough that a serious referee should see it, even if they demand multi-seed runs and larger models in revision.","headline":"Solid empirical grid showing that common data-level debiasing often redistributes stereotype scores across groups; single-seed and proxy-score limits are already flagged by the authors.","tokens_in":25830,"tokens_out":566,"would_cite":true,"duration_ms":6648,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Cleaning stereotypes out of training data for one group often worsens them for others, even unrelated ones.","keywords":["stereotype mitigation","preprocessing debiasing","side effects","StereoSet","CrowS-Pairs","attention rollout","language models","Wikipedia"],"falsifier":"Re-run the same three preprocessing recipes on the same Wikipedia snapshot with multiple random seeds and a second independent stereotype detector; if non-target score movements reverse sign or vanish while target improvements remain, the claimed side-effect phenomenon collapses.","tokens_in":25972,"feed_emoji":"⚖️","tokens_out":604,"duration_ms":6369,"temperature":0.7,"pith_summary":"Preprocessing-based debiasing—removing stereotypical sentences, stripping group mentions, or swapping identity references in Wikipedia before pre- or post-training—reliably lowers stereotype scores for the targeted demographic. The paper shows this frequently comes with side effects: stereotyping or counter-stereotyping rises for non-target groups, including across gender, race, and religion categories that share no obvious content overlap. The pattern holds for both encoder-only and decoder-only models, at full and 5% data scales, and even appears in a larger model after limited post-training. Standard aggregate benchmarks often miss the shifts, and attention-rollout maps change little, so the redistribution is not explained by large attention rewiring. The authors therefore treat data-level mitigation as an intervention whose collateral effects must be measured, not as a clean fix.","feed_headline":"Debiasing one group often worsens stereotypes for others","feed_subtitle":"Removing or swapping identity content for a target demographic redistributes bias across unrelated groups","key_machinery":"Side-effect definition and measurement: a mitigation that improves the target group's stereotype score while moving one or more non-target groups farther from the neutral baseline of 50 on StereoSet (or the corresponding CrowS-Pairs category scores), tracked by directional change relative to the matched base model under fixed seed and scale.","core_discovery":"Interventions that reduce measurable stereotypes for a chosen demographic on Wikipedia-trained models systematically induce side effects: undesired movement of stereotype scores away from neutrality for other groups, including across unrelated categories. These side effects are robust across three common preprocessing recipes, pre- and post-training, two model families, and reduced data scales, and they are poorly captured by overall benchmark scores or by attention-flow diagnostics.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Debiasing one group shifts stereotypes onto others","Stereotype mitigation redistributes bias across groups","Preprocessing debiasing worsens scores for non-targets","Targeted debiasing induces side effects on unrelated groups","Fixing stereotypes for some raises them for others"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Directional shifts of StereoSet and CrowS-Pairs scores away from 50, measured once under a fixed seed on Wikipedia-derived data, are treated as a trustworthy signal of real stereotype redistribution rather than an artifact of benchmark makeup or detector noise.","fun_headline_variants_meta":{"raw":{"variants":["Debiasing one group shifts stereotypes onto others","Stereotype mitigation redistributes bias across groups","Preprocessing debiasing worsens scores for non-targets","Targeted debiasing induces side effects on unrelated groups","Fixing stereotypes for some raises them for others"]},"model":"grok-4.5","effort":"low","cost_usd":0.005142,"raw_usage":{"total_tokens":1393,"prompt_tokens":708,"num_sources_used":0,"completion_tokens":57,"cost_in_usd_ticks":51420000,"prompt_tokens_details":{"text_tokens":708,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":628,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":708,"tokens_out":57,"duration_ms":6301,"temperature":1.0,"reasoning_tokens":628,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T15:05:13.928109+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same three preprocessing recipes on the same Wikipedia snapshot with multiple random seeds and a second independent stereotype detector; if non-target score movements reverse sign or vanish while target improvements remain, the claimed side-effect phenomenon collapses.","supporting_citations":[],"review_version":1}