{"id":"ef61254a-a61e-4194-b56a-63c060e1d433","arxiv_id":"2505.08792","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The paper offers five theory-derived guidelines for intersectionality in ML and evaluates three existing efforts, finding all misalign with the framework's core tenets.","lead":"This paper proposes a preliminary framework of five guidelines for using intersectionality in machine learning pipelines, drawn from foundational Black feminist scholarship. The authors apply the framework to three existing ML papers and report that all three misalign with intersectionality's core tenets of power, relationality, and social justice.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The evaluative instrument is unvalidated: without a demonstration that the five guidelines faithfully operationalize the three C's, the Table 1 alignment ratings cannot carry the paper's normative conclusion.","rationale":"Good-faith reading: the paper is a theory-driven position piece, and its usefulness is as a structured proposal, not as a quantitative proof. The central empirical assertion, however, is evaluative: all three case studies 'misalign' with intersectionality. That assertion is only as strong as the validity of the guidelines used to detect misalignment. The reader's weakest-assumption diagnosis is correct: the theory-to-checklist translation is the load-bearing step. I considered two other candidates—selection bias (n=3, hand-picked) and the contested nature of intersectionality itself—but those are acknowledged limitations of a preliminary study and do not by themselves undercut the within-sample findings. The instrument problem is more fundamental: without criterion validity or reliability, even the within-sample ratings are not reproducible. The proposed check—a traceability audit plus expert judgment of entailment, followed by recomputation of Table 1—directly tests whether the guidelines are a faithful operationalization of the three C's. If ratings change under that audit, the paper's central claim reduces to a statement about the authors' preferred reading of intersectionality; if ratings survive, the concern is answered. Because the authors explicitly label the framework preliminary and the reader already assigned CONDITIONAL, no verdict change is needed.","tokens_in":20148,"tokens_out":6441,"duration_ms":62264,"concrete_test":"Perform a traceability audit of §3.2: create a table mapping each guideline and each subcriterion to specific passages in Crenshaw (1989; 1991), Combahee River Collective (1983), and Collins (2015; 2019). Have two scholars with intersectionality expertise independently judge each mapping as 'directly entailed' or 'not directly entailed' by the cited sources, with disagreements resolved by a third. Delete or reclassify any item not directly entailed, then recompute the Table 1 alignment ratings. If the overall ×/partial/aligned pattern changes for any case study, the Table 1 ratings are an artifact of the authors' interpretive choices rather than a faithful operationalization of the three C's.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central conclusion that CS1–CS3 misalign with intersectionality depends entirely on the five guidelines in §3.2 as the measuring instrument. Those guidelines are asserted to be 'derived from foundational intersectionality scholarship,' but the paper provides no derivation protocol, no mapping from individual criteria to specific passages in Crenshaw, Combahee, or Collins, and no independent validation. The subcriteria use highly interpretive predicates—'acknowledges the social justice facet,' 'contemplates cultural harm,' 'refrains from diluting social identities'—with no decision rules for when a paper satisfies them. The same researchers who designed the instrument then applied it, without inter-rater reliability or blinding, so the ×/partial/aligned codings in Table 1 are inherently vulnerable to confirmation bias. The authors' own §5.4 states the guidelines are preliminary and will be refined, conceding that the instrument is not yet established. If a different analyst applied the same guidelines to the same papers, there is no evidence the ratings would reproduce; if the guidelines themselves do not faithfully operationalize the three C's, the 'gaps and misalignments' reported are an artifact of the authors' interpretive choices rather than a property of the case studies.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a preliminary framework for applying intersectionality in machine learning pipelines, organized as five guidelines with eighteen subcriteria. The guidelines are presented as derived from the foundational scholarship of the Combahee River Collective, Kimberlé Crenshaw, and Patricia Hill Collins. The authors review 23 papers, select three case studies (Wang et al. 2022; Klumbytė et al. 2022; Nelson 2021), and rate each case study's alignment with the guidelines in Table 1. They conclude that all three case studies partially align with the foundations of intersectionality but exhibit gaps and misalignments in relationality, power, social justice, and ethical transparency. The paper closes with recommendations for future ML efforts, emphasizing agenda alignment, positionality, and interdisciplinary collaboration.","tokens_in":20296,"tokens_out":5884,"duration_ms":58226,"significance":"If the proposed framework were validated, it would make a useful contribution by pushing ML researchers to treat power, historical specificity, and social justice as first-order design considerations rather than as decorative citations. The authors deserve credit for grounding the discussion in foundational scholarship, for providing concrete subcriteria, for including a positionality statement in Section 5.2, and for being explicit in Section 5.4 that the guidelines are preliminary. However, as presented, the evaluative results do not yet support the paper's normative conclusions. The measurement instrument is not operationalized, the Table 1 ratings are not shown to be reproducible, and the case-study selection is not demonstrated to be representative. The paper's central claim therefore rests on an unvalidated instrument applied by its own designers.","major_comments":[{"comment":"The evaluation instrument is not sufficiently operationalized. Subcriteria such as \"acknowledges the social justice facet of intersectionality\" (Guideline 3.1), \"contemplates cultural harm caused by reinforcing existing power structures\" (Guideline 3.4), and \"refrains from diluting social identities for computational convenience\" (Guideline 4.1) require substantial interpretive judgment, yet no decision rules are provided. The paper does not report a coding protocol, inter-rater reliability, or blinding. Because the central conclusion that CS1-CS3 misalign with intersectionality is entirely mediated by Table 1, these ratings cannot bear the paper's claims without at least a demonstration that the ratings are reproducible.","section":"Section 3.2 and Table 1"},{"comment":"The guidelines lack a derivation protocol and external validation. The paper states that the guidelines were \"derived from foundational intersectionality scholarship,\" but it does not map each guideline or subcriterion to specific passages in Crenshaw, Combahee, or Collins, nor does it justify why these particular subcriteria constitute a faithful operationalization of the three C's. This creates a concrete risk of circularity: the same interpretive framework used to define \"correct\" intersectionality is then used as the benchmark for judging the case studies. The authors should provide a traceable mapping from each subcriterion to canonical scholarship and, ideally, an independent audit or member checking with intersectionality scholars.","section":"Section 3.2"},{"comment":"The case-study selection does not support the paper's generalizing language. The authors reviewed 23 papers but selected \"three case studies in particular that represented the existing efforts\" without stating the criteria for representativeness, and they then restricted the pool to papers that explicitly cite Crenshaw. This selection strategy biases both RQ1 (how intersectionality is being applied) and RQ2 (alignment with the theoretical foundations). The abstract and conclusion frame the findings as evidence about \"existing efforts\" applying intersectionality in ML, but a non-random sample of three cannot support that generalization. Either justify the sample's theoretical sufficiency or soften the claims to describe three illustrative cases.","section":"Section 3.1 and RQ1/RQ2"},{"comment":"There is no stated rule for aggregating subcriterion ratings into the guideline-level ratings, and the table contains inconsistent use of \"not applicable.\" For CS2, subcriteria 4.1, 4.2, and 4.4 are marked as not applicable, yet overall Guideline 4 is rated \"partially aligned\"; for CS1, subcriteria 4.1 and 4.2 are rated \"×\" while 4.3 and 4.4 are rated as partially aligned, yet the overall rating for Guideline 4 is \"×.\" A reproducible scoring rule (e.g., majority, minimum, or a weighted judgment rubric) is needed before Table 1 can be interpreted. Additionally, the narrative for CS1 Guideline 3.3 contains evidence that sounds like partial alignment (e.g., acknowledgment of underlying inequalities) but is coded as \"no alignment\"; this apparent contradiction should be resolved.","section":"Table 1 and Section 4.1"}],"minor_comments":[{"comment":"There are typographical errors: \"Comhabee River Collectice\" in Section 3.2 should be \"Combahee River Collective,\" and \"Cohambee\" in Section 5.1 should be \"Combahee.\"","section":"Section 3.2 and Section 5.1"},{"comment":"The table legend symbols (\"medbullet,\" \"LEFTCIRCLE,\" \"Circle\") appear to be LaTeX macros that did not render correctly in the submitted version; please ensure the symbols display as intended in the final PDF.","section":"Table 1"},{"comment":"The prose for CS1 Guideline 3.3 includes observations that seem to support partial alignment, yet the table records \"no alignment.\" Please revise the text so that the reasoning is internally consistent with the assigned rating.","section":"Section 4.1, CS1 Guideline 3.3"},{"comment":"The claim that \"machine learning pipelines have generally failed to explicitly consider intersecting identities\" is broad and is stated without a citation or systematic evidence; please either support or qualify this claim early in the paper.","section":"Section 2.3"},{"comment":"The limitation paragraph is clear, but it appears after the findings. Consider placing an abbreviated version of this limitation before Table 1 so that readers interpret the ratings with appropriate caution from the outset.","section":"Section 5.4"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the manuscript reads more like a position or vision piece than a completed empirical study. The core theoretical content is reasonable, but the evidentiary link between the proposed framework and the Table 1 ratings is the main weakness. If the venue welcomes preliminary frameworks, the requested revisions should be sufficient; the paper does not need to be rejected as long as the authors either strengthen the evaluation methodology or substantially soften the empirical claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things: this is a genuine attempt to build a practical bridge from intersectionality scholarship to ML pipeline design, and it is explicitly preliminary—the authors say so in the limitations. The main contribution is a set of five guidelines (relational analysis, social formations, historical/cultural specificity, feature engineering/statistical methods, ethical transparency) with subcriteria, derived from Crenshaw, Combahee, Collins, and others. Applying these to three ML papers that cite Crenshaw, they find persistent misalignments: the papers use intersectionality in motivation but not in methodology. That finding is plausible and, I think, correct in broad strokes.\n\nWhat it does well: it takes foundational theory seriously and tries to make it operational for technical audiences. The case-study discussions are careful and specific—for example, they show how CS1's reallocation of racial identities to \"closest neighbor\" compromises the mutually constitutive principle. The paper is also honest about its limitations and doesn't overclaim novelty; it positions the guidelines as a starting point.\n\nSoft spots: The stress-test note lands. The five guidelines are asserted to be \"derived from\" foundational scholarship, but there's no derivation protocol or mapping from each subcriterion to specific passages. Predicates like \"acknowledges the social justice facet\" and \"contemplates cultural harm\" are interpretive, and the same authors both designed the instrument and applied it, with no inter-rater reliability or blinding. The case-study selection is a convenience sample of 3 from 23, and representativeness is asserted, not shown. However, these are exactly the limitations the authors flag in §5.4; they explicitly say the guidelines will be refined. So the paper is better read as a well-argued position piece than a validated audit.\n\nThe circularity concern is real but mild: any critique of fidelity to a theory measures against the critic's reading of the theory. That's normal in critical theory work; it doesn't invalidate the exercise, but it does mean the ratings are interpretive, not objective.\n\nBottom line: This is a useful paper for fairness/HCI researchers who want to use intersectionality without treating it as a demographic checkbox. It deserves peer review, but a referee should push for a more transparent derivation of the guidelines, a larger and more systematic sample, and possibly an inter-rater reliability check or at least a clear statement that the ratings are the authors' interpretive judgment. I'd send it out.","headline":"A useful, honest, preliminary framework for intersectionality in ML, but the case-study evaluation rests on an unvalidated checklist that needs transparent operationalization.","tokens_in":20873,"tokens_out":2081,"would_cite":true,"duration_ms":20231,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An audit of three machine-learning efforts finds each misapplies intersectionality's core tenets.","keywords":["intersectionality","machine learning","algorithmic fairness","social justice","critical theory","positionality","case study","framework"],"falsifier":"A direct test: assemble a panel of intersectionality scholars, give them the three founding texts but not the five guidelines, and have them independently judge whether the same three ML projects honor intersectionality. If the panel's verdicts diverge substantially from the paper's ratings, the guidelines are not a faithful instrument; if they agree, the misalignment finding holds.","tokens_in":19875,"feed_emoji":"⚖️","tokens_out":6644,"duration_ms":61201,"temperature":0.7,"pith_summary":"This paper claims that machine-learning efforts that invoke intersectionality routinely fall short of the theory's founding commitments, and that the gap is structural rather than incidental: researchers cite Crenshaw, the Combahee River Collective, and Collins accurately, then operationalize intersectionality in ways that drop its core concerns of power, relationality, and social justice. To make that case, the authors distill foundational scholarship into five preliminary guidelines and use them to audit three published ML projects that explicitly ground themselves in intersectionality. The audit finds misalignments in all three: identities are often treated as separable or aggregatable, societal implications are disclaimed, and reflexive attention to researcher positionality is missing. If the guidelines are a faithful translation of the theory, the paper gives researchers a structured instrument for checking whether an ML pipeline actually honors intersectionality rather than merely naming it.","feed_headline":"Audit: ML papers citing intersectionality misapply it in every case","feed_subtitle":"A five-guideline rubric drawn from the theory's founders rates three machine-learning efforts—and all fall short on power and justice.","key_machinery":"The machinery is the five-guideline audit rubric built from foundational intersectionality scholarship, specifically the 'three C's': the Combahee River Collective's interlocking oppressions, Kimberlé Crenshaw's account of how law fails Black women at the race–gender intersection, and Patricia Hill Collins' critical social theory of intersectionality. Each guideline is unpacked into concrete subcriteria—for example, whether the project treats identities as mutually constructive, whether data-imbalance fixes compromise identities, whether the work acknowledges social justice and societal implications, whether statistical methods oversimplify lived experience, and whether researchers disclose positionality. The rubric does the argument's work by converting theory into a checkable standard, so the case-study verdicts in the paper's alignment table are the evidence for its claim.","core_discovery":"The central discovery is an operational diagnosis: when machine learning researchers adopt intersectionality, they adopt its vocabulary and its citations but not its analytic core. The paper's five-guideline framework—relational analysis, social formations of complex inequalities, historical and cultural specificity, feature engineering and statistical methods, and ethical considerations and transparency—is used to score three case studies. In all three, at least some guidelines receive partial or full alignment, yet overall the projects are rated misaligned on key tenets: one project learns about Black women by training on Black men and white women, collapsing mutually constitutive identities; another invokes intersectionality mainly in its introduction and does not connect its design workshops to power; the third maps word embeddings onto nineteenth-century identities without reckoning with its own interpretive position. The authors conclude that current practice defines intersectionality correctly but applies it inconsistently, and they position the five guidelines as a preliminary corrective.","pith_inferences":["One extension the paper leaves implicit: the rubric could be turned into a reviewer checklist for fairness venues, making 'misaligned with intersectionality' a citable reason to request revisions.","The findings suggest a deeper tension than the paper states—if relationality is non-negotiable, some purely quantitative pipeline designs may be incapable of satisfying it, and would need participatory or qualitative components.","A testable extension: have independent intersectionality scholars apply the original founding texts directly to the same three papers; high agreement with the paper's ratings would validate the rubric, and low agreement would localize where the theory-to-checklist translation loses fidelity.","The framework could be extended from three descriptive cases into a design tool, specifying for each pipeline stage which concrete practices satisfy each guideline."],"forward_implications":["An ML project that cites intersectionality now carries a burden of proof: it must show relational feature engineering, not subgroup aggregation or cross-group training that collapses identities.","Researchers who disclaim societal implications in their fairness work are, by this standard, misusing intersectionality even when they define it correctly.","The rubric gives reviewers and practitioners a concrete checklist for auditing pipeline decisions, from data labeling through evaluation metric choice.","Because all three audited projects failed or partially failed on ethical transparency and positionality, the paper implies that equitable ML work should include reflexive statements and interdisciplinary input.","The guidelines are preliminary, and the authors call for refinement through deeper analysis and a larger sample of case studies."],"supporting_citations":[{"why":"Supplies Collins' definitional account of intersectionality as a critical knowledge project, the theoretical base for the guidelines.","marker":"[28]"},{"why":"Supplies Collins' critical social theory framing, especially power and social justice, that the guidelines test for.","marker":"[30]"},{"why":"Crenshaw's mapping of the race–gender intersection anchors the relationality guideline.","marker":"[33]"},{"why":"Combahee River Collective statement supplies the interlocking-oppression tenet behind the social-inequalities guideline.","marker":"[26]"},{"why":"Bowleg supplies the core methodological warning against treating identities additively in quantitative work.","marker":"[16]"},{"why":"Bauer and Scheim provide quantitative intercategorical methods that the feature-engineering guideline draws on.","marker":"[10]"},{"why":"Hancock's argument that intersectionality is a critical theory, not a generalizable paradigm, justifies judging projects on power and justice.","marker":"[44]"},{"why":"First audited case study; its subgroup training design is the concrete target of the relationality critique.","marker":"[86]"},{"why":"Second audited case study; its workshop design is judged partial on social justice and weak on explicit power.","marker":"[51]"},{"why":"Third audited case study; its word-embedding analysis is judged misaligned on historical specificity and reflexivity.","marker":"[66]"}],"fun_headline_variants":["ML papers cite intersectionality but fail its core tenets","Intersectionality in ML: theory cited, practice misaligned","Audit: ML's intersectionality usage fails power and justice","Five guidelines show ML misapplies intersectionality routinely","ML research: intersectionality quoted, not followed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire analysis depends on the idea that the five guidelines really do capture what Crenshaw, the Combahee River Collective, and Collins meant by intersectionality; if that translation from theory to checklist is disputed, the verdicts on the case studies no longer follow.","fun_headline_variants_meta":{"raw":{"variants":["ML papers cite intersectionality but fail its core tenets","Intersectionality in ML: theory cited, practice misaligned","Audit: ML's intersectionality usage fails power and justice","Five guidelines show ML misapplies intersectionality routinely","ML research: intersectionality quoted, not followed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000212,"raw_usage":{"total_tokens":1394,"prompt_tokens":897,"completion_tokens":497,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":417}},"tokens_in":513,"tokens_out":497,"duration_ms":4601,"temperature":1.0,"reasoning_tokens":417,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:44:06.029951+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test: assemble a panel of intersectionality scholars, give them the three founding texts but not the five guidelines, and have them independently judge whether the same three ML projects honor intersectionality. If the panel's verdicts diverge substantially from the paper's ratings, the guidelines are not a faithful instrument; if they agree, the misalignment finding holds.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Collins' definitional account of intersectionality as a critical knowledge project, the theoretical base for the guidelines."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Collins' critical social theory framing, especially power and social justice, that the guidelines test for."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Crenshaw's mapping of the race–gender intersection anchors the relationality guideline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Combahee River Collective statement supplies the interlocking-oppression tenet behind the social-inequalities guideline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Bowleg supplies the core methodological warning against treating identities additively in quantitative work."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Bauer and Scheim provide quantitative intercategorical methods that the feature-engineering guideline draws on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Hancock's argument that intersectionality is a critical theory, not a generalizable paradigm, justifies judging projects on power and justice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"First audited case study; its subgroup training design is the concrete target of the relationality critique."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Second audited case study; its workshop design is judged partial on social justice and weak on explicit power."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Third audited case study; its word-embedding analysis is judged misaligned on historical specificity and reflexivity."}],"review_version":1}