{"id":"27f6375f-a6bf-4650-9635-a82100346619","arxiv_id":"1908.09635","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"This survey catalogs types of bias, fairness definitions, and mitigation strategies across ML domains.","lead":"This survey of over 140 papers organizes sources of algorithmic bias, definitions of fairness, and mitigation methods across machine learning, NLP, and deep learning. It provides a broad reference map for researchers and practitioners entering the fairness field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central taxonomy lacks a reproducible placement rule: Figure 2 arrow assignments are asserted as subjective judgment, and Table 1 mislabels fairness-through-unawareness as individual fairness.","rationale":"The paper is a survey, so its main value is organizational rather than novel technical results. The strongest claim is that the authors created a taxonomy of bias sources and fairness definitions and mapped the state of the art. That claim depends on the assignments in Figure 2 and Table 1 being checkable. Section 3.1 explicitly says the bias definitions were placed where the authors 'thought they were most effective,' and no formal criterion or mapping table is provided. Since several bias types have overlapping definitions, the placement is underdetermined. Table 1's labeling of fairness-through-unawareness as an individual fairness notion is a concrete instance of the risk: that definition only excludes explicit use of protected attributes, whereas individual fairness, as defined in the same paper via fairness through awareness, requires similar predictions for similar individuals. The reader's weakest assumption identified the same general issue, so the reader's CONDITIONAL verdict is appropriate. The concern is substantial enough to require a revision, but it does not invalidate the survey's usefulness as a reference; a clear placement rule and a corrected Table 1 would address it. I do not see an internal contradiction that forces rejection, and because there is no machine-checkable verification, the reproducibility of the taxonomy is the main lever for changing confidence.","tokens_in":27685,"tokens_out":5499,"duration_ms":56524,"concrete_test":"Publish the exact arrow assignment for each of the 23 bias types behind Figure 2, then have two independent annotators reconstruct the placement using only the definitions in Section 3.1 and their sources [99, 118, 8]. Pre-register a threshold such as Cohen's kappa >= 0.6. If agreement falls below the threshold, or if closely related pairs such as sampling/self-selection bias or representation/population bias are placed inconsistently, the taxonomy is not reproducible and needs an explicit, principled assignment rule.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central deliverable is a taxonomy: bias types placed on the data-algorithm-user feedback loop, fairness definitions grouped into individual vs. group, and mitigation methods categorized as pre/in/post-processing. The weakest load-bearing point is the feedback-loop taxonomy in Section 3.1 and Figure 2. The text says the authors grouped bias definitions 'on the arrows of the loop where we thought they were most effective,' but it gives no criterion for assigning a bias type to an arrow, and no table lists which of the 23 definitions goes where. The categories themselves overlap (e.g., sampling bias vs. self-selection bias, representation bias vs. population bias), so the assignment is underdetermined: different readers applying the same definitions could place them on different arrows. This is not just an aesthetic issue because the paper's strongest claim is that it provides a valid organization of the field. The same looseness appears in Table 1, where fairness-through-unawareness is labeled an individual fairness notion even though it only requires that protected attributes not be explicitly used; it does not require similar predictions for similar individuals, which is what individual fairness (Definition 4) demands. This is a concrete symptom that the taxonomy is a set of subjective judgments rather than a derivable classification. The survey remains useful as a reference, but its central organizational claim is not independently checkable from the text.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey maps the landscape of bias and fairness in machine learning. It begins with real-world cases of algorithmic unfairness (COMPAS, ad delivery, facial recognition), then catalogs twenty-three types of bias in data and algorithms, proposes a feedback-loop organization of these bias types in Figure 2, reviews ten fairness definitions and classifies them as group or individual notions, summarizes mitigation approaches under pre-/in-/post-processing categories, and surveys domain-specific work in classification, regression, PCA, community detection, causal inference, representation learning, and natural language processing. It also lists commonly used fairness datasets and discusses open challenges such as synthesizing fairness definitions and moving from equality to equity. The paper offers no new experimental results but aims to provide a broad reference taxonomy of the field.","tokens_in":27904,"tokens_out":3388,"duration_ms":38974,"significance":"If its organizational claims held, the survey would be a useful entry point for researchers: it aggregates a large literature, connects bias types to concrete examples, gives a table of protected attributes, and catalogs datasets and toolkits. The paper is honest about the breadth of the area and identifies underexplored directions. Its main weaknesses are that the central taxonomy is not reproducible from the text and that at least one fairness classification in Table 1 is incorrect. The survey remains potentially valuable as a reference, but its central organizational claim needs to be made checkable before the taxonomy can be relied upon.","major_comments":[{"comment":"The central organizational claim of the survey is not independently checkable. After listing the bias types, the text says the authors \"grouped these definitions on the arrows of the loop where we thought they were most effective,\" but no placement criterion is given and no table or list maps each of the twenty-three definitions to a specific arrow of Figure 2. Because several definitions overlap (e.g., sampling bias vs. self-selection bias, representation bias vs. population bias), different readers could plausibly place the same definition on different arrows. Since the paper's stated contribution is \"a taxonomy for fairness definitions\" and a valid organization of bias sources, the assignment rule and the full mapping should be made explicit.","section":"Section 3.1 / Figure 2"},{"comment":"Table 1 marks fairness through unawareness as an individual fairness notion, but this is inconsistent with Definition 5 in the same section. Fairness through unawareness only requires that protected attributes not be explicitly used in the decision, and it does not require similar predictions for similar individuals, which is what fairness through awareness (Definition 4) and individual fairness demand. This misclassification is a concrete symptom that the fairness taxonomy is a set of subjective judgments rather than a derivable classification. The table should be corrected, or the entry should be explicitly classified as a process-based notion separate from individual and group fairness.","section":"Section 4.2 / Table 1"},{"comment":"The reproduced formal definitions in the fair regression subsection are garbled to the point that the reader cannot verify the claims about the three fairness penalties. The displayed formulas for f1, f2, and f3 contain missing or misplaced parentheses and unclear summation scopes. Since a survey's reliability depends on the fidelity of its reproduced definitions, these equations should be typeset cleanly and checked against the cited source.","section":"Section 5.2.2"}],"minor_comments":[{"comment":"The statement of counterfactual fairness contains a typo (\"(or all y\") and the probability notation is incomplete; it should be rewritten to match the cited source.","section":"Section 4.2 / Definition 8"},{"comment":"Equation 2 uses a Cyrillic \"loд\" instead of \"log\", and the conditioning notation P(w|f), P(w|m) is not defined precisely enough to distinguish female and male word sets.","section":"Section 5.4.3 / Equation 2"},{"comment":"The phrase \"in [41], authors also argued\" and several similar constructions could be smoothed; more importantly, the article sometimes alternates between numbered citations and inline URLs, which should be unified.","section":"Section 2"},{"comment":"The heatmap in Figure 7 is mentioned only briefly; the reader is not told how the entries were scored, which fairness definitions were used, or how the domain categorization was performed.","section":"Section 6.2 / Figure 7"},{"comment":"The list of bias types is useful, but several definitions draw on a non-academic web source (footnote 4); citing a peer-reviewed taxonomy for these entries would strengthen the survey.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The survey is a broad and potentially useful reference, but its central taxonomy needs to become reproducible before publication. Many examples and several cited works come from the authors' own prior papers, especially in Sections 3, 5.2.5, and 6.1; this is not improper, but the selection balance should be checked. The garbled equations and the Table 1 misclassification are fixable within the manuscript's scope, so I am not recommending rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a survey of bias and fairness in ML, and the honest summary is that it is a convenient compilation rather than a research contribution. There are no new experiments, derivations, or mechanisms. The value is organizational: bias types, fairness definitions, mitigation families, domain-specific work, and datasets all collected in one place. As of 2019 this was a useful snapshot, and the paper has been widely cited for exactly that. The domain-coverage table and the dataset list are genuinely handy for someone entering the field.\n\nThe authors do try for a conceptual contribution: they place bias definitions on a data-algorithm-user feedback loop (Figure 2), arguing that biases are intertwined rather than cleanly separable. This framing is reasonable, and it is honestly presented as an editorial choice rather than a theorem.\n\nThe soft spots are real but not fatal. First, the arrow-placement in Figure 2 is subjective. The text says the authors 'grouped these definitions on the arrows of the loop where we thought they were most effective,' but it gives no criterion and no table showing which definition goes on which arrow. Since the categories overlap (sampling vs. self-selection; representation vs. population), the placement is underdetermined, and a reader cannot reproduce or check the taxonomy. If the taxonomy is meant to be a deliverable, it needs a placement rule or at least a complete mapping.\n\nSecond, Table 1 mislabels Fairness through Unawareness as an individual fairness notion. That definition only requires that protected attributes not be explicitly used; it does not require similar predictions for similar individuals, which is what individual fairness (Definition 4) demands. That is a concrete error, not a matter of taste, and it should be corrected.\n\nThere are also scattered typographical issues, especially the garbled regression penalties in Section 5.2.2. Minor, but visible.\n\nThe survey's central claim—that the field contains these bias sources, fairness definitions, and mitigation strategies—is well supported by the citations. Self-citations in a few sections (community detection, invariant representations, public perceptions) don't carry the weight of the survey, and the core does not depend on them.\n\nWho is this for? Someone looking for an entry point into the fairness literature. Not for an expert seeking depth or rigor. A serious referee should look at it, but acceptance should be conditional on fixing the taxonomy and the table error. I would take it for peer review and expect revisions.","headline":"Useful reference survey, not a research contribution; the feedback-loop taxonomy needs a reproducible placement rule and Table 1 contains a concrete labeling error.","tokens_in":28426,"tokens_out":3562,"would_cite":true,"duration_ms":32865,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper maps machine-learning bias onto a data-algorithm-user feedback loop and taxonomizes fairness definitions, arguing that unfairness enters and persists through the whole cycle.","keywords":["algorithmic bias","fairness","machine learning","bias taxonomy","fairness definitions","discrimination","bias mitigation","AI ethics"],"falsifier":"Ask an independent panel to place the 23 bias types from Section 3.1 onto the three-node feedback loop in Figure 2; if the placements diverge sharply across raters, the taxonomy's organizing assumption fails. Equally, a systematic sweep of the fairness literature that turns up a widely used definition absent from Table 1 would falsify the survey's claim of coverage.","tokens_in":27469,"feed_emoji":"⚖️","tokens_out":5108,"duration_ms":52135,"temperature":0.7,"pith_summary":"This paper is a survey that sets out to give researchers a usable map of where unfairness in machine learning comes from and what can be done about it. It collects real-world cases, catalogs bias types, and organizes them on a feedback loop connecting data, algorithms, and user interaction. It then gathers formal fairness definitions into a taxonomy and sorts mitigation methods by whether they intervene before, during, or after training. A sympathetic reader would take the main contribution to be an organizing scheme: bias is not a single defect in data or code, but a cycle that can enter and reinforce itself at several points.","feed_headline":"AI bias flows through a data-algorithm-user loop","feed_subtitle":"A survey maps 23 bias types and 10 fairness definitions to guide fairer machine learning.","key_machinery":"The load-bearing device is a three-node feedback loop: data, algorithm, and user interaction, with bias types placed on the arrows where the authors judge them most active. Around this loop the paper organizes a second device: a taxonomy of fairness definitions split into group and individual fairness, plus mitigation strategies categorized as pre-processing, in-processing, and post-processing. Together these devices do the argument's work: they turn scattered observations about biased systems into a map that shows where bias enters, which fairness target each definition addresses, and where existing methods intervene.","core_discovery":"On the paper's own terms, the central discovery is a structured inventory: there is no one bias, but a family of distinct bias mechanisms—historical, representation, measurement, aggregation, sampling, behavioral, linking, popularity, algorithmic, and others—and no one fairness notion, but competing formal definitions such as demographic parity, equalized odds, equal opportunity, fairness through awareness and unawareness, and counterfactual fairness. The paper claims that these are not independent: biases are coupled through the feedback loop in which a model's decisions shape future data and user behavior, so a fair system must be examined as a cycle rather than a static artifact. On this view, whether a system is fair cannot be read off a single metric; it depends on the context, the protected attributes, and the stage at which intervention is possible.","pith_inferences":["Editorial inference: the feedback-loop framing predicts that fairness interventions will decay over time unless they account for future data collection; an A/B test comparing a debiased model retrained on its own outputs against one retrained on a fixed dataset would test this.","Editorial inference: the taxonomy invites a matching exercise between bias types and fairness definitions—for example, measurement bias is most naturally addressed by equalized-odds-style criteria—which could turn the survey's categories into an actionable checklist for model audits.","Editorial inference: the same loop applies to generative AI and recommendation systems, where the user interaction arrow is strongest; the survey's categories could be extended to cover biases introduced by human feedback, which the paper mentions only indirectly."],"forward_implications":["Because biases are coupled in a feedback loop, removing bias from training data alone will not guarantee fair decisions if the algorithm or the user interface re-introduces it.","No single fairness definition can serve every application: equalized odds, demographic parity, and calibration can be mutually incompatible, so fairness must be chosen per context.","The pre/in/post-processing distinction gives practitioners a practical way to select mitigation methods based on what they are allowed to modify.","The survey's mapping of domains to fairness definitions points to under-studied areas such as subgroup-level fairness in community detection."],"supporting_citations":[{"why":"Supplies the source-of-bias categorizations and examples for historical, representation, measurement, evaluation, and aggregation bias.","marker":"[118]"},{"why":"Supplies the list of social-data bias types, including population, behavioral, linking, temporal, and content production bias.","marker":"[99]"},{"why":"Provides the fairness definitions that the survey reiterates and taxonomizes.","marker":"[123]"},{"why":"Defines equalized odds and equal opportunity, which anchor the group-fairness side of the taxonomy.","marker":"[55]"},{"why":"Defines counterfactual fairness and relates fairness-through-awareness ideas to machine learning.","marker":"[73]"},{"why":"Provides fairness through awareness and demographic parity, core individual and group notions in the taxonomy.","marker":"[43]"},{"why":"Supplies the web and user-interaction bias types and motivates the feedback-loop perspective.","marker":"[8]"},{"why":"Contributes emergent bias and the broader framing of bias in computer systems.","marker":"[46]"},{"why":"Demonstrates the incompatibility of calibration with balancing positive and negative classes, motivating context-specific fairness choices.","marker":"[70]"}],"fun_headline_variants":["Fair ML isn't a metric—it's a loop","Survey maps 23 bias types and 10 fairness definitions","Bias in AI isn't one problem—it's a family","AI bias: a feedback loop, not a static bug","From biased data to fair systems: a survey's map"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's claim to be a useful map depends on the authors' judgment in assigning each bias type to one arrow of the data-algorithm-user loop and in choosing which papers to include; if these assignments or selections are unrepresentative, the taxonomy could mislead rather than organize.","fun_headline_variants_meta":{"raw":{"variants":["Fair ML isn't a metric—it's a loop","Survey maps 23 bias types and 10 fairness definitions","Bias in AI isn't one problem—it's a family","AI bias: a feedback loop, not a static bug","From biased data to fair systems: a survey's map"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000622,"raw_usage":{"total_tokens":2878,"prompt_tokens":935,"completion_tokens":1943,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":1860}},"tokens_in":551,"tokens_out":1943,"duration_ms":15993,"temperature":1.0,"reasoning_tokens":1860,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:33:10.665870+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask an independent panel to place the 23 bias types from Section 3.1 onto the three-node feedback loop in Figure 2; if the placements diverge sharply across raters, the taxonomy's organizing assumption fails. Equally, a systematic sweep of the fairness literature that turns up a widely used definition absent from Table 1 would falsify the survey's claim of coverage.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the list of social-data bias types, including population, behavioral, linking, temporal, and content production bias."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the fairness definitions that the survey reiterates and taxonomizes."}],"review_version":1}