Pith. sign in

REVIEW 3 major objections 5 minor 77 references

Underrepresentation, Label Bias, and Proxies: Towards Data Bias Profiles for the EU AI Act and Beyond

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Underrepresentation of vulnerable groups in training data is a weaker driver of algorithmic discrimination than label bias and proxy features, and this paper introduces a Data Bias Profile to quantify each.

desk verdict Solid label-bias and proxy results, but the underrepresentation claim is only tested under uniform random subsampling and is undercut by the paper's own Adult (gender) numbers. read the letter →

arxiv 2507.08866 v1 pith:S2CRHZ2O submitted 2025-07-09 cs.LG cs.CYstat.ML

classification cs.LGcs.CYstat.ML
keywords algorithmicfairnessdatabiasunderrepresentationlabelproxyfeaturesdetectionProfileEUAIAct
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to rank the data properties that actually drive algorithmic discrimination rather than take them on faith. It injects controlled doses of three biases — underrepresentation of a disadvantaged group, systematic label corruption against that group, and proxy features carrying information about the protected attribute — into five tabular and two medical imaging datasets, across several model types and three fairness metrics. The paper claims that underrepresentation in training data is overemphasized in fairness research, while label bias, amplified by strong proxies, is the more critical driver of unequal outcomes. To turn this into practice, it proposes detection statistics that need no unbiased reference data and packages them into a Data Bias Profile, a quantitative summary meant to guide dataset documentation and fairness interventions under anti-discrimination regulation.

What carries the argument

The carrying object is the Data Bias Profile (DBP), a quantitative summary that records three bias signals computed without access to an unbiased reference set: the Representation Difference, the difference between advantaged and disadvantaged group prevalence; the Separation Difference, the average of two cross-group AUC gaps measuring how much harder positive examples of the disadvantaged group are to rank than those of the advantaged group; and the proxy factor sAUC, the AUC of a classifier that attempts to predict the protected attribute from non-protected features. The paper shows each statistic responds mainly to its own injected bias, and that the profile's pattern anticipates both the level of model unfairness and whether a proxy-removal intervention will help.

What would settle it

Run the same bias-injection study with underrepresentation implemented as removal of a structurally defined subpopulation of the disadvantaged group, such as all instances sharing an intersectional feature combination, keeping test sets unbiased; if equal-opportunity gaps rise sharply with the underrepresentation factor, the paper's relative ranking of the three biases fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that the relative importance of data biases for algorithmic discrimination is inverted relative to conventional wisdom: removing up to 100% of a disadvantaged group from the training set leaves equal-opportunity gaps roughly constant in most settings, whereas flipping a share of that group's positive labels produces steep, consistent increases in unfairness, especially in datasets where non-protected features strongly predict the protected attribute. The authors interpret this as evidence that representation-driven interventions are overvalued and that label curation and proxy management deserve priority. To support this claim they evaluate models on unbiased test sets, inject biases only in training and validation, and show that label bias can be strong enough that including the disadvantaged group without fixing labels makes outcomes worse for that group.

Load-bearing premise

The claim that underrepresentation is overemphasized assumes it takes the form of uniform random removal of disadvantaged instances; underrepresentation that removes entire feature regions or intersectional subgroups could hurt the group much more.

Editorial extensions

If this is right

  • Dataset balance should not be treated as the primary fairness fix; scarce annotated data from vulnerable groups is better spent on evaluation than on training.
  • Label quality is the first thing to audit: even a 20% flip of disadvantaged-group positive labels can widen equal-opportunity gaps significantly.
  • Proxy strength should guide intervention choice; removing features correlated with the protected attribute helps when sAUC is high and does little when it is low.
  • Including more disadvantaged-group samples without cleaning their labels can actively worsen outcomes for them, so representation efforts and label repair must go together.
  • For deployment documentation, a DBP can flag which bias signals a dataset carries and which fairness-enhancing intervention is likely to pay off.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If underrepresentation were imposed by deleting structured subpopulations rather than uniform random subsampling, the overemphasis conclusion could flip; testing that variant is a direct extension of the injection protocol.
  • The label-bias results reframe label noise as a fairness problem; the paper's injection protocol could serve as a benchmark for group-dependent label-noise correction methods.
  • A DBP with thresholds and multi-group support would be a stronger compliance tool for anti-discrimination audits; the paper explicitly leaves thresholds and multi-group extensions as open work.
  • The finding that label bias can make representation harmful suggests that data collection and labeling efforts should be evaluated jointly, not as separate pipeline stages.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper studies three types of data bias—underrepresentation, label bias, and proxies—by injecting them into training and validation sets of seven datasets, then measuring fairness on unbiased test sets with three fairness metrics and several models. It reports that underrepresentation has only a minor effect on discrimination, while label bias has a large effect that is amplified by strong proxies. The paper then proposes three detection measures (RD, SD, sAUC) and combines them into a preliminary Data Bias Profile (DBP), demonstrated on popular fairness datasets as a proof of concept for documenting bias, predicting discrimination risk, and selecting fairness interventions. The discussion connects these findings to data governance requirements under the EU AI Act.

Significance. If the empirical ranking of biases is correct, this work would usefully reorder fairness practice toward label curation and proxy management rather than representation targets alone, and the DBP is a promising quantitative complement to qualitative dataset documentation. The paper has concrete strengths: a broad experimental matrix (seven datasets, multiple model families, three fairness metrics, ten repetitions with significance tests), evaluation on unbiased test sets, clearly specified injection protocols, and an explicit limitation discussion. The label-bias and proxy-amplification findings are well supported by the reported results. However, the headline underrepresentation claim and the validation of the DBP detection measures are not yet established at the level of generality claimed in the abstract and Section 6; both need substantial qualification or additional experiments.

major comments (3)
  1. [Section 4.2, Eq. (1); Section 4.3, Table 3; Appendix B.1, Tables 9 and 12] The central claim that underrepresentation in training is overemphasized is inferred entirely from an injection mechanism that removes disadvantaged-group points uniformly at random. Real-world underrepresentation often removes structured subsets (feature subspaces, intersectional strata), which shifts P(x|y,s=d) rather than only reducing sample size, and this case is never tested. Even under the tested mechanism, the paper's own results do not uniformly show a minor impact: Adult (gender) EO rises from 0.08±0.02 to 0.21±0.07 for LR, and the appendix shows larger jumps for RF (0.10 to 0.30) and SVC (0.08 to 0.25), with PQP also significantly worsening. The Limitations section lists other bias types and binary attributes but does not acknowledge this operationalization gap. The abstract and Section 6 should either restrict the conclusion to uniform random subsampling or be backed by structured-missingness experiments.
  2. [Section 5.1, Eqs. (10), (13), (14); Section 5.2, Figure 4] The bias detection measures are validated on biases injected through the same mathematical quantities they measure: RD is exactly the prevalence gap manipulated by Eq. (1), sAUC is exactly the AUC of the proxy classifier used to define proxy strength in Eq. (3), and SD is an AUC-based separability gap that label flipping directly changes. The diagonal responses in Figure 4 are therefore partly by construction, so the experiments demonstrate internal consistency rather than the ability to detect real-world biases of these types. Since the practical value of the DBP depends on detection validity, the paper should provide an external validation or test the measures under bias mechanisms not defined by the same formulas, such as structured underrepresentation or naturally occurring group-dependent label noise.
  3. [Section 4.4, Table 5; Section 6] The recommendation that including disadvantaged groups in training can be harmful under weak label bias goes beyond the evidence. The ΔEO values for f=0.2 are negative for most datasets, but several carry standard errors that overlap zero (e.g., Adult-gender -0.01±0.08, Crime -0.11±0.17, Folktables -0.01±0.06, German -0.09±0.17), and NIH and Compas show positive values. The interaction may be real, but the current wording ('hastily adding disadvantaged groups ... can cause more harm than good') is too strong for these point estimates. Please temper the claim or report the uncertainty more prominently.
minor comments (5)
  1. [Table 2 and Section 4.2, Eq. (1)] The notation table defines u = r - 1, while the text and Eq. (1) define u = 1 - r as the underrepresentation factor. Please correct the table to avoid a sign inconsistency.
  2. [Section 4.4, paragraph on the joint effect] The sentence 'as confirmed by the first column of Table 4 (f = 0)' appears to be a reference error; the relevant ΔEO values are in Table 5, column f = 0, not in Table 4.
  3. [Section 5.2, Footnote 5 and Figure 4] Replacing u=1 and f=1 with 0.95 for the detection experiments is reasonable, but the figure axes and the appendix figures still display 1 at the maximum. Please align the axis labels with the actual values used.
  4. [Figure 2 caption and Figure 5] There are minor typos: 'disadvataged' in the Figure 2 caption should be 'disadvantaged', and 'Folkstables' in Figure 5 should be 'Folktables'.
  5. [Appendix A] The NIH disease list contains 'mas', which should likely be 'mass', and the COMPAS description spells 'ProPulica' instead of 'ProPublica'. Please fix these typos.

Circularity Check

3 steps flagged · score 4.0 of 10

Bias-detection validation is partially self-definitional (RD exactly; sAUC and SD largely), while the main underrepresentation/label-bias effect study is empirical and independent.

  1. self definitional [Section 5.1 Eq. (10) and Section 5.2, Figure 4]
    "RD(σ) = |σa| − |σd| / |σ| = Pr σ (s = a) − Pr σ (s = d) (10) ... underrepresentation is suitably captured by RD increasing linearly in the first column"

    Eq. (1) defines underrepresentation as Prσ′(s=d)=r·Prσ(s=d), i.e., only the prevalence gap changes. RD(σ) is exactly that prevalence gap. Therefore RD's linear response to the underrepresentation injection is an algebraic identity: RD(σ′) = (1−Prσ(s=d)) − r·Prσ(s=d). No classifier, label information, or external signal is involved; the detection curve is determined by the definition of the injection, so calling it evidence of detection is circular.

  2. self definitional [Section 4.2 Eqs. (3)-(4) and Section 5.2]
    "We quantify the strength of proxies as their joint ability to predict sensitive attributes. We train a classifier ˆs = h(x) to estimate the protected attribute s and we compute its AUC to measure the strength of proxies. ... An additive protocol adds to the non-sensitive variables X a new feature correlated with sensitive variables xnew = s + v, v ∼ N(0, std2)"

    The proxy injection adds a feature equal to the sensitive attribute plus noise, and the proxy detector sAUC is the AUC of a classifier trained to predict exactly that sensitive attribute from the augmented feature space. As std→0, xnew→s, so sAUC is forced toward 1 by construction. Thus the diagonal 'detection' of proxies in Figure 4 largely restates the injection mechanism rather than providing independent confirmation that the measure tracks an external proxy phenomenon.

1 more flagged steps
  1. self definitional [Section 4.2 Eq. (2), Section 5.1 Eq. (13), and Section 5.2]
    "For label bias, we selectively flip labels. We let f ∈ (0, 1) indicate the proportion of positive instances (y = ⊕) from the vulnerable group whose label is flipped to negative (y = ⊖) ... SD(σ) = ∆xAUCσ + ∆wAUCσ / 2 ... We expect label bias to worsen the separability for the disadvantaged group and therefore yield high values of SD."

    Label-bias injection directly flips group-d positives to negatives, which is exactly the operation that destroys the ranking separability of group-d positives relative to negatives. SD is defined as a difference of cross-group and within-group AUC separability. The injected flip therefore directly lowers the quantities SD is built from, making its increase with f largely entailed by the construction. Some empirical content remains because SD is computed from a trained classifier, but the diagonal response is primarily a restatement of the injection in ranking terms.

full rationale

The main effect analysis in Section 4 is not circular: biases are injected into training data through explicit mechanisms (Eqs. 1-5) and fairness is evaluated on unbiased test sets, so the conclusion that label bias is more critical than underrepresentation is an empirical, falsifiable result that could have gone either way. No fitted parameter is relabeled as a prediction, and no load-bearing self-citation chain is used; the authors' own prior work appears only in contextual citations. The partial circularity lies in Section 5's validation of the detection measures. RD is exactly the prevalence gap manipulated by Eq. (1), so its diagonal response is an algebraic identity. sAUC and SD are validated with injections defined through the same constructs, so their diagonal responses are largely by construction. Because these measures feed the Data Bias Profile, the DBP's detection claims inherit some definitional character, though the DBP case study is qualitative and not a fitted predictor. The uniform-subsampling operationalization of underrepresentation is an external-validity limitation, not a circularity, and does not affect the score. Overall, the central fairness finding stands on independent evidence; the detection-validation claim is partially circular.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters: the paper is an empirical study; u, f, and proxy correlation values are experimental injection levels, not fitted constants. The axioms above are the load-bearing assumptions for the experimental protocol and conclusions.

assumptions (5)
  • domain assumption The original datasets, before bias injection, provide an unbiased test set for the target population.
    Used throughout Section 4; if the original labels or group distributions are biased relative to the real world, the fairness measurements are relative to a biased ground truth, which the authors acknowledge in the Limitations.
  • ad hoc to paper Underrepresentation is modeled by uniform random subsampling of the disadvantaged group.
    Section 4.2, Eq. 1; this operationalization does not cover missing feature subspaces or intersectional absence.
  • ad hoc to paper Proxy strength equals the AUC of a classifier trained to predict the sensitive attribute from non-sensitive features (sAUC).
    Section 4.5, Eq. 3; this reduces a complex construct to one predictability score.
  • ad hoc to paper The three biases can be injected and detected independently.
    Sections 4.2 and 5.2; in real data the biases co-occur and interact.
  • domain assumption Equal opportunity, demographic parity, and predictive quality parity are adequate measures of algorithmic discrimination.
    Section 4.2, Eqs. 7-9; these are standard but contested fairness definitions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Underrepresentation, Label Bias, and Proxies: Towards Data Bias Profiles for the EU AI Act and Beyond." pith.science (2026). https://pith.science/paper/S2CRHZ2O

@misc{pith2026250708866,
  author       = {Pith},
  title        = {Pith review of: Underrepresentation, Label Bias, and Proxies: Towards Data Bias Profiles for the EU AI Act and Beyond},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S2CRHZ2O}},
  note         = {Machine review of arXiv:2507.08866}
}
read the original abstract

Undesirable biases encoded in the data are key drivers of algorithmic discrimination. Their importance is widely recognized in the algorithmic fairness literature, as well as legislation and standards on anti-discrimination in AI. Despite this recognition, data biases remain understudied, hindering the development of computational best practices for their detection and mitigation. In this work, we present three common data biases and study their individual and joint effect on algorithmic discrimination across a variety of datasets, models, and fairness measures. We find that underrepresentation of vulnerable populations in training sets is less conducive to discrimination than conventionally affirmed, while combinations of proxies and label bias can be far more critical. Consequently, we develop dedicated mechanisms to detect specific types of bias, and combine them into a preliminary construct we refer to as the Data Bias Profile (DBP). This initial formulation serves as a proof of concept for how different bias signals can be systematically documented. Through a case study with popular fairness datasets, we demonstrate the effectiveness of the DBP in predicting the risk of discriminatory outcomes and the utility of fairness-enhancing interventions. Overall, this article bridges algorithmic fairness research and anti-discrimination policy through a data-centric lens.

Figures

Figures reproduced from arXiv: 2507.08866 by the authors.

Figure 1
Figure 1. Large underrepresentation induces minor variations in the True positive rates (TPR) of both groups. Boxplots represent the TPR of the advantaged (s = a) and disadvantaged group (s = d), as u varies. In the experiments below, we measure algorithmic fairness according to these metrics as we inject controlled biases the training sets. We present results for logistic regression (on tabular datasets) and equal opportunit… view at source ↗
Figure 2
Figure 2. Label bias induces sizable variations in groupwise true positive rates (TPR); the disadvantaged group is especially affected. Boxplots representing the TPR on the advantaged and disadvantaged group (y axis), as the percentage of disadvataged group items with flipped labels increases in the training set (x axis). Broadly speaking, we distinguish two categories of datasets based on the effect of f on the TPR of the ad… view at source ↗
Figure 3
Figure 3. Proxies exacerbate the risk of algorithmic discrimination caused by label bias. EO (y axis) increases with label bias f (x axis). This effect is mediated by proxies: weaker proxies (lower sAUC) correspond to a lower slope and a weaker effect of label bias on fairness. the model’s reliance on sensitive information. These findings highlight the critical role of proxy strength in exacerbating label bias and influencing… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: The proposed measures capture specific types of bias. Bias detection on Folktables and NIH. Columns correspond to three bias injection mechanisms; rows correspond to bias detection measures. Measures vary when the corresponding bias increases (diagonal) and remain rela…
Figure 5
Figure 5. Figure 5: Data Bias Profiles hint at the risk of algorithmic discrimination and effectiveness of fairness intervention. DBP of Adult (left) and Folktables (center); on the right, model fairness summarized by demographic parity (x axis) and equality of opportunity (y). DBP highli…
Figure 6
Figure 6. Figure 6: Results of the bias detection methods on Adult (gender). The columns represent the three different scenarios [PITH_FULL_IMAGE:figures/full_fig_p025_6.png]
Figure 7
Figure 7. Figure 7: Results of the bias detection methods on Adult (marital status). The columns represent the three different [PITH_FULL_IMAGE:figures/full_fig_p026_7.png]
Figure 8
Figure 8. Figure 8: Results of the bias detection methods on Compas. The columns represent the three different scenarios while [PITH_FULL_IMAGE:figures/full_fig_p026_8.png]
Figure 9
Figure 9. Figure 9: Results of the bias detection methods on Crime. The columns represent the three different scenarios while the [PITH_FULL_IMAGE:figures/full_fig_p027_9.png]
Figure 10
Figure 10. Figure 10: Results of the bias detection methods on German. The columns represent the three different scenarios while [PITH_FULL_IMAGE:figures/full_fig_p027_10.png]
Figure 11
Figure 11. Figure 11: Results of the bias detection methods on Fitzpatrick17k. The columns represent the three different scenarios [PITH_FULL_IMAGE:figures/full_fig_p028_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

77 extracted references · 45 canonical work pages

  1. [1]

    Jos \' e M. \' A lvarez, Alejandra Bringas Colmenarejo, Alaa Elobaid, Simone Fabbrizzi, Miriam Fahimi, Antonio Ferrara, Siamak Ghodsi, Carlos Mougan, Ioanna Papageorgiou, Paula Reyero Lobo, Mayra Russo, Kristen M. Scott, Laura State, Xuan Zhao, and Salvatore Ruggieri. Policy advice and best practices on bias and fairness in AI . Ethics Inf. Technol., 26 0...

  2. [2]

    Machine bias

    Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner. Machine bias. ProPublica, 2016. URL https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing

  3. [3]

    Dissecting racial bias in an algorithm used to manage the health of populations

    Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366 0 (6464): 0 447--453, 2019

  4. [5]

    Towards a standard for identifying and managing bias in artificial intelligence

    Reva Schwartz, Apostol Vassilev, Kristen Greene, Lori Perine, Andrew Burt, and Patrick Hall. Towards a standard for identifying and managing bias in artificial intelligence. US Department of Commerce, National Institute of Standards and Technology, 2022

  5. [6]

    Artificial intelligence act

    European Parliament. Artificial intelligence act. https://www.europarl.europa.eu/doceo/document/TA-9-2024-0138_EN.pdf, 2024

  6. [7]

    Information technology — artificial intelligence (ai) — bias in ai systems and ai aided decision making, 2021

    ISO. Information technology — artificial intelligence (ai) — bias in ai systems and ai aided decision making, 2021. https://www.iso.org/standard/77607.html

  7. [8]

    Gillis, Vitaly Meursault, and Berk Ustun

    Talia B. Gillis, Vitaly Meursault, and Berk Ustun. Operationalizing the search for less discriminatory alternatives in fair lending. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT 2024, Rio de Janeiro, Brazil, June 3-6, 2024 , pages 377--387. ACM , 2024. doi:10.1145/3630106.3658912. URL https://doi.org/10.1145/3630106.3658912

  8. [9]

    Fairness and bias in algorithmic hiring

    Alessandro Fabris, Nina Baranowska, Matthew J Dennis, David Graus, Philipp Hacker, Jorge Saldivar, Frederik Zuiderveen Borgesius, and Asia J Biega. Fairness and bias in algorithmic hiring. ACM Transactions on Intelligent Systems and Technology, 2024. URL https://doi.org/10.1145/3696457

Show all 77 references
  1. [10]

    Baker and Aaron Hawn

    Ryan S. Baker and Aaron Hawn. Algorithmic bias in education. Int. J. Artif. Intell. Educ., 32 0 (4): 0 1052--1092, 2022. doi:10.1007/S40593-021-00285-9. URL https://doi.org/10.1007/s40593-021-00285-9

  2. [11]

    A data quality approach to the identification of discrimination risk in automated decision making systems

    Antonio Vetr \` o , Marco Torchiano, and Mariachiara Mecati. A data quality approach to the identification of discrimination risk in automated decision making systems. Gov. Inf. Q., 38 0 (4): 0 101619, 2021. doi:10.1016/J.GIQ.2021.101619. URL https://doi.org/10.1016/j.giq.2021.101619

  3. [12]

    Properties of fairness measures in the context of varying class imbalance and protected group ratios

    Dariusz Brzezinski, Julia Stachowiak, Jerzy Stefanowski, Izabela Szczech, Robert Susmaga, Sofya Aksenyuk, Uladzimir Ivashka, and Oleksandr Yasinskyi. Properties of fairness measures in the context of varying class imbalance and protected group ratios. ACM Transactions on Knowl...

  4. [13]

    u ller, Conradin Braun, Domenique Zipperling, and Niklas K \

    Luca Deck, Jan-Laurin M \"u ller, Conradin Braun, Domenique Zipperling, and Niklas K \"u hl. Implications of the ai act for non-discrimination law and algorithmic fairness. arXiv preprint arXiv:2403.20089, 2024

  5. [14]

    Auditing fairness under unawareness through counterfactual reasoning

    Giandomenico Cornacchia, Vito Walter Anelli, Giovanni Maria Biancofiore, Fedelucio Narducci, Claudio Pomo, Azzurra Ragone, and Eugenio Di Sciascio. Auditing fairness under unawareness through counterfactual reasoning. Inf. Process. Manag., 60 0 (2): 0 103224, 2023. doi:10.1016...

  6. [15]

    Measuring fairness in credit ratings

    Ying Chen, Paolo Giudici, Kailiang Liu, and Emanuela Raffinetti. Measuring fairness in credit ratings. Expert Syst. Appl., 258: 0 125184, 2024. doi:10.1016/J.ESWA.2024.125184. URL https://doi.org/10.1016/j.eswa.2024.125184

  7. [16]

    Alessandro Fabris, Gianmaria Silvello, Gian Antonio Susto, and Asia J. Biega. Pairwise fairness in ranking as a dissatisfaction measure. In Tat - Seng Chua, Hady W. Lauw, Luo Si, Evimaria Terzi, and Panayiotis Tsaparas, editors, Proceedings of the Sixteenth ACM International C...

  8. [17]

    Feder Cooper, Katherine Lee, Madiha Zahrah Choksi, Solon Barocas, Christopher De Sa, James Grimmelmann, Jon M

    A. Feder Cooper, Katherine Lee, Madiha Zahrah Choksi, Solon Barocas, Christopher De Sa, James Grimmelmann, Jon M. Kleinberg, Siddhartha Sen, and Baobao Zhang. Arbitrariness and social prediction: The confounding role of variance in fair classification. In Michael J. Wooldridge...

  9. [18]

    Long-term fairness with unknown dynamics

    Tongxin Yin, Reilly Raab, Mingyan Liu, and Yang Liu. Long-term fairness with unknown dynamics. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information Processing Systems 36: Annual Conference on Neural...

  10. [19]

    Cruz and Moritz Hardt

    Andr \' e F. Cruz and Moritz Hardt. Unprocessing seven years of algorithmic fairness. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL https://openreview.net/forum?id=jr03SfWsBS

  11. [20]

    Learning fair representations via rebalancing graph structure

    Guixian Zhang, Debo Cheng, Guan Yuan, and Shichao Zhang. Learning fair representations via rebalancing graph structure. Inf. Process. Manag., 61 0 (1): 0 103570, 2024. doi:10.1016/J.IPM.2023.103570. URL https://doi.org/10.1016/j.ipm.2023.103570

  12. [21]

    FAL-CUR: fair active learning using uncertainty and representativeness on fair clustering

    Ricky Maulana Fajri, Akrati Saxena, Yulong Pei, and Mykola Pechenizkiy. FAL-CUR: fair active learning using uncertainty and representativeness on fair clustering. Expert Syst. Appl., 242: 0 122842, 2024. doi:10.1016/J.ESWA.2023.122842. URL https://doi.org/10.1016/j.eswa.2023.122842

  13. [22]

    Toward fairness in artificial intelligence for medical image analysis: identification and mitigation of potential biases in the roadmap from data collection to model deployment

    Karen Drukker, Weijie Chen, Judy Gichoya, Nicholas Gruszauskas, Jayashree Kalpathy-Cramer, Sanmi Koyejo, Kyle Myers, Rui C S \'a , Berkman Sahiner, Heather Whitney, et al. Toward fairness in artificial intelligence for medical image analysis: identification and mitigation of p...

  14. [23]

    Non-discrimination law in europe: a primer

    Frederik Zuiderveen Borgesius, Nina Baranowska, Philipp Hacker, and Alessandro Fabris. Non-discrimination law in europe: a primer. introducing european non-discrimination law to non-lawyers. arXiv preprint arXiv:2404.08519, 2024

  15. [25]

    Bias on demand: A modelling framework that generates synthetic data with bias

    Joachim Baumann, Alessandro Castelnovo, Riccardo Crupi, Nicole Inverardi, and Daniele Regoli. Bias on demand: A modelling framework that generates synthetic data with bias. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, FAccT 2023, Chi...

  16. [26]

    On explaining unfairness: An overview

    Christos Fragkathoulas, Vasiliki Papanikou, Danae Pla Karidi, and Evaggelia Pitoura. On explaining unfairness: An overview. In 40th International Conference on Data Engineering, ICDE 2024 - Workshops, Utrecht, Netherlands, May 13-16, 2024 , pages 226--236. IEEE , 2024. doi:10....

  17. [27]

    Detecting risk of biased output with balance measures

    Mariachiara Mecati, Antonio Vetr \` o , and Marco Torchiano. Detecting risk of biased output with balance measures. ACM J. Data Inf. Qual. , 14 0 (4): 0 25:1--25:7, 2022. doi:10.1145/3530787. URL https://doi.org/10.1145/3530787

  18. [28]

    Measuring imbalance on intersectional protected attributes and on target variable to forecast unfair classifications

    Mariachiara Mecati, Marco Torchiano, Antonio Vetr \` o , and Juan Carlos De Martin. Measuring imbalance on intersectional protected attributes and on target variable to forecast unfair classifications. IEEE Access , 11: 0 26996--27011, 2023. doi:10.1109/ACCESS.2023.3252370. UR...

  19. [29]

    Wallach, Hal Daum \' e III, and Kate Crawford

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna M. Wallach, Hal Daum \' e III, and Kate Crawford. Datasheets for datasets. Commun. ACM , 64 0 (12): 0 86--92, 2021. doi:10.1145/3458723. URL https://doi.org/10.1145/3458723

  20. [30]

    The dataset nutrition label

    Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski. The dataset nutrition label. Data Protection and Privacy, 12 0 (12): 0 1, 2020

  21. [31]

    Data cards: Purposeful and transparent dataset documentation for responsible AI

    Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson. Data cards: Purposeful and transparent dataset documentation for responsible AI . In FAccT '22: 2022 ACM Conference on Fairness, Accountability, and Transparency, Seoul, Republic of Korea, June 21 - 24, 2022 , pages 177...

  22. [32]

    Algorithmic fairness datasets: the story so far

    Alessandro Fabris, Stefano Messina, Gianmaria Silvello, and Gian Antonio Susto. Algorithmic fairness datasets: the story so far. Data Min. Knowl. Discov., 36 0 (6): 0 2074--2152, 2022. doi:10.1007/S10618-022-00854-Z. URL https://doi.org/10.1007/s10618-022-00854-z

  23. [33]

    Ai documentation: A path to accountability

    Florian K \"o nigstorfer and Stefan Thalmann. Ai documentation: A path to accountability. Journal of Responsible Technology, 11: 0 100043, 2022

  24. [34]

    Completeness of datasets documentation on ML/AI repositories: An empirical investigation

    Marco Rondina, Antonio Vetr \` o , and Juan Carlos De Martin. Completeness of datasets documentation on ML/AI repositories: An empirical investigation. In Nuno Moniz, Zita Vale, Jos \' e Cascalho, Catarina Silva, and Raquel Sebasti \ a o, editors, Progress in Artificial Intell...

  25. [35]

    Pandit, Sven Schade, Declan O'Sullivan, and Dave Lewis

    Delaram Golpayegani, Isabelle Hupont, Cecilia Panigutti, Harshvardhan J. Pandit, Sven Schade, Declan O'Sullivan, and Dave Lewis. AI cards: Towards an applied framework for machine-readable AI and risk documentation inspired by the EU AI act. In Meiko Jensen, C \' e dric Laurad...

  26. [36]

    everyone wants to do the model work, not the data work

    Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen K. Paritosh, and Lora Aroyo. "everyone wants to do the model work, not the data work": Data cascades in high-stakes AI . In Yoshifumi Kitamura, Aaron Quigley, Katherine Isbister, Takeo Igarashi, Pernill...

  27. [37]

    Metrics for dataset demographic bias: A case study on facial expression recognition

    Iris Dominguez - Catena, Daniel Paternain, and Mikel Galar. Metrics for dataset demographic bias: A case study on facial expression recognition. IEEE Trans. Pattern Anal. Mach. Intell. , 46 0 (8): 0 5209--5226, 2024. doi:10.1109/TPAMI.2024.3361979. URL https://doi.org/10.1109/...

  28. [38]

    A survey on bias and fairness in machine learning

    Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. A survey on bias and fairness in machine learning. ACM Comput. Surv. , 54 0 (6): 0 115:1--115:35, 2022. doi:10.1145/3457607. URL https://doi.org/10.1145/3457607

  29. [39]

    Harini Suresh and John V. Guttag. A framework for understanding sources of harm throughout the machine learning life cycle. In EAAMO 2021: ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, Virtual Event, USA, October 5 - 9, 2021 , pages 17:1--17:...

  30. [40]

    Invisible women: Data bias in a world designed for men

    Caroline Criado Perez. Invisible women: Data bias in a world designed for men. Abrams, 2019

  31. [41]

    The uncounted

    Alex Cobham. The uncounted. Polity, 2020

  32. [42]

    Shreya Shankar, Yoni Halpern, Eric Breck, James Atwood, Jimbo Wilson, and D. Sculley. No classification without representation: Assessing geodiversity issues in open data sets for the developing world. In NIPS 2017 workshop: Machine Learning for the Developing World, 2017

  33. [43]

    Gender shades: Intersectional accuracy disparities in commercial gender classification

    Joy Buolamwini and Timnit Gebru. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Sorelle A. Friedler and Christo Wilson, editors, Conference on Fairness, Accountability and Transparency, FAT 2018, 23-24 February 2018, New York, NY, US...

  34. [44]

    Fairness and Machine Learning: Limitations and Opportunities

    Solon Barocas, Moritz Hardt, and Arvind Narayanan. Fairness and Machine Learning: Limitations and Opportunities. MIT Press, 2023

  35. [45]

    The impact of group membership bias on the quality and fairness of exposure in ranking

    Ali Vardasbi, Maarten de Rijke, Fernando Diaz, and Mostafa Dehghani. The impact of group membership bias on the quality and fairness of exposure in ranking. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGI...

  36. [46]

    It's compaslicated: The messy relationship between RAI datasets and algorithmic fairness benchmarks

    Michelle Bao, Angela Zhou, Samantha Zottola, Brian Brubach, Sarah Desmarais, Aaron Horowitz, Kristian Lum, and Suresh Venkatasubramanian. It's compaslicated: The messy relationship between RAI datasets and algorithmic fairness benchmarks. In Joaquin Vanschoren and Sai - Kit Ye...

  37. [47]

    Potential biases in machine learning algorithms using electronic health record data

    Milena A Gianfrancesco, Suzanne Tamang, Jinoos Yazdany, and Gabriela Schmajuk. Potential biases in machine learning algorithms using electronic health record data. JAMA internal medicine, 178 0 (11): 0 1544--1547, 2018

  38. [48]

    Unintended bias and identity terms, 2018

    Jigsaw. Unintended bias and identity terms, 2018

  39. [49]

    Data preprocessing techniques for classification without discrimination

    Faisal Kamiran and Toon Calders. Data preprocessing techniques for classification without discrimination. Knowl. Inf. Syst., 33 0 (1): 0 1--33, 2011. doi:10.1007/S10115-011-0463-8. URL https://doi.org/10.1007/s10115-011-0463-8

  40. [50]

    Friedler, John Moeller, Carlos Scheidegger, and Suresh Venkatasubramanian

    Michael Feldman, Sorelle A. Friedler, John Moeller, Carlos Scheidegger, and Suresh Venkatasubramanian. Certifying and removing disparate impact. In Longbing Cao, Chengqi Zhang, Thorsten Joachims, Geoffrey I. Webb, Dragos D. Margineantu, and Graham Williams, editors, Proceeding...

  41. [52]

    Aim: Attributing, interpreting, mitigating data unfairness

    Zhining Liu, Ruizhong Qiu, Zhichen Zeng, Yada Zhu, Hendrik Hamann, and Hanghang Tong. Aim: Attributing, interpreting, mitigating data unfairness. arXiv e-prints, pages arXiv--2406, 2024

  42. [53]

    Jacobs and Hanna Wallach

    Abigail Z. Jacobs and Hanna Wallach. Measurement and fairness. In Madeleine Clare Elish, William Isaac, and Richard S. Zemel, editors, FAccT '21: 2021 ACM Conference on Fairness, Accountability, and Transparency, Virtual Event / Toronto, Canada, March 3-10, 2021 , pages 375--3...

  43. [54]

    Comparison and benchmark of name-to-gender inference services

    Luc \' a Santamar \' a and Helena Mihaljevic. Comparison and benchmark of name-to-gender inference services. PeerJ Comput. Sci., 4: 0 e156, 2018. doi:10.7717/PEERJ-CS.156. URL https://doi.org/10.7717/peerj-cs.156

  44. [55]

    Demographic prediction based on user's browsing behavior

    Jian Hu, Hua - Jun Zeng, Hua Li, Cheng Niu, and Zheng Chen. Demographic prediction based on user's browsing behavior. In Carey L. Williamson, Mary Ellen Zurko, Peter F. Patel - Schneider, and Prashant J. Shenoy, editors, Proceedings of the 16th International Conference on Worl...

  45. [56]

    Equality of opportunity in supervised learning

    Moritz Hardt, Eric Price, and Nati Srebro. Equality of opportunity in supervised learning. In Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett, editors, Advances in Neural Information Processing Systems 29: Annual Conference on Neural Info...

  46. [57]

    Discrimination-aware data mining

    Dino Pedreschi, Salvatore Ruggieri, and Franco Turini. Discrimination-aware data mining. In Ying Li, Bing Liu, and Sunita Sarawagi, editors, Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Las Vegas, Nevada, USA, August 24-27...

  47. [58]

    Big data's disparate impact

    Solon Barocas and Andrew D Selbst. Big data's disparate impact. Calif. L. Rev., 104: 0 671, 2016

  48. [59]

    David Madras, Elliot Creager, Toniann Pitassi, and Richard S. Zemel. Learning adversarially fair and transferable representations. In Jennifer G. Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm \" a s...

  49. [60]

    Harrison Edwards and Amos J. Storkey. Censoring representations with an adversary. In Yoshua Bengio and Yann LeCun, editors, 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings , 2016. URL http...

  50. [61]

    Reducing unintended bias of ML models on tabular and textual data

    Guilherme Alves, Maxime Amblard, Fabien Bernier, Miguel Couceiro, and Amedeo Napoli. Reducing unintended bias of ML models on tabular and textual data. In 8th IEEE International Conference on Data Science and Advanced Analytics, DSAA 2021, Porto, Portugal, October 6-9, 2021 , ...

  51. [62]

    Explainability statement, 2022

    HireVue. Explainability statement, 2022. URL https://hirevue-api.dev-directory.com/wp-content/uploads/2022/04/HV_AI_Short-Form_Explainability_1pager.pdf

  52. [63]

    Unbiased interviews

    Blind Stairs . Unbiased interviews. https://blindstairs.com/en/interviews \; URL visited on 9th May 2024

  53. [64]

    Navigating demographic measurement for fairness and equity

    Miranda Bogen. Navigating demographic measurement for fairness and equity. https://cdt.org/wp-content/uploads/2024/05/2024-04-29-AI-Gov-Lab-Demographic-Data-report-final.pdf, 2024

  54. [65]

    Chen, and Marzyeh Ghassemi

    Laleh Seyyed-Kalantari, Guanxiong Liu, Matthew McDermott, Irene Y. Chen, and Marzyeh Ghassemi. Chexclusion: Fairness gaps in deep chest x-ray classifiers, 2020

  55. [66]

    Evaluating deep neural networks trained on clinical images in dermatology with the fitzpatrick 17k dataset

    Matthew Groh, Caleb Harris, Luis Soenksen, Felix Lau, Rachel Han, Aerin Kim, Arash Koochek, and Omar Badri. Evaluating deep neural networks trained on clinical images in dermatology with the fitzpatrick 17k dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision ...

  56. [67]

    Measuring discrimination in algorithmic decision making

    Indre Zliobaite. Measuring discrimination in algorithmic decision making. Data Min. Knowl. Discov., 31 0 (4): 0 1060--1089, 2017. doi:10.1007/S10618-017-0506-1. URL https://doi.org/10.1007/s10618-017-0506-1

  57. [68]

    Fairness in deep learning: A computational perspective

    Mengnan Du, Fan Yang, Na Zou, and Xia Hu. Fairness in deep learning: A computational perspective. IEEE Intell. Syst. , 36 0 (4): 0 25--34, 2021. doi:10.1109/MIS.2020.3000681. URL https://doi.org/10.1109/MIS.2020.3000681

  58. [69]

    On formalizing fairness in prediction with machine learning, 2018

    Pratik Gajane and Mykola Pechenizkiy. On formalizing fairness in prediction with machine learning, 2018

  59. [70]

    The fairness of risk scores beyond classification: Bipartite ranking and the xauc metric, 2019

    Nathan Kallus and Angela Zhou. The fairness of risk scores beyond classification: Bipartite ranking and the xauc metric, 2019. URL https://arxiv.org/abs/1902.05826

  60. [71]

    Zhang, Mark Harman, and Federica Sarro

    Max Hort, Zhenpeng Chen, Jie M. Zhang, Mark Harman, and Federica Sarro. Bias mitigation for machine learning classifiers: A comprehensive survey. ACM J. Responsib. Comput., 1 0 (2), June 2024. doi:10.1145/3631326. URL https://doi.org/10.1145/3631326

  61. [72]

    Scaling up the accuracy of naive-bayes classifiers: A decision-tree hybrid

    Ron Kohavi. Scaling up the accuracy of naive-bayes classifiers: A decision-tree hybrid. In Evangelos Simoudis, Jiawei Han, and Usama M. Fayyad, editors, Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (KDD-96), Portland, Oregon, USA , ...

  62. [73]

    Empirical risk minimization under fairness constraints

    Michele Donini, Luca Oneto, Shai Ben - David, John Shawe - Taylor, and Massimiliano Pontil. Empirical risk minimization under fairness constraints. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol \` o Cesa - Bianchi, and Roman Garnett, editors, Advanc...

  63. [74]

    Anian Ruoss, Mislav Balunovic, Marc Fischer, and Martin T. Vechev. Learning certified individually fair representations. In Hugo Larochelle, Marc'Aurelio Ranzato, Raia Hadsell, Maria - Florina Balcan, and Hsuan - Tien Lin, editors, Advances in Neural Information Processing Sys...

  64. [75]

    Mislav Balunovic, Anian Ruoss, and Martin T. Vechev. Fair normalizing flows. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net, 2022. URL https://openreview.net/forum?id=BrFIKuxrZE

  65. [76]

    Retiring adult: New datasets for fair machine learning

    Frances Ding, Moritz Hardt, John Miller, and Ludwig Schmidt. Retiring adult: New datasets for fair machine learning. Advances in Neural Information Processing Systems, 34, 2021

  66. [77]

    Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases

    Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and Ronald M Summers. Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Proceedings of the IEEE conference on co...

  67. [78]

    Are sex-based physiological differences the cause of gender bias for chest x-ray diagnosis? In Workshop on Clinical Image-Based Procedures, pages 142--152

    Nina Weng, Siavash Bigdeli, Eike Petersen, and Aasa Feragen. Are sex-based physiological differences the cause of gender bias for chest x-ray diagnosis? In Workshop on Clinical Image-Based Procedures, pages 142--152. Springer, 2023

  68. [79]

    Fitzpatrick

    Thomas B. Fitzpatrick. The validity and practicality of sun-reactive skin types i through vi. Archives of dermatology, 124 6: 0 869--71, 1988. URL https://api.semanticscholar.org/CorpusID:29991932

  69. [80]

    Novoa, Justin M

    Andre Esteva, Brett Kuprel, Roberto A. Novoa, Justin M. Ko, Susan M. Swetter, Helen M. Blau, and Sebastian Thrun. Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542: 0 115--118, 2017. URL https://api.semanticscholar.org/CorpusID:3767412

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.