REVIEW 3 major objections 5 minor 77 references
Underrepresentation, Label Bias, and Proxies: Towards Data Bias Profiles for the EU AI Act and Beyond
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Underrepresentation of vulnerable groups in training data is a weaker driver of algorithmic discrimination than label bias and proxy features, and this paper introduces a Data Bias Profile to quantify each.
desk verdict Solid label-bias and proxy results, but the underrepresentation claim is only tested under uniform random subsampling and is undercut by the paper's own Adult (gender) numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the Data Bias Profile (DBP), a quantitative summary that records three bias signals computed without access to an unbiased reference set: the Representation Difference, the difference between advantaged and disadvantaged group prevalence; the Separation Difference, the average of two cross-group AUC gaps measuring how much harder positive examples of the disadvantaged group are to rank than those of the advantaged group; and the proxy factor sAUC, the AUC of a classifier that attempts to predict the protected attribute from non-protected features. The paper shows each statistic responds mainly to its own injected bias, and that the profile's pattern anticipates both the level of model unfairness and whether a proxy-removal intervention will help.
What would settle it
Run the same bias-injection study with underrepresentation implemented as removal of a structurally defined subpopulation of the disadvantaged group, such as all instances sharing an intersectional feature combination, keeping test sets unbiased; if equal-opportunity gaps rise sharply with the underrepresentation factor, the paper's relative ranking of the three biases fails.
Extended reading notes
Core claim
The paper's central claim is that the relative importance of data biases for algorithmic discrimination is inverted relative to conventional wisdom: removing up to 100% of a disadvantaged group from the training set leaves equal-opportunity gaps roughly constant in most settings, whereas flipping a share of that group's positive labels produces steep, consistent increases in unfairness, especially in datasets where non-protected features strongly predict the protected attribute. The authors interpret this as evidence that representation-driven interventions are overvalued and that label curation and proxy management deserve priority. To support this claim they evaluate models on unbiased test sets, inject biases only in training and validation, and show that label bias can be strong enough that including the disadvantaged group without fixing labels makes outcomes worse for that group.
Load-bearing premise
The claim that underrepresentation is overemphasized assumes it takes the form of uniform random removal of disadvantaged instances; underrepresentation that removes entire feature regions or intersectional subgroups could hurt the group much more.
Editorial extensions
If this is right
- Dataset balance should not be treated as the primary fairness fix; scarce annotated data from vulnerable groups is better spent on evaluation than on training.
- Label quality is the first thing to audit: even a 20% flip of disadvantaged-group positive labels can widen equal-opportunity gaps significantly.
- Proxy strength should guide intervention choice; removing features correlated with the protected attribute helps when sAUC is high and does little when it is low.
- Including more disadvantaged-group samples without cleaning their labels can actively worsen outcomes for them, so representation efforts and label repair must go together.
- For deployment documentation, a DBP can flag which bias signals a dataset carries and which fairness-enhancing intervention is likely to pay off.
Reading between the lines
- If underrepresentation were imposed by deleting structured subpopulations rather than uniform random subsampling, the overemphasis conclusion could flip; testing that variant is a direct extension of the injection protocol.
- The label-bias results reframe label noise as a fairness problem; the paper's injection protocol could serve as a benchmark for group-dependent label-noise correction methods.
- A DBP with thresholds and multi-group support would be a stronger compliance tool for anti-discrimination audits; the paper explicitly leaves thresholds and multi-group extensions as open work.
- The finding that label bias can make representation harmful suggests that data collection and labeling efforts should be evaluated jointly, not as separate pipeline stages.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies three types of data bias—underrepresentation, label bias, and proxies—by injecting them into training and validation sets of seven datasets, then measuring fairness on unbiased test sets with three fairness metrics and several models. It reports that underrepresentation has only a minor effect on discrimination, while label bias has a large effect that is amplified by strong proxies. The paper then proposes three detection measures (RD, SD, sAUC) and combines them into a preliminary Data Bias Profile (DBP), demonstrated on popular fairness datasets as a proof of concept for documenting bias, predicting discrimination risk, and selecting fairness interventions. The discussion connects these findings to data governance requirements under the EU AI Act.
Significance. If the empirical ranking of biases is correct, this work would usefully reorder fairness practice toward label curation and proxy management rather than representation targets alone, and the DBP is a promising quantitative complement to qualitative dataset documentation. The paper has concrete strengths: a broad experimental matrix (seven datasets, multiple model families, three fairness metrics, ten repetitions with significance tests), evaluation on unbiased test sets, clearly specified injection protocols, and an explicit limitation discussion. The label-bias and proxy-amplification findings are well supported by the reported results. However, the headline underrepresentation claim and the validation of the DBP detection measures are not yet established at the level of generality claimed in the abstract and Section 6; both need substantial qualification or additional experiments.
major comments (3)
- [Section 4.2, Eq. (1); Section 4.3, Table 3; Appendix B.1, Tables 9 and 12] The central claim that underrepresentation in training is overemphasized is inferred entirely from an injection mechanism that removes disadvantaged-group points uniformly at random. Real-world underrepresentation often removes structured subsets (feature subspaces, intersectional strata), which shifts P(x|y,s=d) rather than only reducing sample size, and this case is never tested. Even under the tested mechanism, the paper's own results do not uniformly show a minor impact: Adult (gender) EO rises from 0.08±0.02 to 0.21±0.07 for LR, and the appendix shows larger jumps for RF (0.10 to 0.30) and SVC (0.08 to 0.25), with PQP also significantly worsening. The Limitations section lists other bias types and binary attributes but does not acknowledge this operationalization gap. The abstract and Section 6 should either restrict the conclusion to uniform random subsampling or be backed by structured-missingness experiments.
- [Section 5.1, Eqs. (10), (13), (14); Section 5.2, Figure 4] The bias detection measures are validated on biases injected through the same mathematical quantities they measure: RD is exactly the prevalence gap manipulated by Eq. (1), sAUC is exactly the AUC of the proxy classifier used to define proxy strength in Eq. (3), and SD is an AUC-based separability gap that label flipping directly changes. The diagonal responses in Figure 4 are therefore partly by construction, so the experiments demonstrate internal consistency rather than the ability to detect real-world biases of these types. Since the practical value of the DBP depends on detection validity, the paper should provide an external validation or test the measures under bias mechanisms not defined by the same formulas, such as structured underrepresentation or naturally occurring group-dependent label noise.
- [Section 4.4, Table 5; Section 6] The recommendation that including disadvantaged groups in training can be harmful under weak label bias goes beyond the evidence. The ΔEO values for f=0.2 are negative for most datasets, but several carry standard errors that overlap zero (e.g., Adult-gender -0.01±0.08, Crime -0.11±0.17, Folktables -0.01±0.06, German -0.09±0.17), and NIH and Compas show positive values. The interaction may be real, but the current wording ('hastily adding disadvantaged groups ... can cause more harm than good') is too strong for these point estimates. Please temper the claim or report the uncertainty more prominently.
minor comments (5)
- [Table 2 and Section 4.2, Eq. (1)] The notation table defines u = r - 1, while the text and Eq. (1) define u = 1 - r as the underrepresentation factor. Please correct the table to avoid a sign inconsistency.
- [Section 4.4, paragraph on the joint effect] The sentence 'as confirmed by the first column of Table 4 (f = 0)' appears to be a reference error; the relevant ΔEO values are in Table 5, column f = 0, not in Table 4.
- [Section 5.2, Footnote 5 and Figure 4] Replacing u=1 and f=1 with 0.95 for the detection experiments is reasonable, but the figure axes and the appendix figures still display 1 at the maximum. Please align the axis labels with the actual values used.
- [Figure 2 caption and Figure 5] There are minor typos: 'disadvataged' in the Figure 2 caption should be 'disadvantaged', and 'Folkstables' in Figure 5 should be 'Folktables'.
- [Appendix A] The NIH disease list contains 'mas', which should likely be 'mass', and the COMPAS description spells 'ProPulica' instead of 'ProPublica'. Please fix these typos.
Circularity Check
Bias-detection validation is partially self-definitional (RD exactly; sAUC and SD largely), while the main underrepresentation/label-bias effect study is empirical and independent.
-
self definitional
[Section 5.1 Eq. (10) and Section 5.2, Figure 4]
"RD(σ) = |σa| − |σd| / |σ| = Pr σ (s = a) − Pr σ (s = d) (10) ... underrepresentation is suitably captured by RD increasing linearly in the first column"
Eq. (1) defines underrepresentation as Prσ′(s=d)=r·Prσ(s=d), i.e., only the prevalence gap changes. RD(σ) is exactly that prevalence gap. Therefore RD's linear response to the underrepresentation injection is an algebraic identity: RD(σ′) = (1−Prσ(s=d)) − r·Prσ(s=d). No classifier, label information, or external signal is involved; the detection curve is determined by the definition of the injection, so calling it evidence of detection is circular.
-
self definitional
[Section 4.2 Eqs. (3)-(4) and Section 5.2]
"We quantify the strength of proxies as their joint ability to predict sensitive attributes. We train a classifier ˆs = h(x) to estimate the protected attribute s and we compute its AUC to measure the strength of proxies. ... An additive protocol adds to the non-sensitive variables X a new feature correlated with sensitive variables xnew = s + v, v ∼ N(0, std2)"
The proxy injection adds a feature equal to the sensitive attribute plus noise, and the proxy detector sAUC is the AUC of a classifier trained to predict exactly that sensitive attribute from the augmented feature space. As std→0, xnew→s, so sAUC is forced toward 1 by construction. Thus the diagonal 'detection' of proxies in Figure 4 largely restates the injection mechanism rather than providing independent confirmation that the measure tracks an external proxy phenomenon.
1 more flagged steps
-
self definitional
[Section 4.2 Eq. (2), Section 5.1 Eq. (13), and Section 5.2]
"For label bias, we selectively flip labels. We let f ∈ (0, 1) indicate the proportion of positive instances (y = ⊕) from the vulnerable group whose label is flipped to negative (y = ⊖) ... SD(σ) = ∆xAUCσ + ∆wAUCσ / 2 ... We expect label bias to worsen the separability for the disadvantaged group and therefore yield high values of SD."
Label-bias injection directly flips group-d positives to negatives, which is exactly the operation that destroys the ranking separability of group-d positives relative to negatives. SD is defined as a difference of cross-group and within-group AUC separability. The injected flip therefore directly lowers the quantities SD is built from, making its increase with f largely entailed by the construction. Some empirical content remains because SD is computed from a trained classifier, but the diagonal response is primarily a restatement of the injection in ranking terms.
full rationale
The main effect analysis in Section 4 is not circular: biases are injected into training data through explicit mechanisms (Eqs. 1-5) and fairness is evaluated on unbiased test sets, so the conclusion that label bias is more critical than underrepresentation is an empirical, falsifiable result that could have gone either way. No fitted parameter is relabeled as a prediction, and no load-bearing self-citation chain is used; the authors' own prior work appears only in contextual citations. The partial circularity lies in Section 5's validation of the detection measures. RD is exactly the prevalence gap manipulated by Eq. (1), so its diagonal response is an algebraic identity. sAUC and SD are validated with injections defined through the same constructs, so their diagonal responses are largely by construction. Because these measures feed the Data Bias Profile, the DBP's detection claims inherit some definitional character, though the DBP case study is qualitative and not a fitted predictor. The uniform-subsampling operationalization of underrepresentation is an external-validity limitation, not a circularity, and does not affect the score. Overall, the central fairness finding stands on independent evidence; the detection-validation claim is partially circular.
Assumptions & free parameters
assumptions (5)
- domain assumption The original datasets, before bias injection, provide an unbiased test set for the target population.
- ad hoc to paper Underrepresentation is modeled by uniform random subsampling of the disadvantaged group.
- ad hoc to paper Proxy strength equals the AUC of a classifier trained to predict the sensitive attribute from non-sensitive features (sAUC).
- ad hoc to paper The three biases can be injected and detected independently.
- domain assumption Equal opportunity, demographic parity, and predictive quality parity are adequate measures of algorithmic discrimination.
Cite this review
Pith. "Pith review of Underrepresentation, Label Bias, and Proxies: Towards Data Bias Profiles for the EU AI Act and Beyond." pith.science (2026). https://pith.science/paper/S2CRHZ2O
@misc{pith2026250708866,
author = {Pith},
title = {Pith review of: Underrepresentation, Label Bias, and Proxies: Towards Data Bias Profiles for the EU AI Act and Beyond},
year = {2026},
howpublished = {\url{https://pith.science/paper/S2CRHZ2O}},
note = {Machine review of arXiv:2507.08866}
}
read the original abstract
Undesirable biases encoded in the data are key drivers of algorithmic discrimination. Their importance is widely recognized in the algorithmic fairness literature, as well as legislation and standards on anti-discrimination in AI. Despite this recognition, data biases remain understudied, hindering the development of computational best practices for their detection and mitigation. In this work, we present three common data biases and study their individual and joint effect on algorithmic discrimination across a variety of datasets, models, and fairness measures. We find that underrepresentation of vulnerable populations in training sets is less conducive to discrimination than conventionally affirmed, while combinations of proxies and label bias can be far more critical. Consequently, we develop dedicated mechanisms to detect specific types of bias, and combine them into a preliminary construct we refer to as the Data Bias Profile (DBP). This initial formulation serves as a proof of concept for how different bias signals can be systematically documented. Through a case study with popular fairness datasets, we demonstrate the effectiveness of the DBP in predicting the risk of discriminatory outcomes and the utility of fairness-enhancing interventions. Overall, this article bridges algorithmic fairness research and anti-discrimination policy through a data-centric lens.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Jos \' e M. \' A lvarez, Alejandra Bringas Colmenarejo, Alaa Elobaid, Simone Fabbrizzi, Miriam Fahimi, Antonio Ferrara, Siamak Ghodsi, Carlos Mougan, Ioanna Papageorgiou, Paula Reyero Lobo, Mayra Russo, Kristen M. Scott, Laura State, Xuan Zhao, and Salvatore Ruggieri. Policy advice and best practices on bias and fairness in AI . Ethics Inf. Technol., 26 0...
-
[2]
Machine bias
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner. Machine bias. ProPublica, 2016. URL https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing
2016
-
[3]
Dissecting racial bias in an algorithm used to manage the health of populations
Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366 0 (6464): 0 447--453, 2019
2019
-
[5]
Towards a standard for identifying and managing bias in artificial intelligence
Reva Schwartz, Apostol Vassilev, Kristen Greene, Lori Perine, Andrew Burt, and Patrick Hall. Towards a standard for identifying and managing bias in artificial intelligence. US Department of Commerce, National Institute of Standards and Technology, 2022
work page 2022
-
[6]
European Parliament. Artificial intelligence act. https://www.europarl.europa.eu/doceo/document/TA-9-2024-0138_EN.pdf, 2024
work page 2024
-
[7]
ISO. Information technology — artificial intelligence (ai) — bias in ai systems and ai aided decision making, 2021. https://www.iso.org/standard/77607.html
work page 2021
-
[8]
Gillis, Vitaly Meursault, and Berk Ustun
Talia B. Gillis, Vitaly Meursault, and Berk Ustun. Operationalizing the search for less discriminatory alternatives in fair lending. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT 2024, Rio de Janeiro, Brazil, June 3-6, 2024 , pages 377--387. ACM , 2024. doi:10.1145/3630106.3658912. URL https://doi.org/10.1145/3630106.3658912
arXiv 2024
-
[9]
Fairness and bias in algorithmic hiring
Alessandro Fabris, Nina Baranowska, Matthew J Dennis, David Graus, Philipp Hacker, Jorge Saldivar, Frederik Zuiderveen Borgesius, and Asia J Biega. Fairness and bias in algorithmic hiring. ACM Transactions on Intelligent Systems and Technology, 2024. URL https://doi.org/10.1145/3696457
doi:10.1145/3696457 2024
Show all 77 references
-
[10]
Baker and Aaron Hawn
Ryan S. Baker and Aaron Hawn. Algorithmic bias in education. Int. J. Artif. Intell. Educ., 32 0 (4): 0 1052--1092, 2022. doi:10.1007/S40593-021-00285-9. URL https://doi.org/10.1007/s40593-021-00285-9
2022 doi
-
[11]
A data quality approach to the identification of discrimination risk in automated decision making systems
Antonio Vetr \` o , Marco Torchiano, and Mariachiara Mecati. A data quality approach to the identification of discrimination risk in automated decision making systems. Gov. Inf. Q., 38 0 (4): 0 101619, 2021. doi:10.1016/J.GIQ.2021.101619. URL https://doi.org/10.1016/j.giq.2021.101619
2021
-
[12]
Properties of fairness measures in the context of varying class imbalance and protected group ratios
Dariusz Brzezinski, Julia Stachowiak, Jerzy Stefanowski, Izabela Szczech, Robert Susmaga, Sofya Aksenyuk, Uladzimir Ivashka, and Oleksandr Yasinskyi. Properties of fairness measures in the context of varying class imbalance and protected group ratios. ACM Transactions on Knowl...
2024
-
[13]
u ller, Conradin Braun, Domenique Zipperling, and Niklas K \
Luca Deck, Jan-Laurin M \"u ller, Conradin Braun, Domenique Zipperling, and Niklas K \"u hl. Implications of the ai act for non-discrimination law and algorithmic fairness. arXiv preprint arXiv:2403.20089, 2024
2024 arXiv
-
[14]
Auditing fairness under unawareness through counterfactual reasoning
Giandomenico Cornacchia, Vito Walter Anelli, Giovanni Maria Biancofiore, Fedelucio Narducci, Claudio Pomo, Azzurra Ragone, and Eugenio Di Sciascio. Auditing fairness under unawareness through counterfactual reasoning. Inf. Process. Manag., 60 0 (2): 0 103224, 2023. doi:10.1016...
2023
-
[15]
Measuring fairness in credit ratings
Ying Chen, Paolo Giudici, Kailiang Liu, and Emanuela Raffinetti. Measuring fairness in credit ratings. Expert Syst. Appl., 258: 0 125184, 2024. doi:10.1016/J.ESWA.2024.125184. URL https://doi.org/10.1016/j.eswa.2024.125184
2024
-
[16]
Alessandro Fabris, Gianmaria Silvello, Gian Antonio Susto, and Asia J. Biega. Pairwise fairness in ranking as a dissatisfaction measure. In Tat - Seng Chua, Hady W. Lauw, Luo Si, Evimaria Terzi, and Panayiotis Tsaparas, editors, Proceedings of the Sixteenth ACM International C...
2023
-
[17]
Feder Cooper, Katherine Lee, Madiha Zahrah Choksi, Solon Barocas, Christopher De Sa, James Grimmelmann, Jon M
A. Feder Cooper, Katherine Lee, Madiha Zahrah Choksi, Solon Barocas, Christopher De Sa, James Grimmelmann, Jon M. Kleinberg, Siddhartha Sen, and Baobao Zhang. Arbitrariness and social prediction: The confounding role of variance in fair classification. In Michael J. Wooldridge...
2024
-
[18]
Long-term fairness with unknown dynamics
Tongxin Yin, Reilly Raab, Mingyan Liu, and Yang Liu. Long-term fairness with unknown dynamics. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information Processing Systems 36: Annual Conference on Neural...
2023
-
[19]
Cruz and Moritz Hardt
Andr \' e F. Cruz and Moritz Hardt. Unprocessing seven years of algorithmic fairness. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL https://openreview.net/forum?id=jr03SfWsBS
2024
-
[20]
Learning fair representations via rebalancing graph structure
Guixian Zhang, Debo Cheng, Guan Yuan, and Shichao Zhang. Learning fair representations via rebalancing graph structure. Inf. Process. Manag., 61 0 (1): 0 103570, 2024. doi:10.1016/J.IPM.2023.103570. URL https://doi.org/10.1016/j.ipm.2023.103570
2024
-
[21]
FAL-CUR: fair active learning using uncertainty and representativeness on fair clustering
Ricky Maulana Fajri, Akrati Saxena, Yulong Pei, and Mykola Pechenizkiy. FAL-CUR: fair active learning using uncertainty and representativeness on fair clustering. Expert Syst. Appl., 242: 0 122842, 2024. doi:10.1016/J.ESWA.2023.122842. URL https://doi.org/10.1016/j.eswa.2023.122842
2024
-
[22]
Toward fairness in artificial intelligence for medical image analysis: identification and mitigation of potential biases in the roadmap from data collection to model deployment
Karen Drukker, Weijie Chen, Judy Gichoya, Nicholas Gruszauskas, Jayashree Kalpathy-Cramer, Sanmi Koyejo, Kyle Myers, Rui C S \'a , Berkman Sahiner, Heather Whitney, et al. Toward fairness in artificial intelligence for medical image analysis: identification and mitigation of p...
2023
-
[23]
Non-discrimination law in europe: a primer
Frederik Zuiderveen Borgesius, Nina Baranowska, Philipp Hacker, and Alessandro Fabris. Non-discrimination law in europe: a primer. introducing european non-discrimination law to non-lawyers. arXiv preprint arXiv:2404.08519, 2024
2024 arXiv
-
[25]
Bias on demand: A modelling framework that generates synthetic data with bias
Joachim Baumann, Alessandro Castelnovo, Riccardo Crupi, Nicole Inverardi, and Daniele Regoli. Bias on demand: A modelling framework that generates synthetic data with bias. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, FAccT 2023, Chi...
2023
-
[26]
On explaining unfairness: An overview
Christos Fragkathoulas, Vasiliki Papanikou, Danae Pla Karidi, and Evaggelia Pitoura. On explaining unfairness: An overview. In 40th International Conference on Data Engineering, ICDE 2024 - Workshops, Utrecht, Netherlands, May 13-16, 2024 , pages 226--236. IEEE , 2024. doi:10....
2024
-
[27]
Detecting risk of biased output with balance measures
Mariachiara Mecati, Antonio Vetr \` o , and Marco Torchiano. Detecting risk of biased output with balance measures. ACM J. Data Inf. Qual. , 14 0 (4): 0 25:1--25:7, 2022. doi:10.1145/3530787. URL https://doi.org/10.1145/3530787
2022 doi
-
[28]
Measuring imbalance on intersectional protected attributes and on target variable to forecast unfair classifications
Mariachiara Mecati, Marco Torchiano, Antonio Vetr \` o , and Juan Carlos De Martin. Measuring imbalance on intersectional protected attributes and on target variable to forecast unfair classifications. IEEE Access , 11: 0 26996--27011, 2023. doi:10.1109/ACCESS.2023.3252370. UR...
2023
-
[29]
Wallach, Hal Daum \' e III, and Kate Crawford
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna M. Wallach, Hal Daum \' e III, and Kate Crawford. Datasheets for datasets. Commun. ACM , 64 0 (12): 0 86--92, 2021. doi:10.1145/3458723. URL https://doi.org/10.1145/3458723
2021 doi
-
[30]
The dataset nutrition label
Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski. The dataset nutrition label. Data Protection and Privacy, 12 0 (12): 0 1, 2020
2020
-
[31]
Data cards: Purposeful and transparent dataset documentation for responsible AI
Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson. Data cards: Purposeful and transparent dataset documentation for responsible AI . In FAccT '22: 2022 ACM Conference on Fairness, Accountability, and Transparency, Seoul, Republic of Korea, June 21 - 24, 2022 , pages 177...
2022
-
[32]
Algorithmic fairness datasets: the story so far
Alessandro Fabris, Stefano Messina, Gianmaria Silvello, and Gian Antonio Susto. Algorithmic fairness datasets: the story so far. Data Min. Knowl. Discov., 36 0 (6): 0 2074--2152, 2022. doi:10.1007/S10618-022-00854-Z. URL https://doi.org/10.1007/s10618-022-00854-z
2022 doi
-
[33]
Ai documentation: A path to accountability
Florian K \"o nigstorfer and Stefan Thalmann. Ai documentation: A path to accountability. Journal of Responsible Technology, 11: 0 100043, 2022
2022
-
[34]
Completeness of datasets documentation on ML/AI repositories: An empirical investigation
Marco Rondina, Antonio Vetr \` o , and Juan Carlos De Martin. Completeness of datasets documentation on ML/AI repositories: An empirical investigation. In Nuno Moniz, Zita Vale, Jos \' e Cascalho, Catarina Silva, and Raquel Sebasti \ a o, editors, Progress in Artificial Intell...
2023
-
[35]
Pandit, Sven Schade, Declan O'Sullivan, and Dave Lewis
Delaram Golpayegani, Isabelle Hupont, Cecilia Panigutti, Harshvardhan J. Pandit, Sven Schade, Declan O'Sullivan, and Dave Lewis. AI cards: Towards an applied framework for machine-readable AI and risk documentation inspired by the EU AI act. In Meiko Jensen, C \' e dric Laurad...
2024
-
[36]
everyone wants to do the model work, not the data work
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen K. Paritosh, and Lora Aroyo. "everyone wants to do the model work, not the data work": Data cascades in high-stakes AI . In Yoshifumi Kitamura, Aaron Quigley, Katherine Isbister, Takeo Igarashi, Pernill...
2021
-
[37]
Metrics for dataset demographic bias: A case study on facial expression recognition
Iris Dominguez - Catena, Daniel Paternain, and Mikel Galar. Metrics for dataset demographic bias: A case study on facial expression recognition. IEEE Trans. Pattern Anal. Mach. Intell. , 46 0 (8): 0 5209--5226, 2024. doi:10.1109/TPAMI.2024.3361979. URL https://doi.org/10.1109/...
2024
-
[38]
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. A survey on bias and fairness in machine learning. ACM Comput. Surv. , 54 0 (6): 0 115:1--115:35, 2022. doi:10.1145/3457607. URL https://doi.org/10.1145/3457607
2022 doi
-
[39]
Harini Suresh and John V. Guttag. A framework for understanding sources of harm throughout the machine learning life cycle. In EAAMO 2021: ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, Virtual Event, USA, October 5 - 9, 2021 , pages 17:1--17:...
2021
-
[40]
Invisible women: Data bias in a world designed for men
Caroline Criado Perez. Invisible women: Data bias in a world designed for men. Abrams, 2019
2019
-
[41]
The uncounted
Alex Cobham. The uncounted. Polity, 2020
2020
-
[42]
Shreya Shankar, Yoni Halpern, Eric Breck, James Atwood, Jimbo Wilson, and D. Sculley. No classification without representation: Assessing geodiversity issues in open data sets for the developing world. In NIPS 2017 workshop: Machine Learning for the Developing World, 2017
2017
-
[43]
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Sorelle A. Friedler and Christo Wilson, editors, Conference on Fairness, Accountability and Transparency, FAT 2018, 23-24 February 2018, New York, NY, US...
2018
-
[44]
Fairness and Machine Learning: Limitations and Opportunities
Solon Barocas, Moritz Hardt, and Arvind Narayanan. Fairness and Machine Learning: Limitations and Opportunities. MIT Press, 2023
2023
-
[45]
The impact of group membership bias on the quality and fairness of exposure in ranking
Ali Vardasbi, Maarten de Rijke, Fernando Diaz, and Mostafa Dehghani. The impact of group membership bias on the quality and fairness of exposure in ranking. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGI...
2024
-
[46]
It's compaslicated: The messy relationship between RAI datasets and algorithmic fairness benchmarks
Michelle Bao, Angela Zhou, Samantha Zottola, Brian Brubach, Sarah Desmarais, Aaron Horowitz, Kristian Lum, and Suresh Venkatasubramanian. It's compaslicated: The messy relationship between RAI datasets and algorithmic fairness benchmarks. In Joaquin Vanschoren and Sai - Kit Ye...
2021
-
[47]
Potential biases in machine learning algorithms using electronic health record data
Milena A Gianfrancesco, Suzanne Tamang, Jinoos Yazdany, and Gabriela Schmajuk. Potential biases in machine learning algorithms using electronic health record data. JAMA internal medicine, 178 0 (11): 0 1544--1547, 2018
2018
-
[48]
Unintended bias and identity terms, 2018
Jigsaw. Unintended bias and identity terms, 2018
2018
-
[49]
Data preprocessing techniques for classification without discrimination
Faisal Kamiran and Toon Calders. Data preprocessing techniques for classification without discrimination. Knowl. Inf. Syst., 33 0 (1): 0 1--33, 2011. doi:10.1007/S10115-011-0463-8. URL https://doi.org/10.1007/s10115-011-0463-8
2011 doi
-
[50]
Friedler, John Moeller, Carlos Scheidegger, and Suresh Venkatasubramanian
Michael Feldman, Sorelle A. Friedler, John Moeller, Carlos Scheidegger, and Suresh Venkatasubramanian. Certifying and removing disparate impact. In Longbing Cao, Chengqi Zhang, Thorsten Joachims, Geoffrey I. Webb, Dragos D. Margineantu, and Graham Williams, editors, Proceeding...
2015
-
[52]
Aim: Attributing, interpreting, mitigating data unfairness
Zhining Liu, Ruizhong Qiu, Zhichen Zeng, Yada Zhu, Hendrik Hamann, and Hanghang Tong. Aim: Attributing, interpreting, mitigating data unfairness. arXiv e-prints, pages arXiv--2406, 2024
2024
-
[53]
Jacobs and Hanna Wallach
Abigail Z. Jacobs and Hanna Wallach. Measurement and fairness. In Madeleine Clare Elish, William Isaac, and Richard S. Zemel, editors, FAccT '21: 2021 ACM Conference on Fairness, Accountability, and Transparency, Virtual Event / Toronto, Canada, March 3-10, 2021 , pages 375--3...
2021
-
[54]
Comparison and benchmark of name-to-gender inference services
Luc \' a Santamar \' a and Helena Mihaljevic. Comparison and benchmark of name-to-gender inference services. PeerJ Comput. Sci., 4: 0 e156, 2018. doi:10.7717/PEERJ-CS.156. URL https://doi.org/10.7717/peerj-cs.156
2018 doi
-
[55]
Demographic prediction based on user's browsing behavior
Jian Hu, Hua - Jun Zeng, Hua Li, Cheng Niu, and Zheng Chen. Demographic prediction based on user's browsing behavior. In Carey L. Williamson, Mary Ellen Zurko, Peter F. Patel - Schneider, and Prashant J. Shenoy, editors, Proceedings of the 16th International Conference on Worl...
2007
-
[56]
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro. Equality of opportunity in supervised learning. In Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett, editors, Advances in Neural Information Processing Systems 29: Annual Conference on Neural Info...
2016
-
[57]
Discrimination-aware data mining
Dino Pedreschi, Salvatore Ruggieri, and Franco Turini. Discrimination-aware data mining. In Ying Li, Bing Liu, and Sunita Sarawagi, editors, Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Las Vegas, Nevada, USA, August 24-27...
2008
-
[58]
Big data's disparate impact
Solon Barocas and Andrew D Selbst. Big data's disparate impact. Calif. L. Rev., 104: 0 671, 2016
2016
-
[59]
David Madras, Elliot Creager, Toniann Pitassi, and Richard S. Zemel. Learning adversarially fair and transferable representations. In Jennifer G. Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm \" a s...
2018
-
[60]
Harrison Edwards and Amos J. Storkey. Censoring representations with an adversary. In Yoshua Bengio and Yann LeCun, editors, 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings , 2016. URL http...
2016 arXiv
-
[61]
Reducing unintended bias of ML models on tabular and textual data
Guilherme Alves, Maxime Amblard, Fabien Bernier, Miguel Couceiro, and Amedeo Napoli. Reducing unintended bias of ML models on tabular and textual data. In 8th IEEE International Conference on Data Science and Advanced Analytics, DSAA 2021, Porto, Portugal, October 6-9, 2021 , ...
2021
-
[62]
Explainability statement, 2022
HireVue. Explainability statement, 2022. URL https://hirevue-api.dev-directory.com/wp-content/uploads/2022/04/HV_AI_Short-Form_Explainability_1pager.pdf
2022
-
[63]
Unbiased interviews
Blind Stairs . Unbiased interviews. https://blindstairs.com/en/interviews \; URL visited on 9th May 2024
2024
-
[64]
Navigating demographic measurement for fairness and equity
Miranda Bogen. Navigating demographic measurement for fairness and equity. https://cdt.org/wp-content/uploads/2024/05/2024-04-29-AI-Gov-Lab-Demographic-Data-report-final.pdf, 2024
2024
-
[65]
Chen, and Marzyeh Ghassemi
Laleh Seyyed-Kalantari, Guanxiong Liu, Matthew McDermott, Irene Y. Chen, and Marzyeh Ghassemi. Chexclusion: Fairness gaps in deep chest x-ray classifiers, 2020
2020
-
[66]
Evaluating deep neural networks trained on clinical images in dermatology with the fitzpatrick 17k dataset
Matthew Groh, Caleb Harris, Luis Soenksen, Felix Lau, Rachel Han, Aerin Kim, Arash Koochek, and Omar Badri. Evaluating deep neural networks trained on clinical images in dermatology with the fitzpatrick 17k dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision ...
2021
-
[67]
Measuring discrimination in algorithmic decision making
Indre Zliobaite. Measuring discrimination in algorithmic decision making. Data Min. Knowl. Discov., 31 0 (4): 0 1060--1089, 2017. doi:10.1007/S10618-017-0506-1. URL https://doi.org/10.1007/s10618-017-0506-1
2017 doi
-
[68]
Fairness in deep learning: A computational perspective
Mengnan Du, Fan Yang, Na Zou, and Xia Hu. Fairness in deep learning: A computational perspective. IEEE Intell. Syst. , 36 0 (4): 0 25--34, 2021. doi:10.1109/MIS.2020.3000681. URL https://doi.org/10.1109/MIS.2020.3000681
2021
-
[69]
On formalizing fairness in prediction with machine learning, 2018
Pratik Gajane and Mykola Pechenizkiy. On formalizing fairness in prediction with machine learning, 2018
2018
-
[70]
The fairness of risk scores beyond classification: Bipartite ranking and the xauc metric, 2019
Nathan Kallus and Angela Zhou. The fairness of risk scores beyond classification: Bipartite ranking and the xauc metric, 2019. URL https://arxiv.org/abs/1902.05826
2019 arXiv
-
[71]
Zhang, Mark Harman, and Federica Sarro
Max Hort, Zhenpeng Chen, Jie M. Zhang, Mark Harman, and Federica Sarro. Bias mitigation for machine learning classifiers: A comprehensive survey. ACM J. Responsib. Comput., 1 0 (2), June 2024. doi:10.1145/3631326. URL https://doi.org/10.1145/3631326
2024 doi
-
[72]
Scaling up the accuracy of naive-bayes classifiers: A decision-tree hybrid
Ron Kohavi. Scaling up the accuracy of naive-bayes classifiers: A decision-tree hybrid. In Evangelos Simoudis, Jiawei Han, and Usama M. Fayyad, editors, Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (KDD-96), Portland, Oregon, USA , ...
1996
-
[73]
Empirical risk minimization under fairness constraints
Michele Donini, Luca Oneto, Shai Ben - David, John Shawe - Taylor, and Massimiliano Pontil. Empirical risk minimization under fairness constraints. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol \` o Cesa - Bianchi, and Roman Garnett, editors, Advanc...
2018
-
[74]
Anian Ruoss, Mislav Balunovic, Marc Fischer, and Martin T. Vechev. Learning certified individually fair representations. In Hugo Larochelle, Marc'Aurelio Ranzato, Raia Hadsell, Maria - Florina Balcan, and Hsuan - Tien Lin, editors, Advances in Neural Information Processing Sys...
2020
-
[75]
Mislav Balunovic, Anian Ruoss, and Martin T. Vechev. Fair normalizing flows. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net, 2022. URL https://openreview.net/forum?id=BrFIKuxrZE
2022
-
[76]
Retiring adult: New datasets for fair machine learning
Frances Ding, Moritz Hardt, John Miller, and Ludwig Schmidt. Retiring adult: New datasets for fair machine learning. Advances in Neural Information Processing Systems, 34, 2021
2021
-
[77]
Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases
Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and Ronald M Summers. Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Proceedings of the IEEE conference on co...
2017
-
[78]
Are sex-based physiological differences the cause of gender bias for chest x-ray diagnosis? In Workshop on Clinical Image-Based Procedures, pages 142--152
Nina Weng, Siavash Bigdeli, Eike Petersen, and Aasa Feragen. Are sex-based physiological differences the cause of gender bias for chest x-ray diagnosis? In Workshop on Clinical Image-Based Procedures, pages 142--152. Springer, 2023
2023
-
[79]
Fitzpatrick
Thomas B. Fitzpatrick. The validity and practicality of sun-reactive skin types i through vi. Archives of dermatology, 124 6: 0 869--71, 1988. URL https://api.semanticscholar.org/CorpusID:29991932
1988
-
[80]
Novoa, Justin M
Andre Esteva, Brett Kuprel, Roberto A. Novoa, Justin M. Ko, Susan M. Swetter, Helen M. Blau, and Sebastian Thrun. Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542: 0 115--118, 2017. URL https://api.semanticscholar.org/CorpusID:3767412
2017
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.