Pith. sign in

REVIEW 3 major objections 4 minor 43 references

Critical Appraisal of Fairness Metrics in Clinical Predictive AI

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A scoping review finds 62 fairness metrics for clinical AI, but only one measures whether decisions improve patient outcomes.

desk verdict Useful taxonomy and critical appraisal of fairness metrics for clinical AI, but the 'only one clinical utility metric' headline is less robust than it looks. read the letter →

arxiv 2506.17035 v1 pith:BVQPX52H submitted 2025-06-20 cs.LG

classification cs.LG
keywords fairnessmetricsclinicalpredictionmodelsalgorithmicbiasscopingreviewutilitysubgroupnetbenefithealthequitythreshold-dependent
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This scoping review sets out to establish what fairness metrics actually exist for clinical predictive AI and whether they are fit for purpose. Reviewing 41 studies published between 2014 and 2024, the authors identified 62 distinct fairness metrics, classified them by whether they depend on model performance and on decision thresholds, and critically appraised each one. They find that the field is fragmented: only 18 metrics were explicitly developed for healthcare, most metrics rely on arbitrarily chosen thresholds, and only one metric, subgroup net benefit, measures clinical utility. The paper argues that fairness evaluation should move toward clinically meaningful, threshold-independent metrics and should treat subgroup performance, calibration, and clinical utility as the central objects.

What carries the argument

The organizing machinery is a taxonomy that sorts each metric along three axes: performance-independent versus performance-dependent (whether the metric uses outcome labels, i.e., whether it asks if the model's behaviour or its performance should be equalized across groups); probability-based versus threshold-dependent (whether it uses estimated probabilities $\hat{p}$ or predicted classes $\hat{Y}$); and, for performance-dependent metrics, the type of base performance metric (discrimination, calibration, overall performance, partial classification, summary accuracy, or clinical utility). The taxonomy does the work of converting 62 heterogeneous formulas into comparable categories, and the paper pairs it with a decision diagram and a three-level guidance scheme to route users to clinically appropriate metrics.

What would settle it

A repeat search that adds 'bias', 'parity', and 'disparity' and uses two independent reviewers to screen and classify metrics would falsify the central claim if it surfaced several additional healthcare-developed clinical-utility fairness metrics or if reclassification of ambiguous names substantially changed the category counts.

Watch

Extended reading notes

Core claim

The central claim is that the current fairness-metric landscape in clinical predictive AI is fragmented, under-validated, and poorly aligned with clinical decision-making. From 820 screened records the authors retained 41 studies and extracted 62 metrics meeting their definition of a fairness metric: a measure quantifying whether a model's output discriminates, in the societal sense, against individuals or groups defined by a sensitive attribute. Most metrics (47 of 62) are performance-dependent, meaning they compare model performance across groups and therefore tend to preserve the status quo encoded in labels; 33 are threshold-dependent, meaning their value changes with the arbitrary choice of a decision cut-off. Only 18 metrics were explicitly proposed for healthcare, and only one, subgroup net benefit, directly captures clinical utility. The authors conclude that the lack of clear definitions and standardisation undermines systematic fairness assessment, and they provide a three-branch taxonomy and a guidance table to help researchers choose and interpret metrics.

Load-bearing premise

The review's headline counts rest on the assumption that its search, screening, and metric standardization were complete and correct; the search omitted the terms 'bias', 'parity', and 'disparity', and screening and extraction were done by a single reviewer.

Editorial extensions

If this is right

  • Reporting guidelines such as TRIPOD+AI and FUTURE-AI will need to be more specific about fairness, since the review shows the current guidance is cautious and non-specific.
  • Evaluators of clinical prediction models should pair discrimination-based fairness metrics such as AUROC parity with calibration-based parity metrics, and should treat threshold-dependent metrics as descriptive for a single threshold only.
  • Subgroup net benefit is the only threshold-dependent fairness metric the review advises on its own, because it captures the harms-and-benefits trade-off at a clinically meaningful threshold.
  • Performance-independent metrics such as conditional statistical parity have a role when outcome labels themselves may encode inequities, but they should be used only when no legitimate risk differences exist between subgroups.
  • Fairness evaluation should be reframed from statistical parity at all costs to minimum acceptable performance for all groups, guided by clinical relevance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A search that added the omitted terms 'bias', 'parity', and 'disparity' would likely surface additional metrics, so the precise counts (62 metrics, one clinical-utility metric) should be read as lower bounds rather than an exhaustive census.
  • The reclassification of ambiguously named metrics, such as renaming 'Discrimination Index' as 'F1-Score Parity', was done by a single reviewer; a second independent extraction could shift some category counts at the margins.
  • The review's emphasis on probability-based and clinical-utility metrics suggests a concrete research agenda: compare how often parity-based and minimum-performance-based criteria select different models on the same clinical datasets.
  • Subgroup net benefit could be extended into a family of utility-based fairness measures, for example net-benefit differences and subgroup decision-curve analysis, which the paper identifies as promising but does not itself develop.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper reports a PRISMA-ScR scoping review of fairness metrics for clinical predictive AI. The authors searched five databases from 2014 to 2024, supplemented by backward citation tracking, and included 41 studies from which they extracted 62 metrics satisfying their Box 1 definition of a fairness metric. They propose a taxonomy that distinguishes performance-independent versus performance-dependent metrics, probability-based versus threshold-dependent metrics, and base performance metric type (calibration, discrimination, overall, partial, summary, clinical utility), and they critically appraise each metric with qualitative guidance levels (recommended, use with caution, inadvisable). The headline findings are that the fairness metric landscape is fragmented, most metrics lack clinical validation and empirical evaluation, 33/62 are threshold-dependent, and only one clinical utility metric (subgroup net benefit) was identified. The paper concludes that future work should prioritise clinically meaningful, probability-based metrics and addresses gaps in uncertainty quantification, intersectionality, and sample size considerations.

Significance. The review addresses a timely and important question, and its structured taxonomy and critical appraisal are useful contributions for a field that is rapidly evolving. Strengths include the detailed search protocol, the OSF supplementary materials with data extraction forms and a metric catalogue, the explicit definitions in Box 1, and the practical guidance in Box 5 that could inform reporting guidelines such as TRIPOD+AI and FUTURE-AI. The conclusion that the field is fragmented and that clinical utility is under-served is potentially valuable, but its significance depends on the accuracy of the specific quantitative claims, especially the statement that only one clinical utility metric was identified, and on the defensibility of the classification choices. Because the protocol is transparent, the headline claims are falsifiable and could be audited or extended in future work, which is a genuine strength.

major comments (3)
  1. [Results, §2.2.3 and Abstract] The claim that only one clinical utility metric (subgroup net benefit, ref 88) was identified is load-bearing for the conclusion that the fairness metric landscape lacks clinical relevance. This count depends on the classification of ref 71 (Pfohl et al., 'Net benefit, calibration, threshold selection, and training objectives for algorithmic fairness in healthcare'), which is included in the review but used only for ACE Parity (§2.1.2 and Box 3). Because the title indicates that the paper explicitly frames net benefit in the context of algorithmic fairness, the review should either extract a clinical utility metric from it (if it compares net benefit across protected groups) or explain why its net benefit analysis does not meet the Box 1 definition of a fairness metric. Please verify this classification and, if necessary, revise the counts and the 'only one' statement in the abstract.
  2. [Limitations and Search Strategy] The search intentionally omitted the terms 'bias', 'parity', 'disparity', as well as 'net benefit' and 'clinical utility'. Since the headline counts (62 metrics, 18 healthcare-specific, one clinical utility) are completeness-sensitive, and many clinical papers express fairness in terms of parity or disparity rather than using the word 'fairness', the omission could materially affect the results. The backward citation tracking partially compensates, but the manuscript does not report how many included studies were found via backward citation versus the primary search, nor does it assess the sensitivity of the counts to the omitted terms. Please provide this information and discuss its implications for the robustness of the quantitative findings.
  3. [Methods, Data Extraction] All screening, data extraction, and classification was performed by a single reviewer (JM), including subjective decisions such as redefining 'Discrimination Index' as 'F1-Score Parity' (ref 52) and classifying equalising disincentives (ref 74) and treatment equality (ref 79) as not clinical utility metrics. These decisions directly affect the headline counts. The manuscript should either provide a reproducibility check (for example, a second reviewer on a random subset with inter-rater reliability) or a detailed coding manual with decision rules, and temper the precision of the counts in the abstract if such a check is not feasible.
minor comments (4)
  1. [Results, Performance-independent metrics] The text reports n=15 for performance-independent metrics but then says 'only 3 of these 16 metrics'; the breakdown (probability-based n=6, threshold-dependent n=7) does not sum to 15 or 16, apparently omitting the two individual fairness metrics. Please harmonise these numbers with Table 1.
  2. [Discussion] The sentence 'The predominance of performance-dependent (46/62) and threshold-dependent (33/62) metrics' is inconsistent with Table 1 and the Results section, which report 47 performance-dependent metrics; please correct to 47/62.
  3. [Methods, Data Extraction] The example of standardisation ('Discrimination Index' redefined as 'F1-Score Parity') is helpful, but the full list of such standardisation decisions should be placed in the supplementary materials so that readers can audit the classification and its potential impact on the counts.
  4. [Supplementary Table 2.4] The comparison with previous reviews would be more informative if the paper also reported which of the 62 metrics overlap with those found in the previous reviews and which are new, to support the claim in the Limitations that the search 'covered metrics identified in previous reviews, as well as new ones'.

Circularity Check

0 steps flagged · score 0.0 of 10

Scoping review's headline counts are empirical synthesis, not derivation; no circularity found.

full rationale

This is a scoping review, not a derivation. The claimed outputs — 62 fairness metrics, 18 healthcare-proposed metrics, and one clinical utility metric — are counts obtained by screening 820 records and extracting metrics from 41 included studies under a stated definition of a 'fairness metric' (Box 1). No step in the paper reduces a predicted quantity to a fitted parameter, nor defines an input in terms of an output. The taxonomy in Figure 1 builds on prior taxonomies (refs 1, 36, 44, 53), some authored by the present team, and the critical appraisal cites the authors' earlier framework on proper scoring rules (ref 1), but these citations supply background concepts such as properness and bias-preserving versus bias-transforming metrics rather than forcing the empirical counts; the counts would not change by algebraic construction if those citations were removed. The standardization of metric names, e.g., redefining 'Discrimination Index' as 'F1-Score Parity', is transparent data extraction, not the presentation of a known result as a novel derivation. The only substantive concerns are search-term omissions and single-reviewer classification, including whether 'equalising disincentives' or 'treatment equality' should count as clinical utility metrics; these are correctness and completeness risks explicitly acknowledged in the Limitations section, not circularity. No step exhibits the characteristic pattern of an equation being equal to its own input by construction or of a fitted parameter being renamed as a prediction.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The review has no free parameters or invented entities. It relies on domain assumptions: the working definition of a fairness metric, the adequacy of the search strategy, and the appropriateness of scoping review methodology.

assumptions (3)
  • domain assumption A 'fairness metric' is defined as a measure quantifying whether a model discriminates against individuals or groups defined by sensitive attributes, excluding causal fairness notions and metrics lacking operational detail.
    This definition determines which metrics are included in the 62; stated in Methods and Box 1.
  • domain assumption Searching PubMed, ACM, IEEE, arXiv, and medRxiv with terms 'fairness', 'metric', 'clinical', and 'model' from 2014 to 2024 is sufficient to capture the relevant fairness metric literature.
    The authors acknowledge the omission of synonyms like 'bias', 'parity', and 'disparity' as a limitation.
  • domain assumption Scoping review methodology per Arksey and O'Malley and PRISMA-ScR is appropriate for mapping fairness metrics.
    These are standard frameworks; the paper follows them.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Critical Appraisal of Fairness Metrics in Clinical Predictive AI." pith.science (2026). https://pith.science/paper/BVQPX52H

@misc{pith2026250617035,
  author       = {Pith},
  title        = {Pith review of: Critical Appraisal of Fairness Metrics in Clinical Predictive AI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BVQPX52H}},
  note         = {Machine review of arXiv:2506.17035}
}
read the original abstract

Predictive artificial intelligence (AI) offers an opportunity to improve clinical practice and patient outcomes, but risks perpetuating biases if fairness is inadequately addressed. However, the definition of "fairness" remains unclear. We conducted a scoping review to identify and critically appraise fairness metrics for clinical predictive AI. We defined a "fairness metric" as a measure quantifying whether a model discriminates (societally) against individuals or groups defined by sensitive attributes. We searched five databases (2014-2024), screening 820 records, to include 41 studies, and extracted 62 fairness metrics. Metrics were classified by performance-dependency, model output level, and base performance metric, revealing a fragmented landscape with limited clinical validation and overreliance on threshold-dependent measures. Eighteen metrics were explicitly developed for healthcare, including only one clinical utility metric. Our findings highlight conceptual challenges in defining and quantifying fairness and identify gaps in uncertainty quantification, intersectionality, and real-world applicability. Future work should prioritise clinically meaningful metrics.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

43 extracted references · 27 canonical work pages

  1. [1]

    Performance evaluation of predictive AI models to support medical decisions: Overview and guidance

    1 Van Calster B, Collins GS, Vickers AJ, et al. Performance evaluation of predictive AI models to support medical decisions: Overview and guidance. arXiv [cs.LG]. 2024; published online Dec

  2. [3]

    93 McDermott MBA, Zhang H, Hansen LH, Angelotti G, Gallifant J

    DOI:10.2139/ssrn.3547922. 93 McDermott MBA, Zhang H, Hansen LH, Angelotti G, Gallifant J. A closer look at AUROC and AUPRC under class imbalance. arXiv [cs.LG]. 2024; published online Jan

  3. [4]

    58 Keya KN, Islam R, Pan S, Stockwell I, Foulds JR

    http://arxiv.org/abs/2310.03146. 58 Keya KN, Islam R, Pan S, Stockwell I, Foulds JR. Equitable allocation of healthcare resources with fair cox models. arXiv preprint arXiv:2010 06820

  4. [6]

    DOI:10.1093/OED/8693190878

    2024; published online Dec. DOI:10.1093/OED/8693190878. 120 justice, n. meanings, etymology and more. https://www.oed.com/dictionary/justice_n?tab=meaning_and_use#40269279 (accessed March 4, 2025). 121 Beigang F. Reconciling algorithmic fairness criteria. Philos Public Aff 2023; 51 : 166–90. 122 Ravindranath R, Stein JD, Hernandez-Boussard T, Fisher AC, W...

  5. [10]

    Understanding algorithmic fairness for clinical prediction in terms of subgroup net benefit and health equity

    http://arxiv.org/abs/2412.07879. 89 Mittelstadt B, Wachter S, Russell C. The Unfairness of Fair Machine Learning: Levelling down and strict egalitarianism by default. https://papers.ssrn.com › sol3 › papershttps://papers.ssrn.com › sol3 › papers. 2023; published online Jan

  6. [11]

    The use of clinical risk factors enhances the performance of BMD in the prediction of hip and osteoporotic fractures in men and women

    6 Kanis JA, Oden A, Johnell O, et al. The use of clinical risk factors enhances the performance of BMD in the prediction of hip and osteoporotic fractures in men and women. Osteoporos Int 2007; 18 : 1033–46. 7 Hippisley-Cox J, Coupland C, Brindle P. Development and validation of QRISK3 risk prediction algorithms to estimate future risk of cardiovascular d...

  7. [12]

    112 Riley RD, Collins GS, Archer L, et al

    http://arxiv.org/abs/2407.09293 (accessed May 16, 2025). 112 Riley RD, Collins GS, Archer L, et al. A decomposition of Fisher’s information to inform sample size for Page 31 of 32 Fairness metrics in clinical predictive AI developing fair and precise clinical prediction models -- Part 2: time-to-event outcomes. arXiv [stat.ME]. 2025; published online Jan

  8. [13]

    Performance evaluation of predictive AI models to support medical decisions: Overview and guidance

    http://arxiv.org/abs/2412.10288. 2 van Smeden M, Reitsma JB, Riley RD, Collins GS, Moons KG. Clinical prediction models: diagnosis versus prognosis. J Clin Epidemiol 2021; 132 : 142–5. 3 Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med 2019; 25 : 44–56. 4 Wynants L, Van Calster B, Collins GS, et al. Predict...

Show all 43 references
  1. [15]

    54 van der Meijden SL, Wang Y, Arbous MS, Geerts BF, Steyerberg EW, Hernandez-Boussard T

    DOI:10.2139/ssrn.3792772. 54 van der Meijden SL, Wang Y, Arbous MS, Geerts BF, Steyerberg EW, Hernandez-Boussard T. Navigating fairness in AI-based prediction models: Theoretical constructs and practical applications. medRxiv. 2025; published online March

  2. [17]

    Fairness-enhancing mixed effects deep learning improves fairness on in- and out-of-distribution clustered (non-iid) data

    57 Nguyen S, Wang A, Montillo A. Fairness-enhancing mixed effects deep learning improves fairness on in- and out-of-distribution clustered (non-iid) data. arXiv [cs.LG]. 2023; published online Oct

  3. [18]

    Ethical limitations of algorithmic fairness solutions in health care machine learning

    98 McCradden MD, Joshi S, Mazwi M, Anderson JA. Ethical limitations of algorithmic fairness solutions in health care machine learning. The Lancet Digital Health 2020; 2 : e221–3. 99 Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW, Topic Group ‘Evaluating dia...

  4. [19]

    37 Kusner MJ, Loftus JR, Russell C, Silva R

    http://arxiv.org/abs/1104.3913 (accessed Feb 17, 2025). 37 Kusner MJ, Loftus JR, Russell C, Silva R. Counterfactual Fairness. arXiv [stat.ML]. 2017; published online March

  5. [20]

    38 Martinez N, Bertran M, Sapiro G

    http://arxiv.org/abs/1703.06856 (accessed Feb 17, 2025). 38 Martinez N, Bertran M, Sapiro G. Minimax Pareto Fairness: A Multi Objective Perspective. In: International Conference on Machine Learning. PMLR, 2020: 6755–64. 39 Lekadir K, Frangi AF, Porras AR, et al. FUTURE-AI: int...

  6. [21]

    Certifying and removing disparate impact

    62 Feldman M, Friedler SA, Moeller J, Scheidegger C, Venkatasubramanian S. Certifying and removing disparate impact. In: proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining. 2015: 259–68. 63 Cheng V, Suriyakumar VM, Dullerud N, Jo...

  7. [22]

    Measuring and Reducing Racial Bias in a Pediatric Urinary Tract Infection Model

    68 Anderson JW, Shaikh N, Visweswaran S. Measuring and Reducing Racial Bias in a Pediatric Urinary Tract Infection Model. AMIA Summits on Translational Science Proceedings 2024; 2024 :

  8. [23]

    Longitudinal fairness with censorship

    70 Zhang W, Weiss JC. Longitudinal fairness with censorship. In: proceedings of the AAAI conference on artificial intelligence. 2022: 12235–43. 71 Pfohl S, Xu Y, Foryciarz A, Ignatiadis N, Genkins J, Shah N. Net benefit, calibration, threshold selection, and training objectives ...

  9. [24]

    55 Foulds JR, Islam R, Keya KN, Pan S

    DOI:10.1101/2025.03.24.25324500. 55 Foulds JR, Islam R, Keya KN, Pan S. An intersectional definition of fairness. In: 2020 IEEE 36th International Conference on Data Engineering (ICDE). IEEE, 2020: 1918–21. 56 Cui S, Pan W, Zhang C, Wang F. Bipartite ranking fairness through a ...

  10. [25]

    Detection and mitigation of algorithmic bias via predictive parity

    77 DiCiccio C, Hsu B, Yu Y, Nandy P, Basu K. Detection and mitigation of algorithmic bias via predictive parity. In: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 2023: 1801–16. 78 Beutel A, Chen J, Doshi T, et al. Putting fairness princ...

  11. [26]

    117 Thomassen D, le Cessie S, van Houwelingen HC, Steyerberg EW

    http://arxiv.org/abs/2401.14893 (accessed April 25, 2025). 117 Thomassen D, le Cessie S, van Houwelingen HC, Steyerberg EW. Effective sample size: A measure of individual uncertainty in predictions. Stat Med 2024; 43 : 1384–96. 118 Efthimiou O, Seo M, Chalkou K, Debray T, Egge...

  12. [27]

    80 Xiao Y, Lim S, Pollard TJ, Ghassemi M

    http://arxiv.org/abs/1703.09207. 80 Xiao Y, Lim S, Pollard TJ, Ghassemi M. In the name of fairness: assessing the bias in clinical record de-identification. In: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 2023: 123–37. 81 Riad R, Denais...

  13. [28]

    90 Coston A, Ramamurthy KN, Wei D, et al

    https://papers.ssrn.com/abstract=4331652 (accessed Dec 12, 2023). 90 Coston A, Ramamurthy KN, Wei D, et al. Fair transfer learning with missing protected attributes. In: Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society. 2019: 91–8. 91 Watkins EA, Chen J. ...

  14. [29]

    92 Wachter S, Mittelstadt B, Russell C

    DOI:10.1145/3630106.3658938. 92 Wachter S, Mittelstadt B, Russell C. Why fairness cannot be automated: Bridging the gap between EU non-discrimination law and AI. SSRN Electron J 2020; published online March

  15. [31]

    42 Pessach D, Shmueli E

    http://arxiv.org/abs/1808.00023. 42 Pessach D, Shmueli E. A review on fairness in machine learning. ACM Computing Surveys (CSUR) 2022; 55 : 1–44. 43 Mehrabi N, Morstatter F, Saxena N, Lerman K, Galstyan A. A survey on bias and fairness in machine learning. ACM Comput Surv 2022...

  16. [32]

    94 Chouldechova A

    http://arxiv.org/abs/2401.06091 (accessed March 27, 2025). 94 Chouldechova A. Fair prediction with disparate impact: A study of bias in recidivism prediction instruments. Big data 2017; 5 : 153–63. 95 Błasiok J, Nakkiran P. Smooth ECE: Principled reliability diagrams via kerne...

  17. [33]

    96 Nam RK, Kattan MW, Chin JL, et al

    http://arxiv.org/abs/2309.12236. 96 Nam RK, Kattan MW, Chin JL, et al. Prospective multi-institutional study evaluating the performance of prostate cancer risk calculators. J Clin Oncol 2011; 29 : 2959–64. 97 Vickers AJ, van Calster B, Steyerberg EW. A simple, step-by-step gui...

  18. [36]

    102 Tal E

    http://arxiv.org/abs/2205.08875. 102 Tal E. Target specification bias, counterfactual prediction, and algorithmic fairness in healthcare. In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society. New York, NY, USA: ACM, 2023: 312–21. 103 Jun I, Ser SE, Cohen S...

  19. [37]

    107 Nielsen MW, Gissi E, Heidari S, et al

    https://www.ijcai.org/proceedings/2023/0742.pdf. 107 Nielsen MW, Gissi E, Heidari S, et al. Intersectional analysis for science and technology. Nature 2025; 640 : 329–37. 108 Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med De...

  20. [39]

    113 Riley RD, Collins GS, Kirton L, et al

    http://arxiv.org/abs/2501.14482 (accessed May 16, 2025). 113 Riley RD, Collins GS, Kirton L, et al. Uncertainty of risk estimates from clinical prediction models: rationale, challenges, and approaches. BMJ 2025; 388 : e080749. 114 intersectionality, n. meanings, etymology and ...

  21. [40]

    116 Herlihy C, Truong K, Chouldechova A, Dudik M

    http://arxiv.org/abs/2304.09270. 116 Herlihy C, Truong K, Chouldechova A, Dudik M. A structured regression approach for evaluating model performance across intersectional subgroups. arXiv [cs.LG]. 2024; published online Jan

  22. [43]

    Evaluating the Impact of Social Determinants on Health Prediction in the Intensive Care Unit

    126 Yang MY, Kwak GH, Pollard T, Celi LA, Ghassemi M. Evaluating the Impact of Social Determinants on Health Prediction in the Intensive Care Unit. In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society. 2023: 333–50. Page 32 of 32

  23. [126]

    Tackling algorithmic bias and promoting transparency in health datasets: the STANDING Together consensus recommendations

    31 Alderman JE, Palmer J, Laws E, et al. Tackling algorithmic bias and promoting transparency in health datasets: the STANDING Together consensus recommendations. Lancet Digit Health 2025; 7 : e64–88. 32 Oakden-Rayner L, Dunnmon J, Carneiro G, Ré C. Hidden stratification causes...

  24. [230]

    Soliciting stakeholders’ fairness notions in child maltreatment predictive systems

    100 Cheng H-F, Stapleton L, Wang R, et al. Soliciting stakeholders’ fairness notions in child maltreatment predictive systems. In: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. New York, NY, USA: ACM,

  25. [243]

    Ethnic classifications in algorithmic fairness: Concepts, measures and implications in practice

    29 Jaime S, Kern C. Ethnic classifications in algorithmic fairness: Concepts, measures and implications in practice. In: The 2024 ACM Conference on Fairness, Accountability, and Transparency. 2024: 237–53. 30 Goetz L, Seedat N, Vandersluis R, van der Schaar M. Generalization-a ...

  26. [488]

    Peeking into a black box, the fairness and generalizability of a MIMIC-III benchmarking model

    69 Röösli E, Bozkurt S, Hernandez-Boussard T. Peeking into a black box, the fairness and generalizability of a MIMIC-III benchmarking model. Scientific Data 2022; 9 :

  27. [2010]

    https://www.legislation.gov.uk/ukpga/2010/15/contents (accessed April 3, 2025). 26 Legislation summary - How is discrimination addressed in EU legislation? https://www.eu-patient.eu/policy/Policy/Anti-discrimination/legislation-summary---how-is-discrimination-addressed-in-eu-l...

  28. [2016]

    Algorithmic decision making and the cost of fairness

    73 Corbett-Davies S, Pierson E, Feller A, Goel S, Huq A. Algorithmic decision making and the cost of fairness. In: Proceedings of the 23rd acm sigkdd international conference on knowledge discovery and data mining. 2017: 797–806. 74 Jung C, Kannan S, Lee C, Pai M, Roth A, Vohr...

  29. [2018]

    41 Corbett-Davies S, Gaebler JD, Nilforoshan H, Shroff R, Goel S

    DOI:10.1145/3194770.3194776. 41 Corbett-Davies S, Gaebler JD, Nilforoshan H, Shroff R, Goel S. The measure and mismeasure of fairness. arXiv [cs.CY]. 2018; published online July

  30. [2019]

    36 Dwork C, Hardt M, Pitassi T, Reingold O, Zemel R

    DOI:10.1145/3287560.3287594. 36 Dwork C, Hardt M, Pitassi T, Reingold O, Zemel R. Fairness Through Awareness. arXiv [cs.CC]. 2011; published online April

  31. [2020]

    Fair and interpretable models for survival analysis

    59 Rahman MM, Purushotham S. Fair and interpretable models for survival analysis. In: Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. New York, NY, USA: ACM, 2022: Page 29 of 32 Fairness metrics in clinical predictive AI 1452–62. 60 Zemel ...

  32. [2021]

    101 Chien I, Deliu N, Turner RE, Weller A, Villar SS, Kilbertus N

    DOI:10.1145/3411764.3445308. 101 Chien I, Deliu N, Turner RE, Weller A, Villar SS, Kilbertus N. Multi-disciplinary fairness considerations in machine learning for clinical trials. arXiv [cs.LG]. 2022; published online May

  33. [2022]

    51 Luo Y, Tian Y, Shi M, et al

    DOI:10.1145/3514094.3534137. 51 Luo Y, Tian Y, Shi M, et al. Harvard glaucoma fairness: a retinal nerve disease dataset for fairness learning and fair identity normalization. IEEE Transactions on Medical Imaging

  34. [2023]

    Fairness in Machine Learning: A survey

    45 Caton S, Haas C. Fairness in Machine Learning: A survey. ACM Comput Surv 2024; 56 : 1–38. 46 Anderson JW, Visweswaran S. Algorithmic individual fairness and healthcare: a scoping review. JAMIA Open 2025; 8 : ooae149. 47 Mienye ID, Swart TG, Obaido G. Fairness Metrics in AI ...

  35. [2024]

    Fairfl: A fair federated learning approach to reducing demographic bias in privacy-sensitive classification models

    52 Zhang DY, Kou Z, Wang D. Fairfl: A fair federated learning approach to reducing demographic bias in privacy-sensitive classification models. In: 2020 IEEE International Conference on Big Data (Big Data). IEEE, 2020: 1051–60. 53 Wachter S, Mittelstadt B, Russell C. Bias preser...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.