Pith. sign in

REVIEW 3 major objections 3 minor 1 cited by

Understanding algorithmic fairness for clinical prediction in terms of subgroup net benefit and health equity

T0 review · 3 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper argues that algorithmic fairness for clinical prediction should be assessed by subgroup net benefit, a utility-based measure of how a model distributes clinical benefit across protected groups.

desk verdict A useful reframing of net benefit for subgroup fairness, but every equity-gap claim rests on one fixed lambda, and the paper never tests it. read the letter →

arxiv 2412.07879 v1 pith:QUNG2VV6 submitted 2024-12-10 stat.AP

classification stat.AP
keywords clinicalpredictionmodelsalgorithmicfairnesssubgroupnetbenefithealthequityinequalitiesdecisioncurveanalysisresourceconstraintsmaximin
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that algorithmic fairness in clinical prediction should be judged by a model's clinical impact in each protected subgroup, not by parity of predictions or error rates. The authors extend net benefit, the standard decision-analytic measure of a model's value, into a subgroup net benefit that adds each group's baseline disease burden and the relative risk reduction of the intervention. Using this measure, they show how a model distributes benefit across ethnic or deprivation groups, how to compare models by whether they narrow or widen the gap between best- and worst-off groups, and why resource limits can force a trade-off between overall benefit and equity. Two worked examples, type 2 diabetes prevention and lung cancer screening, demonstrate that the metric is computable, collapsible across subgroups, and sensitive to modelling choices such as including the protected attribute as a predictor.

What carries the argument

The central object is the subgroup net benefit, a re-scaled decision-curve net benefit that reintroduces the model-independent prevalence term $1-\pi$ and weights the conventional net benefit by $\lambda = RRR$, the relative risk reduction of the intervention at the decision threshold. It carries the argument by converting benefit from a property of predictions into a property of clinical decisions in each group: a group with $\text{sNB}=1$ is unburdened by the outcome, while a group with $\text{sNB}=0$ consists entirely of untreated false negatives. Its collapsibility, meaning the total sNB is the size-weighted mean of subgroup sNBs, is what allows the paper to treat overall benefit and maximin equity as aligned without resource constraints, and to make explicit the Pareto trade-off when capacity is capped.

What would settle it

Estimate subgroup-specific $\lambda$ from randomized trial data or from observational data with treatment assignment for the diabetes prevention and lung cancer screening interventions; if the subgroup-specific relative risk reductions differ from the single published RRRs of 0.58 and 0.20 by more than a few percent, recompute the sNB rankings and gap reductions. If the ordering of models by maximin sNB flips, or if the reported gap-narrowing effects reverse, the paper's fairness conclusions hinge on the single-$\lambda$ assumption.

Watch

Extended reading notes

Core claim

The paper's central claim is that the fairness of a clinical prediction model is a question about the distribution of clinical benefit, and that this distribution can be measured by the subgroup net benefit $\text{sNB}(t^*) = 1 - \pi + \frac{\lambda}{N}\left(TP(t^*) - \frac{t^*}{1-t^*}FP(t^*)\right)$, where $\pi$ is subgroup prevalence, $TP$ and $FP$ are true and false positives at threshold $t^*$, and $\lambda$ is the relative risk reduction of the intervention the model allocates. Because the term $1-\pi$ is retained, a group with high disease burden starts with lower sNB, so treating all groups equally is not enough to identify fairness. The metric is collapsible: the sNB of the whole population is the size-weighted average of subgroup sNBs, so improving the worst-off subgroup (maximin) need not sacrifice overall benefit when resources are unconstrained. Under capacity constraints, however, the paper shows via Pareto fronts that there is a genuine trade-off between overall net benefit and the sNB of the most disadvantaged group, and it proposes using the Pareto curve to make this trade-off explicit for decision-makers.

Load-bearing premise

The load-bearing assumption is that the relative risk reduction $\lambda$ of the intervention is known and identical across subgroups; if the true treatment effect varies by ethnicity or deprivation, or if it differs in the model-identified positive patients from the trial population that supplied $\lambda$, every sNB level, gap reduction, and Pareto trade-off in the paper would change.

Editorial extensions

If this is right

  • Including the protected attribute as a predictor in the diabetes model raised the sNB of Asian and Black groups and narrowed the gap to the white group compared with omitting it, so model-building choices can be evaluated by their equity effect.
  • In the lung cancer example, all three modelling strategies reduced the sNB gap between deprivation quintiles compared to screening no one, and the benefit concentrated in the most-deprived quintile.
  • A random 'fair' screening policy reduced gaps only by lowering sNB in every group, illustrating that naive randomization harms all subgroups rather than levelling up.
  • Under tight capacity caps of screening 3% or 1% of the population, overall and most-deprived-quintile sNB lie on a Pareto front, and the loss in overall benefit across the front is tiny compared to the gain in the worst-off group's sNB.
  • Because sNB is collapsible, without resource constraints one can always achieve both maximum overall benefit and maximin equity by ensembling each subgroup's best model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: Because sNB depends on $\lambda$, the relative risk reduction, the same model can look fair or unfair under different assumptions about treatment effectiveness; reporting sNB as a function of $\lambda$ rather than at a single value would reveal how sensitive equity conclusions are.
  • Editorial inference: The collapsibility of sNB suggests a natural fairness diagnostic for any deployed model: decompose the change in overall sNB into subgroup contributions and attribute the equity gap to differences in prevalence versus differences in predictive performance.
  • Editorial inference: The Pareto-front formulation could be extended from thresholds to other levers such as screening intervals, outreach targeting, or model retraining, so that equity trade-offs are considered across the whole implementation, not just at the decision threshold.
  • Editorial inference: A testable prediction of the framework is that models trained to maximize overall sNB under no constraints will rarely be the ones that minimize the sNB gap; if this holds across many clinical settings, gap-based maximin criteria would be needed in model selection even when resources are abundant.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper proposes a new fairness metric for clinical prediction models, the subgroup net benefit (sNB), defined in Eq. (5) as sNB = 1 - pi + (lambda/N)(TP - (t/(1-t))FP). The sNB reintroduces the prevalence term and a treatment-benefit weight lambda into standard decision-curve net benefit, allowing the clinical impact of a model to be quantified and compared across protected subgroups. The authors argue that fairness should be assessed through the distribution of sNB across groups, using gap reduction and maximin reasoning, and show that resource constraints can create a Pareto trade-off between overall net benefit and the most-underserved subgroup's sNB. The approach is illustrated with two UK Biobank case studies: a 5-year type-2 diabetes risk model compared across ethnicities, and a 6-year lung cancer screening model compared across Townsend deprivation quintiles. The paper finds that models including protected attributes improve sNB in disadvantaged groups and reduce health-inequality gaps.

Significance. If the sNB framework is accepted, it provides a decision-relevant, utility-based alternative to statistical fairness criteria such as demographic parity or equalized odds, and it directly links algorithmic fairness to health equity. Strengths of the paper include a clear conceptual motivation, a derivation from utility theory that is algebraically sound in the main text, the collapsibility property that supports aggregation across subgroups, and two detailed case studies with internal validation and bootstrap-based optimism correction. The code is made available, and the Discussion candidly acknowledges the main assumptions. However, the quantitative fairness conclusions are conditional on externally fixed treatment-benefit weights and clinically optimal thresholds, and the supplementary derivation of lambda as a relative risk reduction contains a step that does not preserve between-subgroup comparisons. These issues are load-bearing because the paper's central claim is that sNB levels and gaps can be meaningfully compared across subgroups.

major comments (3)
  1. [Supplementary 1.1, Eqs. (S1.9)-(S1.10)] The derivation of lambda as an RRR is not algebraically valid as written. Substituting the proxy-outcome utilities a=1-P(R=1|TP), c=1-P(R=1|FN), d=1 into Eq. (S1.3) gives U(t*) = 1 - P(R=1|FN)*pi + [P(R=1|FN)-P(R=1|TP)]/N * (TP - t/(1-t)FP). To obtain sNB = 1 - pi + lambda/N * (TP - t/(1-t)FP), the authors divide by P(R=1|FN) and then "reset the offset term to 1". The resetting step implicitly adds the constant 1 - 1/P(R=1|FN), which depends on the subgroup unless P(R=1|FN) is assumed constant across subgroups. Adding a subgroup-dependent constant changes the very between-subgroup comparisons that sNB is designed to quantify, so this is not a permissible linear transformation. The main-text Eq. (5) remains a coherent definition with a user-specified lambda, but the specific interpretation of lambda as an RRR, and hence the choice lambda=0.58 and lambda=0.20 in the case studies, needs either a corrected derivation or an explicit reframing as a modelling assumption rather than a derived quantity.
  2. [Results, Diabetes prognostic risk model; Results, Lung cancer screening algorithm; Discussion] All reported quantitative fairness claims are proportional to the single, externally imposed lambda. For example, the diabetes gap reduction of 48 (95% CI 32-66) and the average gap reduction of 24 (13-33) are exactly 0.58 times the corresponding differences in (TP - (t/(1-t))FP)/N, and the lung-cancer sNB increases in the most deprived quintile (5.0; 4.2-5.8) are exactly 0.20 times the standard net-benefit differences. The Discussion states that "ranges of values can be explored", but no sensitivity analysis is actually performed. Because a trial-derived RRR is not a full decision-analytic utility ratio (it omits false-positive harms, overdiagnosis, and quality-of-life effects), and because treatment effects may plausibly vary by ethnicity or deprivation, the reported gap reductions, model rankings, and the Pareto-front trade-offs in Figure 5 may not be robust. The authors should add a sensitivity analysis over a plausible range of lambda values, including subgroup-specific values, and report whether the qualitative conclusions survive, or substantially temper the quantitative claims.
  3. [Methods, Eq. (6)] The collapsibility property is presented as a substantive property of sNB, but Eq. (6) actually defines the aggregate sNB as the weighted average of subgroup sNBs. This is not the same as applying Eq. (5) to the pooled population with a single pooled lambda; if lambda differs by subgroup, the pooled sNB with one common lambda is not generally the weighted average of the subgroup sNBs. The statement that "no trade-off is required, provided that there are no resource constraints" follows from this definitional choice, but it should be stated as such. Otherwise readers may infer a stronger invariance property than the metric actually possesses.
minor comments (3)
  1. [Supplementary 1.2, Eq. (S2.2)] The piecewise definition of the weights is difficult to read because the cases are not clearly separated in the typeset equation. Please reformat so that the "if a_i=0" and "if a_i=1" branches are visually distinct, and define the scaling factor Pr(A=0)/Pr(A=1) more explicitly in the text.
  2. [Figure 5 and Supplementary Figure 9] The main text says the Pareto front in Figure 5 was "corrected for in-sample optimism", while Supplementary Figure 9 for the XGBoost model says the front was "found in the training dataset, without cross-validation" and then plotted in the validation dataset. Please reconcile these statements so readers know exactly which optimism-correction procedure applies to which model and figure.
  3. [Results, Diabetes prognostic risk model] The sentence reporting "Difference in the sNB between the models LogNoSA and LogSingleSA: 32; 16-49" would be clearer if it stated explicitly that this is the difference in the Asian subgroup and that a positive value favours LogSingleSA.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: sNB is a linear rescaling of standard decision-curve utility with externally fixed thresholds and λ; all fairness quantities are computed from bootstrap-corrected confusion matrices, not imposed by the metric.

full rationale

The sNB in Eq. (5) is derived from the standard utility expression (Eqs. 1-4 and Supplementary 1.1) by a monotone linear rescaling that reintroduces the prevalence term; it is not fitted to the model outputs being evaluated. The inputs to sNB are the population prevalence π, the confusion-matrix counts TP and FP at the chosen threshold, and the treatment-weight λ. Prevalence is a population descriptor; TP and FP come from models trained and validated on UK Biobank with bootstrap optimism correction; λ is fixed from published external estimates (RRR 0.58 for diabetes from the Diabetes Prevention Program, 0.20 for lung cancer from a policy review), and thresholds come from existing screening programmes. Thus the reported sNB levels, gap reductions, and Pareto fronts are computed from data rather than being equivalent to the paper's definitions by construction. The paper invokes no load-bearing self-citation and no uniqueness theorem from the authors' prior work; references to prior fairness frameworks (e.g., refs 14, 25, 26, 41) are external. The main vulnerability—that all quantitative conclusions scale with the single subgroup-invariant λ and the chosen thresholds—is an acknowledged assumption, stated as a limitation in the Discussion ('Our approach requires assumptions around the optimal threshold and the benefit of true positives'), and the formula explicitly permits subgroup-specific λ_g. That is a sensitivity and external-validity concern, not a circular step, so the appropriate score is 0.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claim depends on utility-theoretic assumptions, external treatment-effect values, and clinical thresholds that are not estimated in the paper. These are acceptable in decision curve analysis but make the fairness conclusions sensitive to user-supplied parameters.

free parameters (3)
  • lambda (treatment benefit weight) = 0.58 (diabetes), 0.20 (lung cancer)
    Chosen from external RRR estimates; affects magnitude of model benefit in sNB and therefore fairness gap sizes.
  • clinically optimal threshold t* = 15% (diabetes), 1.5% (lung cancer)
    Taken from clinical guidelines; sNB is evaluated at these thresholds in the main analyses.
  • screening capacity caps = 3% and 1% of the population
    Policy parameters imposed to illustrate Pareto trade-offs; the frontier shape depends on them.
assumptions (5)
  • domain assumption Clinical decisions can be summarized by a utility function over the confusion matrix with weights a, b, c, d.
    Standard decision theory; invoked at Eq. (1).
  • standard math The optimal threshold satisfies (1-t*)/t* = (a-c)/(d-b), allowing utility to be rewritten as net benefit.
    Derived from utility balance in Supplementary 1.1.
  • domain assumption lambda equals the relative risk reduction of a proxy outcome in treated true positives and is constant across subgroups.
    Central to sNB; if false, cross-subgroup comparisons are biased. Entered in Eq. (5) and Supplementary 1.1.
  • domain assumption Subgroup net benefit values are interpersonally comparable across groups.
    Needed for maximin equity criterion; a value judgment, not empirically testable.
  • domain assumption The clinically optimal thresholds (15% and 1.5%) are known and apply to all subgroups in the main analyses.
    Taken from clinical guidelines; subgroup-specific thresholds would change sNB results.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Understanding algorithmic fairness for clinical prediction in terms of subgroup net benefit and health equity." pith.science (2026). https://pith.science/paper/QUNG2VV6

@misc{pith2026241207879,
  author       = {Pith},
  title        = {Pith review of: Understanding algorithmic fairness for clinical prediction in terms of subgroup net benefit and health equity},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QUNG2VV6}},
  note         = {Machine review of arXiv:2412.07879}
}
read the original abstract

There are concerns about the fairness of clinical prediction models. 'Fair' models are defined as those for which their performance or predictions are not inappropriately influenced by protected attributes such as ethnicity, gender, or socio-economic status. Researchers have raised concerns that current algorithmic fairness paradigms enforce strict egalitarianism in healthcare, levelling down the performance of models in higher-performing subgroups instead of improving it in lower-performing ones. We propose assessing the fairness of a prediction model by expanding the concept of net benefit, using it to quantify and compare the clinical impact of a model in different subgroups. We use this to explore how a model distributes benefit across a population, its impact on health inequalities, and its role in the achievement of health equity. We show how resource constraints might introduce necessary trade-offs between health equity and other objectives of healthcare systems. We showcase our proposed approach with the development of two clinical prediction models: 1) a prognostic type 2 diabetes model used by clinicians to enrol patients into a preventive care lifestyle intervention programme, and 2) a lung cancer screening algorithm used to allocate diagnostic scans across the population. This approach helps modellers better understand if a model upholds health equity by considering its performance in a clinical and social context.

Figures

Figures reproduced from arXiv: 2412.07879 by the authors.

Figure 4
Figure 4. Without screening programme, the sNB outside the least deprived quintile was [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Critical Appraisal of Fairness Metrics in Clinical Predictive AI

    cs.LG 2025-06 accept novelty 4.0 of 10

    A scoping review of 62 fairness metrics for clinical predictive AI finds a fragmented, threshold-dependent landscape with only one clinical utility metric.

Reference graph

Works this paper leans on

59 extracted references · 38 canonical work pages · cited by 1 Pith paper

  1. [1]

    Predictably unequal: understanding and addressing concerns that algorithmic clinical prediction may increase health disparities

    Paulus JK, Kent DM. Predictably unequal: understanding and addressing concerns that algorithmic clinical prediction may increase health disparities. npj Digit Med. 2020;3(1):1-

  2. [2]

    Ensuring Fairness in Machine Learning to Advance Health Equity

    Rajkomar A, Hardt M, Howell MD, Corrado G, Chin MH. Ensuring Fairness in Machine Learning to Advance Health Equity. Ann Intern Med. 2018;169(12):866-872. doi:10.7326/M18-1990

  3. [3]

    Implementing Machine Learning in Health Care — Addressing Ethical Challenges

    Char DS, Shah NH, Magnus D. Implementing Machine Learning in Health Care — Addressing Ethical Challenges. N Engl J Med. 2018;378(11):981-983. doi:10.1056/NEJMp1714229

  4. [4]

    Understanding Potential Sources of Harm throughout the Machine Learning Life Cycle

    Suresh H, Guttag J. Understanding Potential Sources of Harm throughout the Machine Learning Life Cycle. MIT Case Studies in Social and Ethical Responsibilities of Computing. 2021;(Summer 2021). doi:10.21428/2c646de5.c16a07bb

  5. [5]

    Dissecting racial bias in an algorithm used to manage the health of populations

    Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453. doi:10.1126/science.aax2342

  6. [6]

    TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods

    Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378

  7. [7]

    Against Predictive Optimization: On the Legitimacy of Decision-Making Algorithms that Optimize Predictive Accuracy

    Wang A, Kapoor S, Barocas S, Narayanan A. Against Predictive Optimization: On the Legitimacy of Decision-Making Algorithms that Optimize Predictive Accuracy. In: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. FAccT ’23. Association for Computing Machinery; 2023:626. doi:10.1145/3593013.3594030

  8. [8]

    doi:10.1038/s41746-020-0304-9

Show all 59 references
  1. [9]

    A Survey on Bias and Fairness in Machine Learning

    Mehrabi N, Morstatter F, Saxena N, Lerman K, Galstyan A. A Survey on Bias and Fairness in Machine Learning. ACM Comput Surv. 2021;54(6):115:1-115:35. doi:10.1145/3457607

  2. [10]

    How We Analyzed the COMPAS Recidivism Algorithm

    Mattu JL Julia Angwin,Lauren Kirchner,Surya. How We Analyzed the COMPAS Recidivism Algorithm. ProPublica. Accessed June 1, 2022. https://www.propublica.org/article/how-we-analyzed-the-compas-recidivism- algorithm?token=BqO_ITYNAKmQwhj7daSusnn7aJDGaTWE

  3. [11]

    Algorithmic Bias? An Empirical Study into Apparent Gender- Based Discrimination in the Display of STEM Career Ads

    Lambrecht A, Tucker CE. Algorithmic Bias? An Empirical Study into Apparent Gender- Based Discrimination in the Display of STEM Career Ads. Social Science Research Network; 2018. doi:10.2139/ssrn.2852260

  4. [12]

    Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products

    Raji ID, Buolamwini J. Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products. In: Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society. ACM; 2019:429-435. doi:10.1145/3306618.3314244

  5. [13]

    AI Fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias

    Bellamy RKE, Dey K, Hind M, et al. AI Fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias. IBM Journal of Research and Development. 2019;63(4/5):4:1-4:15. doi:10.1147/JRD.2019.2942287 17

  6. [14]

    Fairness definitions explained

    Verma S, Rubin J. Fairness definitions explained. In: Proceedings of the International Workshop on Software Fairness. FairWare ’18. Association for Computing Machinery; 2018:1-7. doi:10.1145/3194770.3194776

  7. [15]

    The Unfairness of Fair Machine Learning: Levelling down and Strict Egalitarianism by Default

    Mittelstadt B, Wachter S, Russell C. The Unfairness of Fair Machine Learning: Levelling down and Strict Egalitarianism by Default. Social Science Research Network; 2023. Accessed February 8, 2023. https://papers.ssrn.com/abstract=4331652

  8. [16]

    Race Bias, Social Class Bias, and Gender Bias in Clinical Judgment

    Garb HN. Race Bias, Social Class Bias, and Gender Bias in Clinical Judgment. Clinical Psychology: Science and Practice. 1997;4(2):99-120. doi:10.1111/j.1468- 2850.1997.tb00104.x

  9. [17]

    A Historical Overview of Health Disparities and the Potential of eHealth Solutions

    Gibbons MC. A Historical Overview of Health Disparities and the Potential of eHealth Solutions. J Med Internet Res. 2005;7(5):e50. doi:10.2196/jmir.7.5.e50

  10. [18]

    The Root Causes of Health Inequity

    National Academies of Sciences E, Division H and M, Practice B on PH and PH, et al. The Root Causes of Health Inequity. In: Communities in Action: Pathways to Health Equity. National Academies Press (US); 2017. Accessed March 1, 2024. https://www.ncbi.nlm.nih.gov/books/NBK425845/

  11. [19]

    Who cares about equity in the NHS? BMJ

    Whitehead M. Who cares about equity in the NHS? BMJ. 1994;308(6939):1284-1287

  12. [20]

    Accessed March 1, 2024

    Health equity and its determinants. Accessed March 1, 2024. https://www.who.int/publications/m/item/health-equity-and-its-determinants

  13. [21]

    The meaning and goals of equity in health

    Chang WC. The meaning and goals of equity in health. Journal of Epidemiology & Community Health. 2002;56(7):488-491. doi:10.1136/jech.56.7.488

  14. [22]

    Health disparities and health equity: concepts and measurement

    Braveman P. Health disparities and health equity: concepts and measurement. Annual review of public health. 2006;27(1):167-194. doi:10.1146/annurev.publhealth.27.021405.102103

  15. [23]

    A glossary for health inequalities

    Kawachi I, Subramanian SV, Almeida-Filho N. A glossary for health inequalities. Journal of Epidemiology & Community Health. 2002;56(9):647-652. doi:10.1136/jech.56.9.647

  16. [24]

    Decision Curve Analysis: A Novel Method for Evaluating Prediction Models

    Vickers AJ, Elkin EB. Decision Curve Analysis: A Novel Method for Evaluating Prediction Models. Med Decis Making. 2006;26(6):565-574. doi:10.1177/0272989X06295361

  17. [25]

    Decision Making in Health and Medicine: Integrating Evidence and Values

    Hunink MGM, Weinstein MC, Wittenberg E, et al. Decision Making in Health and Medicine: Integrating Evidence and Values. 2nd ed. Cambridge University Press; 2014. doi:10.1017/CBO9781139506779

  18. [26]

    From Utilitarian to Rawlsian Designs for Algorithmic Fairness

    Rigobon DE. From Utilitarian to Rawlsian Designs for Algorithmic Fairness. Published online February 7, 2023. doi:10.48550/arXiv.2302.03567

  19. [27]

    Causal Conceptions of Fairness and their Consequences

    Nilforoshan H, Gaebler JD, Shroff R, Goel S. Causal Conceptions of Fairness and their Consequences. In: Proceedings of the 39th International Conference on Machine Learning. PMLR; 2022:16848-16887. Accessed December 16, 2022. https://proceedings.mlr.press/v162/nilforoshan22a.html

  20. [28]

    The Leicester Risk Assessment score for detecting undiagnosed Type 2 diabetes and impaired glucose regulation for use in a multiethnic UK setting

    Gray LJ, Taub NA, Khunti K, et al. The Leicester Risk Assessment score for detecting undiagnosed Type 2 diabetes and impaired glucose regulation for use in a multiethnic UK setting. Diabetic Medicine. 2010;27(8):887-895. doi:10.1111/j.1464- 5491.2010.03037.x 18

  21. [29]

    Accessed March 14, 2024

    The Leicester Diabetes Risk Score | arc-em.nihr.ac.uk. Accessed March 14, 2024. https://arc-em.nihr.ac.uk/clahrcs-store/leicester-diabetes-risk-score

  22. [30]

    Diabetes Care

    The Diabetes Prevention Program (DPP). Diabetes Care. 2002;25(12):2165-2171

  23. [31]

    Ethnicity and Type 2 diabetes in the UK

    Goff LM. Ethnicity and Type 2 diabetes in the UK. Diabetic Medicine. 2019;36(8):927-

  24. [32]

    Risk Prediction Model Versus United States Preventive Services Task Force Lung Cancer Screening Eligibility Criteria: Reducing Race Disparities

    Pasquinelli MM, Tammemägi MC, Kovitz KL, et al. Risk Prediction Model Versus United States Preventive Services Task Force Lung Cancer Screening Eligibility Criteria: Reducing Race Disparities. J Thorac Oncol. 2020;15(11):1738-1747. doi:10.1016/j.jtho.2020.08.006

  25. [33]

    Second round results from the Manchester ‘Lung Health Check’ community-based targeted lung cancer screening pilot

    Crosbie PA, Balata H, Evison M, et al. Second round results from the Manchester ‘Lung Health Check’ community-based targeted lung cancer screening pilot. Thorax. 2019;74(7):700-704. doi:10.1136/thoraxjnl-2018-212547

  26. [34]

    Implementing lung cancer screening: baseline results from a community-based ‘Lung Health Check’ pilot in deprived areas of Manchester

    Crosbie PA, Balata H, Evison M, et al. Implementing lung cancer screening: baseline results from a community-based ‘Lung Health Check’ pilot in deprived areas of Manchester. Thorax. 2019;74(4):405-409. doi:10.1136/thoraxjnl-2017-211377

  27. [35]

    Systematic Review of Lung Cancer Screening: Advancements and Strategies for Implementation

    Amicizia D, Piazza MF, Marchini F, et al. Systematic Review of Lung Cancer Screening: Advancements and Strategies for Implementation. Healthcare (Basel). 2023;11(14):2085. doi:10.3390/healthcare11142085

  28. [36]

    mice: Multivariate Imputation by Chained Equations in R

    Buuren S van, Groothuis-Oudshoorn K. mice: Multivariate Imputation by Chained Equations in R. Journal of Statistical Software. 2011;45:1-67. doi:10.18637/jss.v045.i03

  29. [37]

    UK Biobank: An Open Access Resource for Identifying the Causes of a Wide Range of Complex Diseases of Middle and Old Age

    Sudlow C, Gallacher J, Allen N, et al. UK Biobank: An Open Access Resource for Identifying the Causes of a Wide Range of Complex Diseases of Middle and Old Age. PLOS Medicine. 2015;12(3):e1001779. doi:10.1371/journal.pmed.1001779

  30. [38]

    Improving predictive inference under covariate shift by weighting the log- likelihood function

    Shimodaira H. Improving predictive inference under covariate shift by weighting the log- likelihood function. Journal of Statistical Planning and Inference. 2000;90(2):227-244. doi:10.1016/S0378-3758(00)00115-4

  31. [39]

    Transporting a Prediction Model for Use in a New Target Population

    Steingrimsson JA, Gatsonis C, Li B, Dahabreh IJ. Transporting a Prediction Model for Use in a New Target Population. Am J Epidemiol. 2023;192(2):296-304. doi:10.1093/aje/kwac128

  32. [40]

    Escaping the Impossibility of Fairness: From Formal to Substantive Algorithmic Fairness

    Green B. Escaping the Impossibility of Fairness: From Formal to Substantive Algorithmic Fairness. Philos Technol. 2022;35(4):90. doi:10.1007/s13347-022-00584-6

  33. [41]

    Evaluation of clinical prediction models (part 1): from development to external validation

    Collins GS, Dhiman P, Ma J, et al. Evaluation of clinical prediction models (part 1): from development to external validation. BMJ. 2024;384:e074819. doi:10.1136/bmj-2023- 074819

  34. [42]

    Assessing the Clinical Impact of Risk Prediction Models With Decision Curves: Guidance for Correct Interpretation and Appropriate Use

    Kerr KF, Brown MD, Zhu K, Janes H. Assessing the Clinical Impact of Risk Prediction Models With Decision Curves: Guidance for Correct Interpretation and Appropriate Use. J Clin Oncol. 2016;34(21):2534-2540. doi:10.1200/JCO.2015.65.5654 19

  35. [43]

    Net benefit, calibration, threshold selection, and training objectives for algorithmic fairness in healthcare

    Pfohl S, Xu Y, Foryciarz A, Ignatiadis N, Genkins J, Shah N. Net benefit, calibration, threshold selection, and training objectives for algorithmic fairness in healthcare. In: 2022 ACM Conference on Fairness, Accountability, and Transparency. ACM; 2022:1039-1052. doi:10.1145/3...

  36. [44]

    Race Corrections in Clinical Algorithms Can Help Correct for Racial Disparities in Data Quality

    Zink A, Obermeyer Z, Pierson E. Race Corrections in Clinical Algorithms Can Help Correct for Racial Disparities in Data Quality. Published online October 10, 2023:2023.03.31.23287926. doi:10.1101/2023.03.31.23287926

  37. [45]

    Assessing the net benefit of machine learning models in the presence of resource constraints

    Singh K, Shah NH, Vickers AJ. Assessing the net benefit of machine learning models in the presence of resource constraints. Journal of the American Medical Informatics Association. 2023;30(4):668-673. doi:10.1093/jamia/ocad006

  38. [46]

    Race and ethnicity – a part of the equation for personalized clinical decision making? Circ Cardiovasc Qual Outcomes

    Paulus JK, Kent DM. Race and ethnicity – a part of the equation for personalized clinical decision making? Circ Cardiovasc Qual Outcomes. 2017;10(7):e003823. doi:10.1161/CIRCOUTCOMES.117.003823

  39. [47]

    Adding social deprivation and family history to cardiovascular risk assessment: the ASSIGN score from the Scottish Heart Health Extended Cohort (SHHEC)

    Woodward M, Brindle P, Tunstall-Pedoe H, SIGN group on risk estimation. Adding social deprivation and family history to cardiovascular risk assessment: the ASSIGN score from the Scottish Heart Health Extended Cohort (SHHEC). Heart. 2007;93(2):172-176. doi:10.1136/hrt.2006.108167

  40. [48]

    Hidden in Plain Sight — Reconsidering the Use of Race Correction in Clinical Algorithms

    Vyas DA, Eisenstein LG, Jones DS. Hidden in Plain Sight — Reconsidering the Use of Race Correction in Clinical Algorithms. New England Journal of Medicine. 2020;383(9):874-882. doi:10.1056/NEJMms2004740

  41. [49]

    All else being equal, men and women are still not the same: using risk models to understand gender disparities in care

    Paulus JK, Shah ND, Kent DM. All else being equal, men and women are still not the same: using risk models to understand gender disparities in care. Circ Cardiovasc Qual Outcomes. 2015;8(3):317-320. doi:10.1161/CIRCOUTCOMES.115.001842

  42. [50]

    AI recognition of patient race in medical imaging: a modelling study

    Gichoya JW, Banerjee I, Bhimireddy AR, et al. AI recognition of patient race in medical imaging: a modelling study. The Lancet Digital Health. 2022;4(6):e406-e414. doi:10.1016/S2589-7500(22)00063-2

  43. [51]

    Toward the elimination of race-based medicine: replace race with racism as preeclampsia risk factor

    Ukoha EP, Snavely ME, Hahn MU, Steinauer JE, Bryant AS. Toward the elimination of race-based medicine: replace race with racism as preeclampsia risk factor. American Journal of Obstetrics & Gynecology. 2022;227(4):593-596. doi:10.1016/j.ajog.2022.05.048

  44. [52]

    Sample size for developing a prediction model with a binary outcome: targeting precise individual risk estimates to improve clinical decisions and fairness

    Riley RD, Collins GS, Whittle R, et al. Sample size for developing a prediction model with a binary outcome: targeting precise individual risk estimates to improve clinical decisions and fairness. Published online July 12, 2024. doi:10.48550/arXiv.2407.09293

  45. [53]

    The expected value of sample information calculations for external validation of risk prediction models

    Sadatsafavi M, Vickers AJ, Lee TY, Gustafson P, Wynants L. The expected value of sample information calculations for external validation of risk prediction models. Published online January 6, 2024. doi:10.48550/arXiv.2401.01849

  46. [55]

    A framework for digital health equity

    Richardson S, Lawrence K, Schoenthaler AM, Mann D. A framework for digital health equity. npj Digit Med. 2022;5(1):119. doi:10.1038/s41746-022-00663-0 20 Figure 1: Performance in 5-year T2 diabetes risk prediction in different ethnicities and overall, for three models: a logis...

  47. [56]

    Train a propensity score model, a logistic regression model with LASSO regularisation, to estimate each individual’s probability of belonging to the target subgroup, Pr (𝐴 = 1|𝑋 = 𝑥𝑖)

  48. [57]

    Use the propensity score model to generate weights 𝑤𝑖 for each individual, according to formula (S1.2). 31

  49. [58]

    Train the clinical prediction model using maximum likelihood minimisation, weighted with 𝑤𝑖

  50. [59]

    Validation metrics are calculated as usual, without weighting

    Repeat this for each subgroup of the population in order to create an ensemble of propensity-weighted models. Validation metrics are calculated as usual, without weighting. When correcting for optimism with bootstrapping for the logistic regression models, we only train the pr...

  51. [938]

    doi:10.1111/dme.13895

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.