REVIEW 3 major objections 3 minor 1 cited by
Understanding algorithmic fairness for clinical prediction in terms of subgroup net benefit and health equity
T0 review · 3 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper argues that algorithmic fairness for clinical prediction should be assessed by subgroup net benefit, a utility-based measure of how a model distributes clinical benefit across protected groups.
desk verdict A useful reframing of net benefit for subgroup fairness, but every equity-gap claim rests on one fixed lambda, and the paper never tests it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the subgroup net benefit, a re-scaled decision-curve net benefit that reintroduces the model-independent prevalence term $1-\pi$ and weights the conventional net benefit by $\lambda = RRR$, the relative risk reduction of the intervention at the decision threshold. It carries the argument by converting benefit from a property of predictions into a property of clinical decisions in each group: a group with $\text{sNB}=1$ is unburdened by the outcome, while a group with $\text{sNB}=0$ consists entirely of untreated false negatives. Its collapsibility, meaning the total sNB is the size-weighted mean of subgroup sNBs, is what allows the paper to treat overall benefit and maximin equity as aligned without resource constraints, and to make explicit the Pareto trade-off when capacity is capped.
What would settle it
Estimate subgroup-specific $\lambda$ from randomized trial data or from observational data with treatment assignment for the diabetes prevention and lung cancer screening interventions; if the subgroup-specific relative risk reductions differ from the single published RRRs of 0.58 and 0.20 by more than a few percent, recompute the sNB rankings and gap reductions. If the ordering of models by maximin sNB flips, or if the reported gap-narrowing effects reverse, the paper's fairness conclusions hinge on the single-$\lambda$ assumption.
Extended reading notes
Core claim
The paper's central claim is that the fairness of a clinical prediction model is a question about the distribution of clinical benefit, and that this distribution can be measured by the subgroup net benefit $\text{sNB}(t^*) = 1 - \pi + \frac{\lambda}{N}\left(TP(t^*) - \frac{t^*}{1-t^*}FP(t^*)\right)$, where $\pi$ is subgroup prevalence, $TP$ and $FP$ are true and false positives at threshold $t^*$, and $\lambda$ is the relative risk reduction of the intervention the model allocates. Because the term $1-\pi$ is retained, a group with high disease burden starts with lower sNB, so treating all groups equally is not enough to identify fairness. The metric is collapsible: the sNB of the whole population is the size-weighted average of subgroup sNBs, so improving the worst-off subgroup (maximin) need not sacrifice overall benefit when resources are unconstrained. Under capacity constraints, however, the paper shows via Pareto fronts that there is a genuine trade-off between overall net benefit and the sNB of the most disadvantaged group, and it proposes using the Pareto curve to make this trade-off explicit for decision-makers.
Load-bearing premise
The load-bearing assumption is that the relative risk reduction $\lambda$ of the intervention is known and identical across subgroups; if the true treatment effect varies by ethnicity or deprivation, or if it differs in the model-identified positive patients from the trial population that supplied $\lambda$, every sNB level, gap reduction, and Pareto trade-off in the paper would change.
Editorial extensions
If this is right
- Including the protected attribute as a predictor in the diabetes model raised the sNB of Asian and Black groups and narrowed the gap to the white group compared with omitting it, so model-building choices can be evaluated by their equity effect.
- In the lung cancer example, all three modelling strategies reduced the sNB gap between deprivation quintiles compared to screening no one, and the benefit concentrated in the most-deprived quintile.
- A random 'fair' screening policy reduced gaps only by lowering sNB in every group, illustrating that naive randomization harms all subgroups rather than levelling up.
- Under tight capacity caps of screening 3% or 1% of the population, overall and most-deprived-quintile sNB lie on a Pareto front, and the loss in overall benefit across the front is tiny compared to the gain in the worst-off group's sNB.
- Because sNB is collapsible, without resource constraints one can always achieve both maximum overall benefit and maximin equity by ensembling each subgroup's best model.
Reading between the lines
- Editorial inference: Because sNB depends on $\lambda$, the relative risk reduction, the same model can look fair or unfair under different assumptions about treatment effectiveness; reporting sNB as a function of $\lambda$ rather than at a single value would reveal how sensitive equity conclusions are.
- Editorial inference: The collapsibility of sNB suggests a natural fairness diagnostic for any deployed model: decompose the change in overall sNB into subgroup contributions and attribute the equity gap to differences in prevalence versus differences in predictive performance.
- Editorial inference: The Pareto-front formulation could be extended from thresholds to other levers such as screening intervals, outreach targeting, or model retraining, so that equity trade-offs are considered across the whole implementation, not just at the decision threshold.
- Editorial inference: A testable prediction of the framework is that models trained to maximize overall sNB under no constraints will rarely be the ones that minimize the sNB gap; if this holds across many clinical settings, gap-based maximin criteria would be needed in model selection even when resources are abundant.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a new fairness metric for clinical prediction models, the subgroup net benefit (sNB), defined in Eq. (5) as sNB = 1 - pi + (lambda/N)(TP - (t/(1-t))FP). The sNB reintroduces the prevalence term and a treatment-benefit weight lambda into standard decision-curve net benefit, allowing the clinical impact of a model to be quantified and compared across protected subgroups. The authors argue that fairness should be assessed through the distribution of sNB across groups, using gap reduction and maximin reasoning, and show that resource constraints can create a Pareto trade-off between overall net benefit and the most-underserved subgroup's sNB. The approach is illustrated with two UK Biobank case studies: a 5-year type-2 diabetes risk model compared across ethnicities, and a 6-year lung cancer screening model compared across Townsend deprivation quintiles. The paper finds that models including protected attributes improve sNB in disadvantaged groups and reduce health-inequality gaps.
Significance. If the sNB framework is accepted, it provides a decision-relevant, utility-based alternative to statistical fairness criteria such as demographic parity or equalized odds, and it directly links algorithmic fairness to health equity. Strengths of the paper include a clear conceptual motivation, a derivation from utility theory that is algebraically sound in the main text, the collapsibility property that supports aggregation across subgroups, and two detailed case studies with internal validation and bootstrap-based optimism correction. The code is made available, and the Discussion candidly acknowledges the main assumptions. However, the quantitative fairness conclusions are conditional on externally fixed treatment-benefit weights and clinically optimal thresholds, and the supplementary derivation of lambda as a relative risk reduction contains a step that does not preserve between-subgroup comparisons. These issues are load-bearing because the paper's central claim is that sNB levels and gaps can be meaningfully compared across subgroups.
major comments (3)
- [Supplementary 1.1, Eqs. (S1.9)-(S1.10)] The derivation of lambda as an RRR is not algebraically valid as written. Substituting the proxy-outcome utilities a=1-P(R=1|TP), c=1-P(R=1|FN), d=1 into Eq. (S1.3) gives U(t*) = 1 - P(R=1|FN)*pi + [P(R=1|FN)-P(R=1|TP)]/N * (TP - t/(1-t)FP). To obtain sNB = 1 - pi + lambda/N * (TP - t/(1-t)FP), the authors divide by P(R=1|FN) and then "reset the offset term to 1". The resetting step implicitly adds the constant 1 - 1/P(R=1|FN), which depends on the subgroup unless P(R=1|FN) is assumed constant across subgroups. Adding a subgroup-dependent constant changes the very between-subgroup comparisons that sNB is designed to quantify, so this is not a permissible linear transformation. The main-text Eq. (5) remains a coherent definition with a user-specified lambda, but the specific interpretation of lambda as an RRR, and hence the choice lambda=0.58 and lambda=0.20 in the case studies, needs either a corrected derivation or an explicit reframing as a modelling assumption rather than a derived quantity.
- [Results, Diabetes prognostic risk model; Results, Lung cancer screening algorithm; Discussion] All reported quantitative fairness claims are proportional to the single, externally imposed lambda. For example, the diabetes gap reduction of 48 (95% CI 32-66) and the average gap reduction of 24 (13-33) are exactly 0.58 times the corresponding differences in (TP - (t/(1-t))FP)/N, and the lung-cancer sNB increases in the most deprived quintile (5.0; 4.2-5.8) are exactly 0.20 times the standard net-benefit differences. The Discussion states that "ranges of values can be explored", but no sensitivity analysis is actually performed. Because a trial-derived RRR is not a full decision-analytic utility ratio (it omits false-positive harms, overdiagnosis, and quality-of-life effects), and because treatment effects may plausibly vary by ethnicity or deprivation, the reported gap reductions, model rankings, and the Pareto-front trade-offs in Figure 5 may not be robust. The authors should add a sensitivity analysis over a plausible range of lambda values, including subgroup-specific values, and report whether the qualitative conclusions survive, or substantially temper the quantitative claims.
- [Methods, Eq. (6)] The collapsibility property is presented as a substantive property of sNB, but Eq. (6) actually defines the aggregate sNB as the weighted average of subgroup sNBs. This is not the same as applying Eq. (5) to the pooled population with a single pooled lambda; if lambda differs by subgroup, the pooled sNB with one common lambda is not generally the weighted average of the subgroup sNBs. The statement that "no trade-off is required, provided that there are no resource constraints" follows from this definitional choice, but it should be stated as such. Otherwise readers may infer a stronger invariance property than the metric actually possesses.
minor comments (3)
- [Supplementary 1.2, Eq. (S2.2)] The piecewise definition of the weights is difficult to read because the cases are not clearly separated in the typeset equation. Please reformat so that the "if a_i=0" and "if a_i=1" branches are visually distinct, and define the scaling factor Pr(A=0)/Pr(A=1) more explicitly in the text.
- [Figure 5 and Supplementary Figure 9] The main text says the Pareto front in Figure 5 was "corrected for in-sample optimism", while Supplementary Figure 9 for the XGBoost model says the front was "found in the training dataset, without cross-validation" and then plotted in the validation dataset. Please reconcile these statements so readers know exactly which optimism-correction procedure applies to which model and figure.
- [Results, Diabetes prognostic risk model] The sentence reporting "Difference in the sNB between the models LogNoSA and LogSingleSA: 32; 16-49" would be clearer if it stated explicitly that this is the difference in the Asian subgroup and that a positive value favours LogSingleSA.
Circularity Check
No circularity: sNB is a linear rescaling of standard decision-curve utility with externally fixed thresholds and λ; all fairness quantities are computed from bootstrap-corrected confusion matrices, not imposed by the metric.
full rationale
The sNB in Eq. (5) is derived from the standard utility expression (Eqs. 1-4 and Supplementary 1.1) by a monotone linear rescaling that reintroduces the prevalence term; it is not fitted to the model outputs being evaluated. The inputs to sNB are the population prevalence π, the confusion-matrix counts TP and FP at the chosen threshold, and the treatment-weight λ. Prevalence is a population descriptor; TP and FP come from models trained and validated on UK Biobank with bootstrap optimism correction; λ is fixed from published external estimates (RRR 0.58 for diabetes from the Diabetes Prevention Program, 0.20 for lung cancer from a policy review), and thresholds come from existing screening programmes. Thus the reported sNB levels, gap reductions, and Pareto fronts are computed from data rather than being equivalent to the paper's definitions by construction. The paper invokes no load-bearing self-citation and no uniqueness theorem from the authors' prior work; references to prior fairness frameworks (e.g., refs 14, 25, 26, 41) are external. The main vulnerability—that all quantitative conclusions scale with the single subgroup-invariant λ and the chosen thresholds—is an acknowledged assumption, stated as a limitation in the Discussion ('Our approach requires assumptions around the optimal threshold and the benefit of true positives'), and the formula explicitly permits subgroup-specific λ_g. That is a sensitivity and external-validity concern, not a circular step, so the appropriate score is 0.
Assumptions & free parameters
free parameters (3)
- lambda (treatment benefit weight) =
0.58 (diabetes), 0.20 (lung cancer)
- clinically optimal threshold t* =
15% (diabetes), 1.5% (lung cancer)
- screening capacity caps =
3% and 1% of the population
assumptions (5)
- domain assumption Clinical decisions can be summarized by a utility function over the confusion matrix with weights a, b, c, d.
- standard math The optimal threshold satisfies (1-t*)/t* = (a-c)/(d-b), allowing utility to be rewritten as net benefit.
- domain assumption lambda equals the relative risk reduction of a proxy outcome in treated true positives and is constant across subgroups.
- domain assumption Subgroup net benefit values are interpersonally comparable across groups.
- domain assumption The clinically optimal thresholds (15% and 1.5%) are known and apply to all subgroups in the main analyses.
Cite this review
Pith. "Pith review of Understanding algorithmic fairness for clinical prediction in terms of subgroup net benefit and health equity." pith.science (2026). https://pith.science/paper/QUNG2VV6
@misc{pith2026241207879,
author = {Pith},
title = {Pith review of: Understanding algorithmic fairness for clinical prediction in terms of subgroup net benefit and health equity},
year = {2026},
howpublished = {\url{https://pith.science/paper/QUNG2VV6}},
note = {Machine review of arXiv:2412.07879}
}
read the original abstract
There are concerns about the fairness of clinical prediction models. 'Fair' models are defined as those for which their performance or predictions are not inappropriately influenced by protected attributes such as ethnicity, gender, or socio-economic status. Researchers have raised concerns that current algorithmic fairness paradigms enforce strict egalitarianism in healthcare, levelling down the performance of models in higher-performing subgroups instead of improving it in lower-performing ones. We propose assessing the fairness of a prediction model by expanding the concept of net benefit, using it to quantify and compare the clinical impact of a model in different subgroups. We use this to explore how a model distributes benefit across a population, its impact on health inequalities, and its role in the achievement of health equity. We show how resource constraints might introduce necessary trade-offs between health equity and other objectives of healthcare systems. We showcase our proposed approach with the development of two clinical prediction models: 1) a prognostic type 2 diabetes model used by clinicians to enrol patients into a preventive care lifestyle intervention programme, and 2) a lung cancer screening algorithm used to allocate diagnostic scans across the population. This approach helps modellers better understand if a model upholds health equity by considering its performance in a clinical and social context.
Figures
Forward citations
Cited by 1 Pith paper
-
Critical Appraisal of Fairness Metrics in Clinical Predictive AI
A scoping review of 62 fairness metrics for clinical predictive AI finds a fragmented, threshold-dependent landscape with only one clinical utility metric.
Reference graph
Works this paper leans on
-
[1]
Paulus JK, Kent DM. Predictably unequal: understanding and addressing concerns that algorithmic clinical prediction may increase health disparities. npj Digit Med. 2020;3(1):1-
work page 2020
-
[2]
Ensuring Fairness in Machine Learning to Advance Health Equity
Rajkomar A, Hardt M, Howell MD, Corrado G, Chin MH. Ensuring Fairness in Machine Learning to Advance Health Equity. Ann Intern Med. 2018;169(12):866-872. doi:10.7326/M18-1990
doi:10.7326/m18-1990 2018
-
[3]
Implementing Machine Learning in Health Care — Addressing Ethical Challenges
Char DS, Shah NH, Magnus D. Implementing Machine Learning in Health Care — Addressing Ethical Challenges. N Engl J Med. 2018;378(11):981-983. doi:10.1056/NEJMp1714229
-
[4]
Understanding Potential Sources of Harm throughout the Machine Learning Life Cycle
Suresh H, Guttag J. Understanding Potential Sources of Harm throughout the Machine Learning Life Cycle. MIT Case Studies in Social and Ethical Responsibilities of Computing. 2021;(Summer 2021). doi:10.21428/2c646de5.c16a07bb
-
[5]
Dissecting racial bias in an algorithm used to manage the health of populations
Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453. doi:10.1126/science.aax2342
-
[6]
Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378
-
[7]
Wang A, Kapoor S, Barocas S, Narayanan A. Against Predictive Optimization: On the Legitimacy of Decision-Making Algorithms that Optimize Predictive Accuracy. In: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. FAccT ’23. Association for Computing Machinery; 2023:626. doi:10.1145/3593013.3594030
arXiv 2023
-
[8]
doi:10.1038/s41746-020-0304-9
Show all 59 references
-
[9]
A Survey on Bias and Fairness in Machine Learning
Mehrabi N, Morstatter F, Saxena N, Lerman K, Galstyan A. A Survey on Bias and Fairness in Machine Learning. ACM Comput Surv. 2021;54(6):115:1-115:35. doi:10.1145/3457607
2021 doi
-
[10]
How We Analyzed the COMPAS Recidivism Algorithm
Mattu JL Julia Angwin,Lauren Kirchner,Surya. How We Analyzed the COMPAS Recidivism Algorithm. ProPublica. Accessed June 1, 2022. https://www.propublica.org/article/how-we-analyzed-the-compas-recidivism- algorithm?token=BqO_ITYNAKmQwhj7daSusnn7aJDGaTWE
2022
-
[11]
Algorithmic Bias? An Empirical Study into Apparent Gender- Based Discrimination in the Display of STEM Career Ads
Lambrecht A, Tucker CE. Algorithmic Bias? An Empirical Study into Apparent Gender- Based Discrimination in the Display of STEM Career Ads. Social Science Research Network; 2018. doi:10.2139/ssrn.2852260
2018 doi
-
[12]
Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products
Raji ID, Buolamwini J. Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products. In: Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society. ACM; 2019:429-435. doi:10.1145/3306618.3314244
2019
-
[13]
AI Fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias
Bellamy RKE, Dey K, Hind M, et al. AI Fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias. IBM Journal of Research and Development. 2019;63(4/5):4:1-4:15. doi:10.1147/JRD.2019.2942287 17
2019
-
[14]
Fairness definitions explained
Verma S, Rubin J. Fairness definitions explained. In: Proceedings of the International Workshop on Software Fairness. FairWare ’18. Association for Computing Machinery; 2018:1-7. doi:10.1145/3194770.3194776
2018
-
[15]
The Unfairness of Fair Machine Learning: Levelling down and Strict Egalitarianism by Default
Mittelstadt B, Wachter S, Russell C. The Unfairness of Fair Machine Learning: Levelling down and Strict Egalitarianism by Default. Social Science Research Network; 2023. Accessed February 8, 2023. https://papers.ssrn.com/abstract=4331652
2023
-
[16]
Race Bias, Social Class Bias, and Gender Bias in Clinical Judgment
Garb HN. Race Bias, Social Class Bias, and Gender Bias in Clinical Judgment. Clinical Psychology: Science and Practice. 1997;4(2):99-120. doi:10.1111/j.1468- 2850.1997.tb00104.x
1997
-
[17]
A Historical Overview of Health Disparities and the Potential of eHealth Solutions
Gibbons MC. A Historical Overview of Health Disparities and the Potential of eHealth Solutions. J Med Internet Res. 2005;7(5):e50. doi:10.2196/jmir.7.5.e50
2005 doi
-
[18]
The Root Causes of Health Inequity
National Academies of Sciences E, Division H and M, Practice B on PH and PH, et al. The Root Causes of Health Inequity. In: Communities in Action: Pathways to Health Equity. National Academies Press (US); 2017. Accessed March 1, 2024. https://www.ncbi.nlm.nih.gov/books/NBK425845/
2017
-
[19]
Who cares about equity in the NHS? BMJ
Whitehead M. Who cares about equity in the NHS? BMJ. 1994;308(6939):1284-1287
1994
-
[20]
Accessed March 1, 2024
Health equity and its determinants. Accessed March 1, 2024. https://www.who.int/publications/m/item/health-equity-and-its-determinants
2024
-
[21]
The meaning and goals of equity in health
Chang WC. The meaning and goals of equity in health. Journal of Epidemiology & Community Health. 2002;56(7):488-491. doi:10.1136/jech.56.7.488
2002 doi
-
[22]
Health disparities and health equity: concepts and measurement
Braveman P. Health disparities and health equity: concepts and measurement. Annual review of public health. 2006;27(1):167-194. doi:10.1146/annurev.publhealth.27.021405.102103
2006
-
[23]
A glossary for health inequalities
Kawachi I, Subramanian SV, Almeida-Filho N. A glossary for health inequalities. Journal of Epidemiology & Community Health. 2002;56(9):647-652. doi:10.1136/jech.56.9.647
2002 doi
-
[24]
Decision Curve Analysis: A Novel Method for Evaluating Prediction Models
Vickers AJ, Elkin EB. Decision Curve Analysis: A Novel Method for Evaluating Prediction Models. Med Decis Making. 2006;26(6):565-574. doi:10.1177/0272989X06295361
2006 doi
-
[25]
Decision Making in Health and Medicine: Integrating Evidence and Values
Hunink MGM, Weinstein MC, Wittenberg E, et al. Decision Making in Health and Medicine: Integrating Evidence and Values. 2nd ed. Cambridge University Press; 2014. doi:10.1017/CBO9781139506779
2014 doi
- [26]
-
[27]
Causal Conceptions of Fairness and their Consequences
Nilforoshan H, Gaebler JD, Shroff R, Goel S. Causal Conceptions of Fairness and their Consequences. In: Proceedings of the 39th International Conference on Machine Learning. PMLR; 2022:16848-16887. Accessed December 16, 2022. https://proceedings.mlr.press/v162/nilforoshan22a.html
2022
-
[28]
The Leicester Risk Assessment score for detecting undiagnosed Type 2 diabetes and impaired glucose regulation for use in a multiethnic UK setting
Gray LJ, Taub NA, Khunti K, et al. The Leicester Risk Assessment score for detecting undiagnosed Type 2 diabetes and impaired glucose regulation for use in a multiethnic UK setting. Diabetic Medicine. 2010;27(8):887-895. doi:10.1111/j.1464- 5491.2010.03037.x 18
2010
-
[29]
Accessed March 14, 2024
The Leicester Diabetes Risk Score | arc-em.nihr.ac.uk. Accessed March 14, 2024. https://arc-em.nihr.ac.uk/clahrcs-store/leicester-diabetes-risk-score
2024
-
[30]
Diabetes Care
The Diabetes Prevention Program (DPP). Diabetes Care. 2002;25(12):2165-2171
2002
-
[31]
Ethnicity and Type 2 diabetes in the UK
Goff LM. Ethnicity and Type 2 diabetes in the UK. Diabetic Medicine. 2019;36(8):927-
2019
-
[32]
Risk Prediction Model Versus United States Preventive Services Task Force Lung Cancer Screening Eligibility Criteria: Reducing Race Disparities
Pasquinelli MM, Tammemägi MC, Kovitz KL, et al. Risk Prediction Model Versus United States Preventive Services Task Force Lung Cancer Screening Eligibility Criteria: Reducing Race Disparities. J Thorac Oncol. 2020;15(11):1738-1747. doi:10.1016/j.jtho.2020.08.006
2020 doi
-
[33]
Second round results from the Manchester ‘Lung Health Check’ community-based targeted lung cancer screening pilot
Crosbie PA, Balata H, Evison M, et al. Second round results from the Manchester ‘Lung Health Check’ community-based targeted lung cancer screening pilot. Thorax. 2019;74(7):700-704. doi:10.1136/thoraxjnl-2018-212547
2019 doi
-
[34]
Implementing lung cancer screening: baseline results from a community-based ‘Lung Health Check’ pilot in deprived areas of Manchester
Crosbie PA, Balata H, Evison M, et al. Implementing lung cancer screening: baseline results from a community-based ‘Lung Health Check’ pilot in deprived areas of Manchester. Thorax. 2019;74(4):405-409. doi:10.1136/thoraxjnl-2017-211377
2019 doi
-
[35]
Systematic Review of Lung Cancer Screening: Advancements and Strategies for Implementation
Amicizia D, Piazza MF, Marchini F, et al. Systematic Review of Lung Cancer Screening: Advancements and Strategies for Implementation. Healthcare (Basel). 2023;11(14):2085. doi:10.3390/healthcare11142085
2023 doi
-
[36]
mice: Multivariate Imputation by Chained Equations in R
Buuren S van, Groothuis-Oudshoorn K. mice: Multivariate Imputation by Chained Equations in R. Journal of Statistical Software. 2011;45:1-67. doi:10.18637/jss.v045.i03
2011 doi
-
[37]
UK Biobank: An Open Access Resource for Identifying the Causes of a Wide Range of Complex Diseases of Middle and Old Age
Sudlow C, Gallacher J, Allen N, et al. UK Biobank: An Open Access Resource for Identifying the Causes of a Wide Range of Complex Diseases of Middle and Old Age. PLOS Medicine. 2015;12(3):e1001779. doi:10.1371/journal.pmed.1001779
2015 doi
-
[38]
Improving predictive inference under covariate shift by weighting the log- likelihood function
Shimodaira H. Improving predictive inference under covariate shift by weighting the log- likelihood function. Journal of Statistical Planning and Inference. 2000;90(2):227-244. doi:10.1016/S0378-3758(00)00115-4
-
[39]
Transporting a Prediction Model for Use in a New Target Population
Steingrimsson JA, Gatsonis C, Li B, Dahabreh IJ. Transporting a Prediction Model for Use in a New Target Population. Am J Epidemiol. 2023;192(2):296-304. doi:10.1093/aje/kwac128
2023 doi
-
[40]
Escaping the Impossibility of Fairness: From Formal to Substantive Algorithmic Fairness
Green B. Escaping the Impossibility of Fairness: From Formal to Substantive Algorithmic Fairness. Philos Technol. 2022;35(4):90. doi:10.1007/s13347-022-00584-6
2022 doi
-
[41]
Evaluation of clinical prediction models (part 1): from development to external validation
Collins GS, Dhiman P, Ma J, et al. Evaluation of clinical prediction models (part 1): from development to external validation. BMJ. 2024;384:e074819. doi:10.1136/bmj-2023- 074819
2024 doi
-
[42]
Assessing the Clinical Impact of Risk Prediction Models With Decision Curves: Guidance for Correct Interpretation and Appropriate Use
Kerr KF, Brown MD, Zhu K, Janes H. Assessing the Clinical Impact of Risk Prediction Models With Decision Curves: Guidance for Correct Interpretation and Appropriate Use. J Clin Oncol. 2016;34(21):2534-2540. doi:10.1200/JCO.2015.65.5654 19
2016 doi
-
[43]
Net benefit, calibration, threshold selection, and training objectives for algorithmic fairness in healthcare
Pfohl S, Xu Y, Foryciarz A, Ignatiadis N, Genkins J, Shah N. Net benefit, calibration, threshold selection, and training objectives for algorithmic fairness in healthcare. In: 2022 ACM Conference on Fairness, Accountability, and Transparency. ACM; 2022:1039-1052. doi:10.1145/3...
2022
-
[44]
Race Corrections in Clinical Algorithms Can Help Correct for Racial Disparities in Data Quality
Zink A, Obermeyer Z, Pierson E. Race Corrections in Clinical Algorithms Can Help Correct for Racial Disparities in Data Quality. Published online October 10, 2023:2023.03.31.23287926. doi:10.1101/2023.03.31.23287926
2023 doi
-
[45]
Assessing the net benefit of machine learning models in the presence of resource constraints
Singh K, Shah NH, Vickers AJ. Assessing the net benefit of machine learning models in the presence of resource constraints. Journal of the American Medical Informatics Association. 2023;30(4):668-673. doi:10.1093/jamia/ocad006
2023 doi
-
[46]
Race and ethnicity – a part of the equation for personalized clinical decision making? Circ Cardiovasc Qual Outcomes
Paulus JK, Kent DM. Race and ethnicity – a part of the equation for personalized clinical decision making? Circ Cardiovasc Qual Outcomes. 2017;10(7):e003823. doi:10.1161/CIRCOUTCOMES.117.003823
2017 doi
-
[47]
Adding social deprivation and family history to cardiovascular risk assessment: the ASSIGN score from the Scottish Heart Health Extended Cohort (SHHEC)
Woodward M, Brindle P, Tunstall-Pedoe H, SIGN group on risk estimation. Adding social deprivation and family history to cardiovascular risk assessment: the ASSIGN score from the Scottish Heart Health Extended Cohort (SHHEC). Heart. 2007;93(2):172-176. doi:10.1136/hrt.2006.108167
2007 arXiv
-
[48]
Hidden in Plain Sight — Reconsidering the Use of Race Correction in Clinical Algorithms
Vyas DA, Eisenstein LG, Jones DS. Hidden in Plain Sight — Reconsidering the Use of Race Correction in Clinical Algorithms. New England Journal of Medicine. 2020;383(9):874-882. doi:10.1056/NEJMms2004740
2020 doi
-
[49]
All else being equal, men and women are still not the same: using risk models to understand gender disparities in care
Paulus JK, Shah ND, Kent DM. All else being equal, men and women are still not the same: using risk models to understand gender disparities in care. Circ Cardiovasc Qual Outcomes. 2015;8(3):317-320. doi:10.1161/CIRCOUTCOMES.115.001842
2015 doi
-
[50]
AI recognition of patient race in medical imaging: a modelling study
Gichoya JW, Banerjee I, Bhimireddy AR, et al. AI recognition of patient race in medical imaging: a modelling study. The Lancet Digital Health. 2022;4(6):e406-e414. doi:10.1016/S2589-7500(22)00063-2
2022 doi
-
[51]
Toward the elimination of race-based medicine: replace race with racism as preeclampsia risk factor
Ukoha EP, Snavely ME, Hahn MU, Steinauer JE, Bryant AS. Toward the elimination of race-based medicine: replace race with racism as preeclampsia risk factor. American Journal of Obstetrics & Gynecology. 2022;227(4):593-596. doi:10.1016/j.ajog.2022.05.048
2022 doi
-
[52]
Sample size for developing a prediction model with a binary outcome: targeting precise individual risk estimates to improve clinical decisions and fairness
Riley RD, Collins GS, Whittle R, et al. Sample size for developing a prediction model with a binary outcome: targeting precise individual risk estimates to improve clinical decisions and fairness. Published online July 12, 2024. doi:10.48550/arXiv.2407.09293
-
[53]
The expected value of sample information calculations for external validation of risk prediction models
Sadatsafavi M, Vickers AJ, Lee TY, Gustafson P, Wynants L. The expected value of sample information calculations for external validation of risk prediction models. Published online January 6, 2024. doi:10.48550/arXiv.2401.01849
-
[55]
A framework for digital health equity
Richardson S, Lawrence K, Schoenthaler AM, Mann D. A framework for digital health equity. npj Digit Med. 2022;5(1):119. doi:10.1038/s41746-022-00663-0 20 Figure 1: Performance in 5-year T2 diabetes risk prediction in different ethnicities and overall, for three models: a logis...
2022 doi
-
[56]
Train a propensity score model, a logistic regression model with LASSO regularisation, to estimate each individual’s probability of belonging to the target subgroup, Pr (𝐴 = 1|𝑋 = 𝑥𝑖)
-
[57]
Use the propensity score model to generate weights 𝑤𝑖 for each individual, according to formula (S1.2). 31
-
[58]
Train the clinical prediction model using maximum likelihood minimisation, weighted with 𝑤𝑖
-
[59]
Validation metrics are calculated as usual, without weighting
Repeat this for each subgroup of the population in order to create an ensemble of propensity-weighted models. Validation metrics are calculated as usual, without weighting. When correcting for optimism with bootstrapping for the logistic regression models, we only train the pr...
-
[938]
doi:10.1111/dme.13895
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.