{"id":"748b6168-2484-46ee-b40d-11cf609ff7e2","arxiv_id":"2412.07879","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Subgroup net benefit reparameterizes net benefit by including baseline prevalence, allowing clinical prediction models to be compared across protected subgroups on a utility scale.","lead":"This paper proposes a fairness metric for clinical prediction models called the 'subgroup net benefit', which adds each group's baseline disease burden to the standard net benefit of a model. It applies the metric to type 2 diabetes risk and lung cancer screening, and shows how resource limits force trade-offs between overall benefit and equity.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"All quantitative fairness claims depend on a single externally imposed, subgroup-invariant treatment weight λ (Eq. 5, Suppl. 1.1), yet no sensitivity analysis is provided to show conclusions survive plausible subgroup-specific values.","rationale":"The reader's conditional verdict is appropriate. The central derivation from utility Eq. (1) to sNB Eq. (5) is algebraically correct, and the two UK Biobank analyses are reproducible from provided code. The collapsibility property is a trivial consequence of linearity, so it adds little; the main contribution is the framing and the application, not the mathematical novelty. The most serious risk is the external value of λ. Because λ multiplies the entire model-benefit term, all quantitative effect sizes are proportional to it; the sign of a gap change is robust to a constant λ, but the magnitude, the maximin comparisons, and the resource-constrained Pareto front are not. Moreover, equating λ with an RRR of a surrogate outcome is a strong substantive assumption about utilities, not just a tuning parameter. The authors acknowledge this limitation in the Discussion but do not provide the promised sensitivity analysis. A conditional acceptance requiring that sensitivity analysis is therefore the right call; this concern does not invalidate the conceptual framework, because the framework is explicitly parameter-dependent and can be re-run.","tokens_in":19273,"tokens_out":8493,"duration_ms":90155,"concrete_test":"Use the public GitHub code to recompute the diabetes (Figure 2) and lung-cancer (Figure 5) results with subgroup-specific λ_g values: for diabetes, draw λ from published RRR estimates by ethnicity or vary λ over 0.2–0.8; for lung cancer, vary over 0.1–0.3. Report whether the direction and 95% CI of the sNB gap reductions (diabetes) and the set of Pareto-optimal thresholds (lung cancer) are preserved. If any conclusion reverses or loses significance, the central claim is conditional on an unverified parameter.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The sNB in Eq. 5 is sNB = 1 − π + (λ/N)(TP − (t/(1−t))FP). Every reported benefit level, gap reduction, and Pareto front therefore scales with λ; all comparisons between no-treatment and model policies, and all gap reductions (e.g., 48 and 24 in the diabetes example), are exactly λ times the corresponding standard net-benefit differences. The paper fixes λ to a single published RRR (0.58 for diabetes, 0.20 for lung cancer) and assumes it applies equally to every model-identified true positive in every ethnic or deprivation subgroup. This is load-bearing because: (1) RRR is a relative risk reduction from a trial population, not a decision-analytic utility ratio; it omits false-positive harms, overdiagnosis, and quality-of-life effects, especially in lung cancer screening. (2) Treatment effects are plausibly heterogeneous across subgroups (ethnicity, deprivation, comorbidity), and if λ differs by group, the ordering of sNB levels and the shape of the Pareto-optimal trade-off in Figure 5 can change. The paper states that ranges of values ‘can be explored’ but never performs that exploration. The Discussion explicitly lists the assumption about true-positive benefit as a limitation, so this is an acknowledged, untested premise rather than a hidden error. The mathematical derivation is sound conditional on that premise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a new fairness metric for clinical prediction models, the subgroup net benefit (sNB), defined in Eq. (5) as sNB = 1 - pi + (lambda/N)(TP - (t/(1-t))FP). The sNB reintroduces the prevalence term and a treatment-benefit weight lambda into standard decision-curve net benefit, allowing the clinical impact of a model to be quantified and compared across protected subgroups. The authors argue that fairness should be assessed through the distribution of sNB across groups, using gap reduction and maximin reasoning, and show that resource constraints can create a Pareto trade-off between overall net benefit and the most-underserved subgroup's sNB. The approach is illustrated with two UK Biobank case studies: a 5-year type-2 diabetes risk model compared across ethnicities, and a 6-year lung cancer screening model compared across Townsend deprivation quintiles. The paper finds that models including protected attributes improve sNB in disadvantaged groups and reduce health-inequality gaps.","tokens_in":19503,"tokens_out":9152,"duration_ms":105101,"significance":"If the sNB framework is accepted, it provides a decision-relevant, utility-based alternative to statistical fairness criteria such as demographic parity or equalized odds, and it directly links algorithmic fairness to health equity. Strengths of the paper include a clear conceptual motivation, a derivation from utility theory that is algebraically sound in the main text, the collapsibility property that supports aggregation across subgroups, and two detailed case studies with internal validation and bootstrap-based optimism correction. The code is made available, and the Discussion candidly acknowledges the main assumptions. However, the quantitative fairness conclusions are conditional on externally fixed treatment-benefit weights and clinically optimal thresholds, and the supplementary derivation of lambda as a relative risk reduction contains a step that does not preserve between-subgroup comparisons. These issues are load-bearing because the paper's central claim is that sNB levels and gaps can be meaningfully compared across subgroups.","major_comments":[{"comment":"The derivation of lambda as an RRR is not algebraically valid as written. Substituting the proxy-outcome utilities a=1-P(R=1|TP), c=1-P(R=1|FN), d=1 into Eq. (S1.3) gives U(t*) = 1 - P(R=1|FN)*pi + [P(R=1|FN)-P(R=1|TP)]/N * (TP - t/(1-t)FP). To obtain sNB = 1 - pi + lambda/N * (TP - t/(1-t)FP), the authors divide by P(R=1|FN) and then \"reset the offset term to 1\". The resetting step implicitly adds the constant 1 - 1/P(R=1|FN), which depends on the subgroup unless P(R=1|FN) is assumed constant across subgroups. Adding a subgroup-dependent constant changes the very between-subgroup comparisons that sNB is designed to quantify, so this is not a permissible linear transformation. The main-text Eq. (5) remains a coherent definition with a user-specified lambda, but the specific interpretation of lambda as an RRR, and hence the choice lambda=0.58 and lambda=0.20 in the case studies, needs either a corrected derivation or an explicit reframing as a modelling assumption rather than a derived quantity.","section":"Supplementary 1.1, Eqs. (S1.9)-(S1.10)"},{"comment":"All reported quantitative fairness claims are proportional to the single, externally imposed lambda. For example, the diabetes gap reduction of 48 (95% CI 32-66) and the average gap reduction of 24 (13-33) are exactly 0.58 times the corresponding differences in (TP - (t/(1-t))FP)/N, and the lung-cancer sNB increases in the most deprived quintile (5.0; 4.2-5.8) are exactly 0.20 times the standard net-benefit differences. The Discussion states that \"ranges of values can be explored\", but no sensitivity analysis is actually performed. Because a trial-derived RRR is not a full decision-analytic utility ratio (it omits false-positive harms, overdiagnosis, and quality-of-life effects), and because treatment effects may plausibly vary by ethnicity or deprivation, the reported gap reductions, model rankings, and the Pareto-front trade-offs in Figure 5 may not be robust. The authors should add a sensitivity analysis over a plausible range of lambda values, including subgroup-specific values, and report whether the qualitative conclusions survive, or substantially temper the quantitative claims.","section":"Results, Diabetes prognostic risk model; Results, Lung cancer screening algorithm; Discussion"},{"comment":"The collapsibility property is presented as a substantive property of sNB, but Eq. (6) actually defines the aggregate sNB as the weighted average of subgroup sNBs. This is not the same as applying Eq. (5) to the pooled population with a single pooled lambda; if lambda differs by subgroup, the pooled sNB with one common lambda is not generally the weighted average of the subgroup sNBs. The statement that \"no trade-off is required, provided that there are no resource constraints\" follows from this definitional choice, but it should be stated as such. Otherwise readers may infer a stronger invariance property than the metric actually possesses.","section":"Methods, Eq. (6)"}],"minor_comments":[{"comment":"The piecewise definition of the weights is difficult to read because the cases are not clearly separated in the typeset equation. Please reformat so that the \"if a_i=0\" and \"if a_i=1\" branches are visually distinct, and define the scaling factor Pr(A=0)/Pr(A=1) more explicitly in the text.","section":"Supplementary 1.2, Eq. (S2.2)"},{"comment":"The main text says the Pareto front in Figure 5 was \"corrected for in-sample optimism\", while Supplementary Figure 9 for the XGBoost model says the front was \"found in the training dataset, without cross-validation\" and then plotted in the validation dataset. Please reconcile these statements so readers know exactly which optimism-correction procedure applies to which model and figure.","section":"Figure 5 and Supplementary Figure 9"},{"comment":"The sentence reporting \"Difference in the sNB between the models LogNoSA and LogSingleSA: 32; 16-49\" would be clearer if it stated explicitly that this is the difference in the Asian subgroup and that a positive value favours LogSingleSA.","section":"Results, Diabetes prognostic risk model"}],"recommendation":"major_revision","confidential_remarks":"This is a promising methods paper that connects decision curve analysis to algorithmic fairness in a clinically meaningful way. The main-text derivation is sound conditional on the utility assumptions, and the case studies are analysed with appropriate internal validation. However, the supplementary derivation of lambda as an RRR contains a step that does not preserve subgroup comparisons, and the absence of any sensitivity analysis for lambda undermines the robustness of the quantitative fairness claims. Both issues are fixable within the scope of the manuscript, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Readable and honest. The contribution is Eq. (5): sNB = 1 − π + (λ/N)(TP − (t/(1−t))FP). It is standard decision-curve net benefit with the prevalence offset put back in, and it is a sensible way to compare subgroups on a common utility scale. The derivation is correct, collapsibility follows immediately from linearity, and the two UK Biobank examples are worked carefully with optimism-corrected internal validation and available code. The Pareto-front analysis under screening caps is the most useful part: it makes the equity/efficiency trade-off concrete.\n\nThe main soft spot is the one the stress test flags. Lambda is load-bearing and externally fixed. The RRR from a trial is not obviously the right utility ratio for every subgroup, and if true treatment benefit differs by ethnicity or deprivation, the sNB levels, gap reductions, and the Pareto front can all change. The authors know this; the Discussion says ranges of values can be explored and future work should investigate the consequences. But they never run that exploration, and the text repeatedly states that models “reduced healthcare inequalities” based on that single lambda. This is an acknowledged, untested premise, not a hidden error.\n\nSecond, in the lung cancer example the absolute differences are tiny: a shift of about 5 true negatives per 10,000 at best, and overall benefit differences along the Pareto front below 1 per 10,000. Calling that “reduced health inequalities” is more confident than the numbers support. The diabetes example is more informative, with meaningful movement in the Asian group and a credible difference between including and excluding ethnicity.\n\nMinor: collapsibility is presented as an important property but is just weighted averaging of a linear quantity. The supplementary figure captions also have copy-paste errors, like Supplementary Figure 9 naming XGB models as logistic regression. Those should be cleaned up.\n\nOverall this is a modest but real contribution, strongest as a framing device and as a reporting recommendation aligned with TRIPOD+AI. I would send it to peer review. A serious referee should ask for a sensitivity analysis over lambda, including subgroup-specific values; temper the “reduces inequalities” language; and ask for tighter supplementary text.","headline":"A useful reframing of net benefit for subgroup fairness, but every equity-gap claim rests on one fixed lambda, and the paper never tests it.","tokens_in":20080,"tokens_out":2357,"would_cite":false,"duration_ms":24045,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that algorithmic fairness for clinical prediction should be assessed by subgroup net benefit, a utility-based measure of how a model distributes clinical benefit across protected groups.","keywords":["clinical prediction models","algorithmic fairness","subgroup net benefit","health equity","health inequalities","decision curve analysis","resource constraints","maximin fairness"],"falsifier":"Estimate subgroup-specific $\\lambda$ from randomized trial data or from observational data with treatment assignment for the diabetes prevention and lung cancer screening interventions; if the subgroup-specific relative risk reductions differ from the single published RRRs of 0.58 and 0.20 by more than a few percent, recompute the sNB rankings and gap reductions. If the ordering of models by maximin sNB flips, or if the reported gap-narrowing effects reverse, the paper's fairness conclusions hinge on the single-$\\lambda$ assumption.","tokens_in":19052,"feed_emoji":"⚖️","tokens_out":8107,"duration_ms":67879,"temperature":0.7,"pith_summary":"This paper argues that algorithmic fairness in clinical prediction should be judged by a model's clinical impact in each protected subgroup, not by parity of predictions or error rates. The authors extend net benefit, the standard decision-analytic measure of a model's value, into a subgroup net benefit that adds each group's baseline disease burden and the relative risk reduction of the intervention. Using this measure, they show how a model distributes benefit across ethnic or deprivation groups, how to compare models by whether they narrow or widen the gap between best- and worst-off groups, and why resource limits can force a trade-off between overall benefit and equity. Two worked examples, type 2 diabetes prevention and lung cancer screening, demonstrate that the metric is computable, collapsible across subgroups, and sensitive to modelling choices such as including the protected attribute as a predictor.","feed_headline":"Clinical model fairness reduces to subgroup net benefit","feed_subtitle":"A single formula turns how much each protected group gains from a model into a number that can be optimized.","key_machinery":"The central object is the subgroup net benefit, a re-scaled decision-curve net benefit that reintroduces the model-independent prevalence term $1-\\pi$ and weights the conventional net benefit by $\\lambda = RRR$, the relative risk reduction of the intervention at the decision threshold. It carries the argument by converting benefit from a property of predictions into a property of clinical decisions in each group: a group with $\\text{sNB}=1$ is unburdened by the outcome, while a group with $\\text{sNB}=0$ consists entirely of untreated false negatives. Its collapsibility, meaning the total sNB is the size-weighted mean of subgroup sNBs, is what allows the paper to treat overall benefit and maximin equity as aligned without resource constraints, and to make explicit the Pareto trade-off when capacity is capped.","core_discovery":"The paper's central claim is that the fairness of a clinical prediction model is a question about the distribution of clinical benefit, and that this distribution can be measured by the subgroup net benefit $\\text{sNB}(t^*) = 1 - \\pi + \\frac{\\lambda}{N}\\left(TP(t^*) - \\frac{t^*}{1-t^*}FP(t^*)\\right)$, where $\\pi$ is subgroup prevalence, $TP$ and $FP$ are true and false positives at threshold $t^*$, and $\\lambda$ is the relative risk reduction of the intervention the model allocates. Because the term $1-\\pi$ is retained, a group with high disease burden starts with lower sNB, so treating all groups equally is not enough to identify fairness. The metric is collapsible: the sNB of the whole population is the size-weighted average of subgroup sNBs, so improving the worst-off subgroup (maximin) need not sacrifice overall benefit when resources are unconstrained. Under capacity constraints, however, the paper shows via Pareto fronts that there is a genuine trade-off between overall net benefit and the sNB of the most disadvantaged group, and it proposes using the Pareto curve to make this trade-off explicit for decision-makers.","pith_inferences":["Editorial inference: Because sNB depends on $\\lambda$, the relative risk reduction, the same model can look fair or unfair under different assumptions about treatment effectiveness; reporting sNB as a function of $\\lambda$ rather than at a single value would reveal how sensitive equity conclusions are.","Editorial inference: The collapsibility of sNB suggests a natural fairness diagnostic for any deployed model: decompose the change in overall sNB into subgroup contributions and attribute the equity gap to differences in prevalence versus differences in predictive performance.","Editorial inference: The Pareto-front formulation could be extended from thresholds to other levers such as screening intervals, outreach targeting, or model retraining, so that equity trade-offs are considered across the whole implementation, not just at the decision threshold.","Editorial inference: A testable prediction of the framework is that models trained to maximize overall sNB under no constraints will rarely be the ones that minimize the sNB gap; if this holds across many clinical settings, gap-based maximin criteria would be needed in model selection even when resources are abundant."],"forward_implications":["Including the protected attribute as a predictor in the diabetes model raised the sNB of Asian and Black groups and narrowed the gap to the white group compared with omitting it, so model-building choices can be evaluated by their equity effect.","In the lung cancer example, all three modelling strategies reduced the sNB gap between deprivation quintiles compared to screening no one, and the benefit concentrated in the most-deprived quintile.","A random 'fair' screening policy reduced gaps only by lowering sNB in every group, illustrating that naive randomization harms all subgroups rather than levelling up.","Under tight capacity caps of screening 3% or 1% of the population, overall and most-deprived-quintile sNB lie on a Pareto front, and the loss in overall benefit across the front is tiny compared to the gain in the worst-off group's sNB.","Because sNB is collapsible, without resource constraints one can always achieve both maximum overall benefit and maximin equity by ensembling each subgroup's best model."],"supporting_citations":[{"why":"Supplies the decision-curve net benefit formula that subgroup net benefit extends.","marker":"[23]"},{"why":"Documents the levelling-down critique of strict egalitarian fairness that motivates an equity-based measure.","marker":"[14]"},{"why":"Argues that clinical prediction models can increase disparities and frames the tension between patient-level and system-level benefit.","marker":"[1]"},{"why":"Reporting guidelines that urge developers to consider underserved populations, providing the adoption context for subgroup reporting.","marker":"[6]"},{"why":"Prior use of net benefit for algorithmic fairness in healthcare; the paper extends it to subgroup comparisons and resource constraints.","marker":"[41]"},{"why":"Provides the method for computing net benefit under resource constraints that the paper applies to fairness trade-offs.","marker":"[43]"},{"why":"Supplies the 0.58 relative risk reduction used as lambda in the diabetes example.","marker":"[29]"},{"why":"Supplies the 0.20 relative risk reduction used as lambda in the lung cancer example.","marker":"[33]"},{"why":"Provides the cohort data for both clinical use cases.","marker":"[35]"}],"fun_headline_variants":["Fairness is net benefit per subgroup, not equal performance","Subgroup net benefit: the real fairness metric for clinical AI","Fair models maximize the worst-off group's net benefit","Clinical AI fairness: a trade-off under resource constraints","Collapsible metric: total benefit = weighted sum of subgroup benefits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the relative risk reduction $\\lambda$ of the intervention is known and identical across subgroups; if the true treatment effect varies by ethnicity or deprivation, or if it differs in the model-identified positive patients from the trial population that supplied $\\lambda$, every sNB level, gap reduction, and Pareto trade-off in the paper would change.","fun_headline_variants_meta":{"raw":{"variants":["Fairness is net benefit per subgroup, not equal performance","Subgroup net benefit: the real fairness metric for clinical AI","Fair models maximize the worst-off group's net benefit","Clinical AI fairness: a trade-off under resource constraints","Collapsible metric: total benefit = weighted sum of subgroup benefits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1396,"prompt_tokens":1010,"completion_tokens":386,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":626,"completion_tokens_details":{"reasoning_tokens":304}},"tokens_in":626,"tokens_out":386,"duration_ms":3943,"temperature":1.0,"reasoning_tokens":304,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:28:33.272904+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Estimate subgroup-specific $\\lambda$ from randomized trial data or from observational data with treatment assignment for the diabetes prevention and lung cancer screening interventions; if the subgroup-specific relative risk reductions differ from the single published RRRs of 0.58 and 0.20 by more than a few percent, recompute the sNB rankings and gap reductions. If the ordering of models by maximin sNB flips, or if the reported gap-narrowing effects reverse, the paper's fairness conclusions hinge on the single-$\\lambda$ assumption.","supporting_citations":[{"cited_title":"Second round results from the Manchester ‘Lung Health Check’ community-based targeted lung cancer screening pilot","cited_arxiv_id":null,"evidence_quote":"Supplies the 0.20 relative risk reduction used as lambda in the lung cancer example."},{"cited_title":"Systematic Review of Lung Cancer Screening: Advancements and Strategies for Implementation","cited_arxiv_id":null,"evidence_quote":"Provides the cohort data for both clinical use cases."},{"cited_title":"A glossary for health inequalities","cited_arxiv_id":null,"evidence_quote":"Supplies the decision-curve net benefit formula that subgroup net benefit extends."},{"cited_title":"Predictably unequal: understanding and addressing concerns that algorithmic clinical prediction may increase health disparities","cited_arxiv_id":null,"evidence_quote":"Argues that clinical prediction models can increase disparities and frames the tension between patient-level and system-level benefit."},{"cited_title":"Accessed March 14, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the 0.58 relative risk reduction used as lambda in the diabetes example."}],"review_version":1}