{"id":"24fc5deb-eb92-4884-8f1f-de22285b8e9d","arxiv_id":"2505.04886","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A survey of 83 Prolific workers finds that people evaluating a kidney transplant prediction model prefer separation and sufficiency fairness metrics over independence, and view the model as fair by gender and race but unfair by age.","lead":"This paper proposes three fairness measures for regression models and tests which measure crowdsourced participants prefer when evaluating a kidney transplant prediction tool. The participants favored separation and sufficiency over independence, and judged the tool fair for gender and race but unfair for age.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Recovered social preference weights may be non-identified; the paper's own Initialization 3 simulation shows SAFF yields inconsistent beta* when no notion dominates, yet Table 5 reports beta* as point estimates without uniqueness or robustness checks.","rationale":"The reader's weakest assumption was the Weibull distributional assumption. That concern is legitimate and should be tested, for example by recomputing the KL divergences nonparametrically from the actual UPAT outputs. However, the most load-bearing issue for the paper's headline claim is the identifiability of the preference weights. The paper itself demonstrates in Section 5.1, Initialization 3, that when no fairness notion has a clear majority, SAFF produces inconsistent estimates across runs because the loss has multiple minima. The survey estimates, especially for age and gender, are close to that equal-preference regime, and no uncertainty or robustness analysis is reported. This means the reported beta* values in Table 5 may not be unique, and the conclusion that participants strongly prefer separation and sufficiency is not established. The paper's contribution remains valuable, and the concern can be settled by a multi-start and bootstrap analysis; hence the verdict stays conditional rather than moving to accept or reject.","tokens_in":16095,"tokens_out":10541,"duration_ms":120090,"concrete_test":"On the actual survey responses, rerun SAFF from 100 random beta* initializations and on 100 bootstrap resamples of the 83 participants, recording the final beta* and the achieved loss. If the final beta* vectors cluster around one point and every cluster has beta2+beta3 substantially above beta1, the concern is resolved. If multiple distinct beta* with near-equal loss appear, or clusters disagree on the preferred fairness notion, the reported point estimates are not identified and the headline preference claim should be withdrawn or re-estimated with a regularized or otherwise identifiable model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central 'strong preference' claim rests on the estimated preference vector beta* in Table 5. For this claim to hold, beta* must be identified from the survey loss. The paper's simulation results (Section 5.1) show that under Initialization 3, a population split roughly equally among the three fairness notions, the learned beta* is inconsistent across runs because the loss landscape has multiple minima. The survey estimates are not far from that regime (e.g., age beta*=(0.27,0.39,0.34); gender (0.29,0.31,0.40)), and the paper reports them with no confidence intervals, bootstrap replicates, or a multi-start analysis. If the survey loss has multiple minima with different implied rankings, the reported preference for separation/sufficiency over independence is an artifact of the initial beta*0 and not a stable social preference. The age/gender/race fairness-score pattern in Table 5 could survive, but the 'strong preference' component of the abstract is unsupported. Note also that the Weibull and independence assumptions feeding phi_l are unvalidated, but the identifiability problem is more direct: it attacks the preference estimates even if every phi_l is exactly right.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes three KL-divergence-based group fairness notions for regression-based predictive models (independence, separation, and sufficiency), applies them to a real kidney-transplant decision-support tool (UPAT), and uses a Prolific survey with 85 participants to estimate social preferences over these notions via a mixed-logit model and a projected-gradient learning algorithm (SAFF). The authors report that participants weight separation and sufficiency more heavily than independence, that UPAT is perceived as fair with respect to gender and race, and that it is perceived as unfair with respect to age. They also provide simulation results on convergence and discuss clinical justifications for age-based disparities.","tokens_in":16321,"tokens_out":4156,"duration_ms":46318,"significance":"If the reported findings are reliable, the paper would be a meaningful contribution to fairness in regression and to human-factors research on algorithmic fairness in a high-stakes medical setting. Its strengths include the use of a deployed clinical prediction tool, the extension of classification fairness notions to continuous regression outputs, a substantial crowd-sourced preference-elicitation study, careful modeling of discrete-choice responses through mixed logit, and an explicit limitations section. The claim that the public cares more about decision-conditional fairness notions than about independence in a regression-based organ-placement tool is both novel and actionable for deployment decisions. However, the central preference estimate rests on distributional assumptions and an identification assumption that are not adequately validated, so the headline claim is currently not fully supported.","major_comments":[{"comment":"The closed-form fairness scores rely on two unvalidated assumptions: that the predictions y_T and y_D are mutually independent, and that each marginally follows a Weibull distribution. Section 2 states that TTNO is produced by a log-logistic accelerated failure time model and mortality by a Cox proportional hazards model, and Section 6 concedes that the Weibull assumption may not fit. Because the phi_l values in Table 5 propagate through Eqs. (10)-(11) into every estimated preference weight, any distributional misspecification directly affects the paper's central 'strong preference' claim. A sensitivity analysis using empirical KL estimates or alternative fitted distributions is needed before the numeric preference weights can be trusted.","section":"Section 3.1, Eqs. (2)-(4) and Table 5"},{"comment":"The simulation section explicitly reports that under Initialization 3, where no single fairness notion dominates, SAFF produces inconsistent socially preferred notions across runs because the loss landscape has multiple minima. The survey estimates in Table 5, such as gender (0.29, 0.31, 0.40) and age (0.27, 0.39, 0.34), lie precisely in the near-equal-weight regime where this identification problem occurs, yet they are reported as deterministic point estimates without confidence intervals, bootstrap replicates, or a multi-start analysis. The abstract's claim of a strong preference for separation and sufficiency over independence therefore requires an identification or robustness analysis that the manuscript does not provide.","section":"Section 5.1, Initialization 3 and Table 5"},{"comment":"The sufficiency score phi_3 is defined as a KL divergence between two Bernoulli distributions conditional on a fixed prediction vector y, but the manuscript does not specify how this pointwise divergence is integrated or averaged over the empirical distribution of prediction vectors to produce the single number reported in Table 5. Without this aggregation step, the reported sufficiency scores are not reproducible, and the comparison of phi_3 across gender, race, and age is not well defined.","section":"Section 3.1, Definition 3 and Eq. (8); Table 5"}],"minor_comments":[{"comment":"Section 2.3 states that 83 participants remained after attention-check exclusions, while Section 4 reports N=75 in the survey experiment; this discrepancy should be resolved.","section":"Section 2.3 vs. Section 4"},{"comment":"The notation 'tan^{-1}' appears in several beta-derivative derivations where 't^{a_n-1}' is evidently intended; this is likely a typesetting error and should be corrected for readability.","section":"Appendix B, Eqs. (24)-(27)"},{"comment":"The denominator in Eq. (8) should consistently display the group condition X_{m'}; the current mixed notation makes the compared distributions harder to parse.","section":"Section 3.1, Eq. (8)"},{"comment":"The table reports no uncertainty measures for beta* or s*, which is especially important given the identifiability concern raised above; adding bootstrap or multi-start summaries would substantially improve the manuscript.","section":"Table 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the journal's scope and the empirical study is valuable, but the headline preference result is not yet robust to the two load-bearing issues of distributional misspecification and non-identification. I would be willing to review a revision that adds sensitivity analyses and uncertainty quantification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the headline: this is a genuine first attempt to measure public fairness preferences for a regression-based clinical decision-support tool, and it uses real UPAT outputs and STAR data. That alone makes it worth reading. The abstract's 'strong preference' claim, though, is not supported by the estimates as reported.\n\nWhat's good: The three fairness notions are straightforward extensions of independence/separation/sufficiency to regression via KL divergence, and while not deeply novel they are cleanly defined and useful. The survey is carefully designed with attention checks, demographic comparison, and real clinical data. The mixed-logit/Beta model is a sensible way to map Likert responses to latent preferences, and the SAFF algorithm is a reasonable projected-gradient approach. The simulator results for Initialization 1 and 2 show the method can recover known preferences in the easy cases.\n\nSoft spots: the Weibull assumption for TTNO and mortality predictions is acknowledged but load-bearing; the actual models are log-logistic AFT and Cox PH, so all phi scores in Table 5 could be miscalibrated. The asserted independence of y_T and y_D is not proven and both depend on shared recipient features. More importantly, the stress-test concern lands: under Initialization 3, where the population is split evenly, the learned beta* is inconsistent across runs because the loss has multiple minima. The survey beta* values (e.g., age 0.27/0.39/0.34; gender 0.29/0.31/0.40) are close to that equal-weight regime, yet they are reported as point estimates with no confidence intervals, bootstraps, or multi-start checks. So the 'strong preference' for separation/sufficiency is not yet established; the age/gender/race fairness-score pattern may survive, but the preference ranking could be an artifact of initialization. The demographic skew (younger, more educated, fewer Hispanic participants) is disclosed but further limits generalizability.\n\nBottom line: This paper deserves serious peer review. It asks a new question and brings real data to bear, but the central claim needs robustness analysis. A referee should ask for identifiability checks (multi-start, profile likelihood, bootstrap), validation of the distributional assumptions, and a softer interpretation of the preference weights. I'd recommend conditional accept or major revision.","headline":"A serious first attempt at measuring fairness preferences for a regression-based clinical tool, but the headline preference claim needs identifiability and distributional robustness checks before it can be trusted.","tokens_in":16858,"tokens_out":1940,"would_cite":true,"duration_ms":19860,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A regression-fairness framework based on KL divergence, applied to a kidney allocation tool, finds that the public prefers separation and sufficiency over independence and views age-based disparities as unfair.","keywords":["algorithmic fairness","regression fairness","KL divergence","fairness perceptions","mixed-logit discrete choice","kidney transplantation","predictive analytics","group fairness"],"falsifier":"Take the same 10 donor-recipient data tuples the survey used, recompute the three fairness scores from the actual prediction outputs using nonparametric density estimates instead of the Weibull closed forms, and rerun the preference-learning algorithm; if the age-group divergences drop to the gender/race range (around 0.1) or the recovered social weights stop favoring separation and sufficiency, the reported age-unfairness verdict is an artifact of the distributional assumption.","tokens_in":15861,"feed_emoji":"⚖️","tokens_out":13369,"duration_ms":125527,"temperature":0.7,"pith_summary":"This paper sets out to show that fairness in regression-based predictive models can be quantified with three distribution-comparison criteria—independence, separation, and sufficiency—and that real people's fairness ratings can be decomposed into weights over those criteria. Applied to a kidney-allocation decision-support tool, the measures show large divergences between older and younger candidates on every criterion (independence 4.10, separation 3.66, sufficiency 1.08), while gender and race groups score close to zero on independence and sufficiency. An online survey (83 participants after attention-check exclusions) feeding a mixed-logit choice model yields social preference weights that favor separation and sufficiency over independence, and the estimated social feedback score rates the tool completely fair on gender and race but completely unfair on age. If these findings hold, fairness audits of regression-based clinical tools should report conditional error parity and calibration, not just demographic balance, and transplant organizations face a public-trust problem around age-stratified risk prediction.","feed_headline":"Public deems kidney-transplant AI unfair to older patients","feed_subtitle":"Survey participants favored error-based separation and calibration-based sufficiency, scoring the tool fair for race and gender.","key_machinery":"The core mechanism is the fairness-score triple φ_ℓ(d_m), each a Kullback-Leibler divergence between two social groups' conditional distributions: φ₁ compares prediction distributions P(ŷ|group) (independence), φ₂ compares predictions conditioned on the surgeon's decision P(ŷ|z, group) (separation), and φ₃ compares decision distributions conditioned on predictions P(z|ŷ, group) (sufficiency). Since the tool emits two conditionally independent predictions—time-to-next-offer and mortality likelihood—each divergence splits into a sum over the two components; the paper evaluates the sums with the closed-form KL divergence for Weibull distributions (Eqs. 3–4), and the authors flag in the limitations that the Weibull assumption may not match the underlying log-logistic and Cox models. The scores are min-max normalized per data tuple, and each participant combines them into a latent aggregated score ψ = Σ β φ̄, which is modeled as Beta-distributed and mapped onto the 7-point Likert regions; a mixed-logit softmax turns region utilities into response probabilities. The social preference vector β* is then fitted by projected-gradient minimization of mean-squared feedback regret (the SAFF algorithm). This chain—divergence scores, weighted aggregation, stochastic choice, regret minimization—is what turns subjective fairness ratings into the reported social weights.","core_discovery":"The paper's central claim is that the standard classification fairness families—independence, separation, and sufficiency—carry over to regression when each is expressed as a zero Kullback-Leibler divergence between group-conditioned distributions, and that social preference over these notions can be learned from Likert-scale feedback. On the transplant tool, the recovered social weights are pronounced: sufficiency and separation dominate independence, most clearly for race (β*₃=0.42, β*₂=0.44, β*₁=0.14), indicating that participants judge fairness conditionally on the surgeon's decision and on calibration. The same data classify gender and race as completely fair (social feedback score 7) and age as completely unfair (score 1), because the age-group divergences are roughly an order of magnitude larger. The authors conclude that transparency about age-stratified clinical risk is needed to preserve public trust, and that conditional fairness metrics must be monitored even when overall verdicts are fair.","pith_inferences":["Beyond the paper: since the time-to-next-offer model deliberately excludes demographic attributes while the mortality model includes age, the age-group divergence most likely originates in the mortality model and in surgeons' decisions about older candidates; a follow-up presenting the two predictions separately could isolate which one drives the unfairness rating.","Beyond the paper: recomputing the KL divergences with nonparametric density estimates could move the age scores substantially; if they fall toward the gender/race range, the age-unfairness verdict would be partly an artifact of the Weibull assumption rather than a property of the tool.","Beyond the paper: the participant pool skews younger, more educated, and less Hispanic than the U.S. population, so the learned social weights may not generalize; a preference-elicitation study with transplant patients, older adults, and clinicians could reveal heterogeneous fairness standards.","Beyond the paper: if public preference for decision-conditional fairness holds more broadly, then fairness regulation for clinical AI should mandate subgroup calibration and error-rate parity reporting, not just demographic balance of predictions."],"forward_implications":["Fairness reporting for regression-based clinical decision support should include separation and sufficiency metrics alongside demographic outcome parity, because those are the criteria the surveyed public weights most heavily.","The transplant network should treat the age-group unfairness verdict as a public-trust issue even if the underlying age-stratified risk is clinically justified; the survey response landed at 'completely unfair' for age.","Monitoring efforts should track the separation score φ₂ for gender and race, since it is non-negligible (0.30 and 0.23) even where the overall social verdict is 'completely fair,' and a rise could erode trust.","The same KL-based fairness scoring and preference-learning pipeline can be applied to other regression-based healthcare prediction tools, such as liver allocation or cancer risk models, after checking the distributional assumptions."],"supporting_citations":[{"why":"Describes the deployed predictive analytics tool (UPAT) and the pilot showing offer-acceptance rising from 16.8% to 19.7%; this is the system whose fairness is evaluated.","marker":"McCulloh et al., 2023"},{"why":"Supplies the three classification fairness frameworks—independence, separation, and sufficiency—that the paper generalizes to regression via KL divergence.","marker":"Barocas et al., 2019"},{"why":"Provides the decision-conditional fairness foundations (equalized odds) that motivate the separation notion and conditioning on the surgeon's decision.","marker":"Hardt et al., 2016"},{"why":"Gives the closed-form KL divergence between Weibull distributions that the paper uses to compute the fairness scores in Eqs. (3)-(4).","marker":"Bauckhage and Manshaei, 2014"},{"why":"Demonstrates that multiple fairness criteria cannot generally be satisfied simultaneously, framing the need to measure which notion the public prefers.","marker":"Chouldechova, 2017"},{"why":"Proves inherent trade-offs among fairness criteria, justifying the choice-model approach to identifying a socially preferred criterion.","marker":"Kleinberg et al., 2017"},{"why":"Introduces the conditional/mixed-logit discrete choice framework used to model participants' fairness feedback.","marker":"McFadden et al., 1973"},{"why":"Provides the mixed-logit estimation methodology underpinning the recovery of heterogeneous preference weights.","marker":"Train, 2009"},{"why":"Prior work learning social fairness preferences in kidney placement from non-expert stakeholders; this paper extends that approach from classification to regression.","marker":"Telukunta et al., 2024"}],"fun_headline_variants":["Transplant AI unfair to elderly, fair to race and gender","Crowd votes: kidney AI age-biased, race and gender fair","Public rejects transplant model's age fairness, prefers conditional","Survey: transplant predictions age-unfair despite race, gender OK","Age bias in transplant AI tops social fairness concerns"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All the fairness scores feeding the public-preference estimate are computed by assuming each prediction follows a particular statistical curve (a Weibull distribution), although the tool's own prediction models are built on different curve types; the paper concedes this may not hold, and if it does not, the scores, weights, and the age-unfairness verdict all shift.","fun_headline_variants_meta":{"raw":{"variants":["Transplant AI unfair to elderly, fair to race and gender","Crowd votes: kidney AI age-biased, race and gender fair","Public rejects transplant model's age fairness, prefers conditional","Survey: transplant predictions age-unfair despite race, gender OK","Age bias in transplant AI tops social fairness concerns"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1224,"prompt_tokens":909,"completion_tokens":315,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":231}},"tokens_in":525,"tokens_out":315,"duration_ms":4081,"temperature":1.0,"reasoning_tokens":231,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:18:27.827277+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same 10 donor-recipient data tuples the survey used, recompute the three fairness scores from the actual prediction outputs using nonparametric density estimates instead of the Weibull closed forms, and rerun the preference-learning algorithm; if the age-group divergences drop to the gender/race range (around 0.1) or the recovered social weights stop favoring separation and sufficiency, the reported age-unfairness verdict is an artifact of the distributional assumption.","supporting_citations":[{"cited_title":"An Experiment on the Impact of Predictive Analytics on Kidney Offer Acceptance Decisions","cited_arxiv_id":null,"evidence_quote":"Describes the deployed predictive analytics tool (UPAT) and the pilot showing offer-acceptance rising from 16.8% to 19.7%; this is the system whose fairness is evaluated."},{"cited_title":"Fairness and machine learning","cited_arxiv_id":null,"evidence_quote":"Supplies the three classification fairness frameworks—independence, separation, and sufficiency—that the paper generalizes to regression via KL divergence."},{"cited_title":"Kernel archetypal analysis for clustering web search frequency time series","cited_arxiv_id":null,"evidence_quote":"Gives the closed-form KL divergence between Weibull distributions that the paper uses to compute the fairness scores in Eqs. (3)-(4)."},{"cited_title":"Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments","cited_arxiv_id":null,"evidence_quote":"Demonstrates that multiple fairness criteria cannot generally be satisfied simultaneously, framing the need to measure which notion the public prefers."},{"cited_title":"Inherent Trade-Offs in the Fair Determination of Risk Scores","cited_arxiv_id":null,"evidence_quote":"Proves inherent trade-offs among fairness criteria, justifying the choice-model approach to identifying a socially preferred criterion."},{"cited_title":"Conditional Logit Analysis of Qualitative Choice Behavior","cited_arxiv_id":null,"evidence_quote":"Introduces the conditional/mixed-logit discrete choice framework used to model participants' fairness feedback."},{"cited_title":"Learning Social Fairness Preferences from Non-Expert Stakeholder Opinions in Kidney Placement","cited_arxiv_id":null,"evidence_quote":"Prior work learning social fairness preferences in kidney placement from non-expert stakeholders; this paper extends that approach from classification to regression."}],"review_version":1}