Pith. sign in

REVIEW 2 major objections 4 minor 96 references

Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments

T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Delayed race reporting can materially distort timely healthcare disparity assessments.

desk verdict A well-executed, novel retrospective analysis showing delayed race reporting materially distorts timely disparity estimates, with one unresolved construct-validity concern about registry ingestion lag. read the letter →

arxiv 2506.13735 v1 pith:PBC7DQDM submitted 2025-06-16 cs.CY

classification cs.CY
keywords healthdisparitiesreportingdelayraceandethnicitydatamissingdisparityassessmentselectronicrecordsalgorithmicfairnessimputation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that delayed demographic reporting—race and ethnicity information that is missing when an assessment is run but appears later—can distort healthcare disparity assessments even though the data eventually arrives. Using timestamped electronic health records for over 5 million patients across more than 1,000 U.S. primary care practices, the authors show that over 73% of patients experience some delay in race reporting, and that delays are unevenly distributed across racial groups and correlate with health outcomes. They simulate a disparity assessment for a fixed 2018 Q1 cohort at the earliest possible moment, when only 59.06% of patients have race recorded, and again once race is complete, holding health outcomes fixed; national average prevalence error is 2.15 percentage points, and 13.39% of practice-level prevalence estimates are wrong by more than 10 percentage points. They also find that a widely used name-and-geography imputation method, BIFSG, reduces average prevalence error but does not reliably reduce disparity error. If the paper is right, routine equity dashboards and audit reports built on incomplete race data can mislead decisions about where and how to intervene.

What carries the argument

The load-bearing mechanism is the timestamped longitudinal race report: every update to a patient's demographic record is logged with a date, so the authors can observe, for each of 5,310,700 patients, the gap between the earliest possible reporting date (the later of the first date-of-birth timestamp and the first practice-level race report) and the date race actually becomes available. This operationalization converts delay from an abstract concern into a continuous, measurable variable. The retrospective analysis then computes prevalence and disparity estimates for a fixed 2018 Q1 cohort at $t_{\text{initial}}$ and $t_{\text{final}}$, with error metrics defined as prevalence error, disparity error, and sign-switch classification, and uncertainty quantified by bootstrapping over practices and patients. The imputation test uses Bayesian Improved First Name Surname Geocoding (BIFSG), a name-and-geography method that assigns each patient a posterior probability of belonging to each racial group, which the paper then treats as weights in prevalence estimates.

What would settle it

Compare registry timestamps against a source EHR's own audit log, or study a practice that switches from pull-only to push-based submission: if the delays and their association with health outcomes largely disappear without any change in patient intake behavior, the delay construct is dominated by ingestion cadence rather than point-of-care missingness.

Watch

Extended reading notes

Core claim

The central claim is that reporting delay is a distinct form of missingness—one that is temporary but can still bias disparity assessments when monitoring is time-sensitive. On the paper's own terms, delayed race reporting is not a static data gap: it is a temporal process the authors measure day-by-day from the earliest possible reporting date to the date race appears in the registry. Retrospective comparisons at $t_{\text{initial}}$ versus $t_{\text{final}}$, with health outcomes held fixed for 2018 Q1, attribute all observed changes in prevalence and pairwise disparity estimates to the delayed arrival of race information. The authors find that these changes are consequential: national average prevalence error is 2.15 percentage points across groups and outcomes, state-level errors vary widely, over 10-percentage-point errors occur in 13.39% of practice-level estimates, and pairwise disparity sign switches occur in up to 14% of comparisons for electrocardiogram procedures. The paper further claims that BIFSG imputation, despite decent individual-level accuracy, does not consistently correct disparity errors, so it cannot replace timely self-reported race data.

Load-bearing premise

The load-bearing premise is that the measured 'delay' reflects real missingness in patient data rather than the registry's data-ingestion schedule; if most of the gap is just how often practices push or the registry pulls data, the conclusion would not transfer to settings that submit race data promptly.

Editorial extensions

If this is right

  • Routine quarterly disparity assessments in this data would exclude most patients: over half of the cohort has race delays of 60 days or more, so early assessments are built on a minority of patients.
  • Assessment granularity matters: national-level distortions are modest on average, but state-level and practice-level prevalence errors are much larger, with 13.39% of practice-level estimates off by 10 percentage points or more.
  • Delays can reverse the apparent direction of a disparity: up to 14% of pairwise comparisons for electrocardiograms switch signs between $t_{\text{initial}}$ and $t_{\text{final}}$, meaning an early audit can name the wrong group as worse off.
  • Imputation is not a reliable fix: BIFSG lowers average prevalence error for all outcomes but significantly improves average disparity error for only three of six outcomes, and it over-estimates prevalence for several minority groups.
  • Because delays correlate with practice-level data-submission modes (pull-only versus push) and specific EHR systems, improving data pipeline design may be as important as improving statistical methods.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Going beyond the paper: if registry ingestion cadence is a major driver of the measured gaps, then switching practices from pull-only to push-based submission could shrink delay-induced disparity distortion without any change in patient self-reporting behavior; this is a testable intervention.
  • Going beyond the paper: the same delay lens applies to other attributes usually treated as static—gender, insurance type, marital status, pre-existing diagnoses—and to fairness audits outside healthcare where administrative data pipelines lag behind real-world events.
  • Going beyond the paper: delay-induced missingness is a form of missing-not-at-random that changes which patients are visible at each assessment, so continuous audits should report data-maturity thresholds and re-run conclusions once race completion passes a prespecified level.
  • Going beyond the paper: because the visible cohort changes over time, ML fairness metrics beyond prevalence—such as calibration or equalized odds—could also shift with delay; the paper's framework invites that extension even though it does not compute those metrics.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper studies temporal missingness of race and ethnicity information in a large U.S. primary-care registry (AFC). It defines a reporting delay as the gap between a patient's earliest reporting opportunity and the date when race information is recorded in the registry, documents that over 73% of patients experience such delays, and shows that delay rates vary by race, health status, and practice infrastructure. The core retrospective analysis fixes health outcomes for 2018 Q1 and varies only the availability of race information, measuring how prevalence and disparity estimates at an early assessment time differ from the same estimates once race reporting is complete. The authors report national-level average prevalence errors of 2.15 percentage points, practice-level prevalence errors exceeding 10 points for 13.39% of estimates, and sign switches in pairwise disparities in up to 14% of comparisons. They also show that BIFSG imputation improves average prevalence error but does not consistently correct disparity errors. The paper concludes that delayed demographic reporting can materially distort timely disparity assessments and argues for pipeline-aware fairness approaches.

Significance. If the delay construct is valid, this is a novel and practically important contribution. The retrospective design is clean: outcomes are fixed for 2018 Q1, and only race availability changes across simulated assessment times. The dataset is unusually rich—5.3M adult patients, 1,000+ practices, all 50 states, and timestamped race updates—and the authors make code available. The paper is also transparent about the provenance of its timestamps and about its exclusion of patients who never report race. The findings speak directly to state-level health equity reporting mandates and to the broader algorithmic fairness literature on missing and delayed sensitive attributes. The main weakness is that the measured delay is anchored to registry ingestion timestamps, not point-of-care reporting, which limits the generalizability of the headline quantitative claims until that construct is interrogated further.

major comments (2)
  1. [Section 4.1; Table 5] The load-bearing assumption is the construct validity of the delay variable. The paper acknowledges in Section 4.1 that "our date information reflects the date when this information was shared or updated with the data provider responsible for producing the AFC data, not the clinical encounter when the patient may have self-reported race." Table 5 shows that delayed patients are substantially more likely to come from "pull only" practices (82.76% vs. 76.39%) and that delay rates vary sharply by EHR system. The headline numbers—73% of patients delayed, 2.15 percentage points average national prevalence error, 13.39% of practice-level estimates off by more than 10 points—therefore conflate point-of-care reporting delay with registry data-ingestion cadence. Because the abstract and conclusion generalize to reporting delays in healthcare and beyond, this is not a minor caveat. I recommend a sensitivity analysis that stratifies the main retrospective estimates by data-update mode (push vs. pull) and EHR system, or that re-estimates the headline errors using only push-only practices. If the errors persist under prompt transmission, the generalization is supported; if they shrink, the conclusions should be explicitly scoped to registry-based retrospective assessments rather than to point-of-care reporting behavior.
  2. [Section 4.2.2; Section 4.3; Table 3] The definition of the 2018 Q1 cohort is not precise enough to support the prevalence claims. The text states that the cohort of 1,776,729 patients is obtained by restricting to patients whose date of birth is reported before 2018 and practices that report race before 2018, but it does not state whether patients are also required to have at least one visit or record in 2018 Q1. If inactive patients are included, the prevalence denominators include patients with no opportunity to have the outcome in the quarter, and the reported values in Table 3 are not period prevalences. If an encounter restriction is applied, it should be stated explicitly in Section 4.2.2. This is important because all of the paper's central prevalence and disparity error metrics depend on the denominator.
minor comments (4)
  1. [Section 5] The sentence "Kruskal-Wallis tests for all pairwise comparisons are statistically significant" is imprecise: the Kruskal-Wallis test is a global test across groups, not a pairwise test. The authors presumably performed pairwise Mann-Whitney U tests with Benjamini-Hochberg correction and should say so.
  2. [Section 6] There is a typo in "vassopressor administration" in the future-work paragraph; it should be "vasopressor administration."
  3. [Section 4.4] The bootstrap uses only 50 resamples, but the main tables and figures do not report confidence intervals for the headline metrics. Given that some subgroup estimates have wide uncertainty (visible in Figures 3 and 4), reporting bootstrap confidence intervals for the average prevalence error and the practice-level 13.39% figure would strengthen the paper.
  4. [Section 4.2.1] The paper notes that detection rates for multiracial patients are likely lower because only simple regex parsing is applied. Since multiracial patients are included in the analyses, a quantitative assessment of the parsing error rate, or at least a sensitivity check excluding multiracial patients, would be useful.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's headline estimates are computed directly from observed timestamps and outcomes, with no fitted parameter or self-citation chain doing the work.

full rationale

The paper's central quantities—delay rates, prevalence errors, disparity errors, and sign-switch frequencies—are direct empirical comparisons of observed data at different time points. Section 4.1 defines delay from timestamps; Section 4.4 defines prevalence error as prevalence(j, t_final) − prevalence(j, t_initial); Section 4.3 holds health outcomes fixed for a 2018 Q1 cohort and only changes which patients have race available. The headline numbers (over 73% delayed patients, 13.39% practice-level prevalence errors over 10 percentage points, up to 14% sign switches) are measurements, not outputs of a model fitted to reproduce them. There is no fitted parameter that is later renamed a prediction, no uniqueness theorem imported from prior work, and no ansatz smuggled in via citation. The only author-overlapping citation is Cheng et al. [22] for the race/ethnicity text-parsing pipeline; that is auxiliary preprocessing used to map free text to OMB categories, and the delay-distortion conclusion does not reduce to that pipeline. The paper's own stated limitation—that timestamps reflect when data was shared with the AFC data provider rather than the clinical encounter—is a construct-validity and generalizability concern about the delay measure, not a circularity in the derivation. Under the hard rules, construct validity concerns belong in correctness risk, not circularity. Accordingly, the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The paper is empirical and contains no fitted model parameters. Its central claims rest on domain assumptions about the validity of registry timestamps as a measure of reporting delay, the completeness of coded health outcomes, and the treatment of never-reporting patients, all acknowledged in the text. No new physical or conceptual entities are postulated beyond the operational definition of 'delay' which is directly measured from timestamps.

assumptions (6)
  • domain assumption Race and ethnicity can be mapped to 1997 OMB categories using the parsing pipeline of Cheng et al. applied to free-text and categorical fields.
    Section 4.2.1 states the mapping follows Cheng et al.; detection rates for multiracial patients are acknowledged to be lower.
  • domain assumption The first timestamp at which race appears in the registry reflects the time of race reporting for the purpose of defining delay.
    Section 4.1 explicitly says timestamps reflect when data was shared with the provider, not the clinical encounter; all delay measures depend on this.
  • domain assumption Clinical codes for outcomes are accurate (a code entered in the system is true).
    Appendix B: 'We assume that a code, if entered in the system, is accurate.'
  • domain assumption Patients who never report mappable race can be excluded without invalidating the comparison; the t_final benchmark is the eventual race for included patients.
    Section 4.2.2 excludes ~900k never-reporting patients; Section 6 acknowledges this may bias disparity estimates.
  • domain assumption Health outcomes in 2018 Q1 are fully observed for included patients despite outcome reporting also being subject to delay.
    Section 4.2.3 treats attributes as fixed if they ever appear and does not consider outcome delays.
  • standard math Bootstrapping with 50 resamples provides valid uncertainty estimates.
    Section 4.4 describes bootstrap uncertainty; a Monte Carlo assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments." pith.science (2026). https://pith.science/paper/PBC7DQDM

@misc{pith2026250613735,
  author       = {Pith},
  title        = {Pith review of: Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PBC7DQDM}},
  note         = {Machine review of arXiv:2506.13735}
}
read the original abstract

Conducting disparity assessments at regular time intervals is critical for surfacing potential biases in decision-making and improving outcomes across demographic groups. Because disparity assessments fundamentally depend on the availability of demographic information, their efficacy is limited by the availability and consistency of available demographic identifiers. While prior work has considered the impact of missing data on fairness, little attention has been paid to the role of delayed demographic data. Delayed data, while eventually observed, might be missing at the critical point of monitoring and action -- and delays may be unequally distributed across groups in ways that distort disparity assessments. We characterize such impacts in healthcare, using electronic health records of over 5M patients across primary care practices in all 50 states. Our contributions are threefold. First, we document the high rate of race and ethnicity reporting delays in a healthcare setting and demonstrate widespread variation in rates at which demographics are reported across different groups. Second, through a set of retrospective analyses using real data, we find that such delays impact disparity assessments and hence conclusions made across a range of consequential healthcare outcomes, particularly at more granular levels of state-level and practice-level assessments. Third, we find limited ability of conventional methods that impute missing race in mitigating the effects of reporting delays on the accuracy of timely disparity assessments. Our insights and methods generalize to many domains of algorithmic fairness where delays in the availability of sensitive information may confound audits, thus deserving closer attention within a pipeline-aware machine learning framework.

Figures

Figures reproduced from arXiv: 2506.13735 by the authors.

Figure 1
Figure 1. On the left, we show the conventional approach to disparity assessments, which results in a static measure of disparity. On the right, we present the three core components of our analysis. First, we leverage access to a comprehensive dataset of over 1,000 primary care practices, 5M patients from all 50 states, and 100M patient interactions from 2010 to 2024. Second, we use timestamped records to identify and measure… view at source ↗
Figure 2
Figure 2. Reporting rates differ by race and ethnicity. On the [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Simulations at the national level for one condition (hypertension), one procedure (electrocardiograms), and one [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Simulations at the state level for hypertension diagnoses, electrocardiogram procedures, and HbA1c tests. Changes in [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: BIFSG over-estimates prevalence for several minority race groups ( [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Cohort size vs average delay of race reporting for [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Simulations at the national level for two conditions (diabetes and depression) and one procedure (depression screens). [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: Simulations at the state level for two conditions (diabetes and depression) and one procedure (depression screens). [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: Simulations at the national level for a fixed cohort of patients from 2017 Q1. 0 1 2 3 16.0 18.0 20.0 22.0 24.0 26.0 Prevalence condition : hypertension 0 1 2 3 Years since initial assessment 18.0 19.0 20.0 21.0 22.0 procedure : electrocardiogram 0 1 2 3 12.0 13.0 14.0…
Figure 10
Figure 10. Figure 10: Simulations at the national level for a fixed cohort of patients from 2019 Q1. Sign changes (i.e., the direction of the disparity changes, which can be either exacerbations or minimizations) are a problem in either setting as they would distort any meaningful takeaway…
Figure 13
Figure 13. Figure 13: BIFSG mitigates error in disparity assessment (dis [PITH_FULL_IMAGE:figures/full_fig_p018_13.png]
Figure 12
Figure 12. Figure 12: BIFSG reduces the average prevalence error for [PITH_FULL_IMAGE:figures/full_fig_p018_12.png]
Figure 14
Figure 14. Figure 14: BIFSG over-estimates prevalence for several minority race groups ( [PITH_FULL_IMAGE:figures/full_fig_p019_14.png]
Figure 15
Figure 15. Figure 15: Comparison of prevalence estimates for different health outcomes at quarterly intervals starting in 2018. This [PITH_FULL_IMAGE:figures/full_fig_p019_15.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

96 extracted references · 77 canonical work pages

  1. [1]

    Revisions to OMB’s Statistical Policy Directive No

    2021. Revisions to OMB’s Statistical Policy Directive No. 15: Standards for Maintaining, Collecting, and Presenting Federal Data on Race and Ethnicity. https://www.federalregister.gov/documents/2024/03/29/2024- 06469/revisions-to-ombs-statistical-policy-directive-no-15-standards-for- maintaining-collecting-and

  2. [2]

    Hammaad Adam, Fan Yin, Huibin Hu, Neil Tenenholtz, Lorin Crawford, Lester Mackey, and Allison Koenecke. 2023. Should I stop or should I go: early stopping with heterogeneous populations.Advances in Neural Information Processing Systems36 (2023), 15799–15832

  3. [3]

    Nil-Jana Akpinar, Zachary Lipton, and Alexandra Chouldechova. 2024. The Impact of Differential Feature Under-reporting on Algorithmic Fairness. InThe 2024 ACM Conference on Fairness, Accountability, and Transparency. 1355–1382

  4. [4]

    Nil-Jana Akpinar, Manish Nagireddy, Logan Stapleton, Hao-Fei Cheng, Haiyi Zhu, Steven Wu, and Hoda Heidari. 2022. A sandbox tool to bias (stress)-test fairness algorithms.arXiv preprint arXiv:2204.10233(2022)

  5. [5]

    Kelli D Allen, Eugene Z Oddone, Cynthia J Coffman, Francis J Keefe, Jennifer H Lindquist, and Hayden B Bosworth. 2010. Racial differences in osteoarthritis pain and function: potential explanatory factors.Osteoarthritis and cartilage18, 2 (2010), 160–167

  6. [6]

    Karen O Anderson, Carmen R Green, and Richard Payne. 2009. Racial and ethnic disparities in pain: causes and consequences of unequal care.The Journal of Pain 10, 12 (2009), 1187–1204

  7. [7]

    McKane Andrus, Elena Spitzer, Jeffrey Brown, and Alice Xiang. 2021. What we can’t measure, we can’t understand: Challenges to demographic data procurement in the pursuit of fairness. InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. 249–260

  8. [8]

    McKane Andrus and Sarah Villeneuve. 2022. Demographic-reliant algorithmic fairness: Characterizing the risks of demographic data collection in the pursuit of fairness. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 1709–1721

Show all 96 references
  1. [9]

    MJ Arnett, Roland J Thorpe, DJ Gaskin, Janice V Bowie, and Thomas A LaVeist

  2. [10]

    Carolyn Ashurst and Adrian Weller. 2023. Fairness without demographic data: A survey of approaches. InProceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. 1–12

  3. [11]

    American Medical Association. [n. d.]. How to select a practice management system. https://www.ama-assn.org/practice-management/claims-processing/ how-select-practice-management-system

  4. [12]

    Pranjal Awasthi, Alex Beutel, Matthäus Kleindessner, Jamie Morgenstern, and Xuezhi Wang. 2021. Evaluating fairness of machine learning models under uncertain and incomplete information. InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. 206–214

  5. [13]

    Brett Beaulieu-Jones, Samuel G Finlayson, Corey Chivers, Irene Chen, Matthew McDermott, Jaz Kandola, Adrian V Dalca, Andrew Beam, Madalina Fiterau, and Tristan Naumann. 2019. Trends and focus of machine learning applications for health research.JAMA network open2, 10 (2019), e...

  6. [14]

    Asia J Biega, Peter Potash, Hal Daumé, Fernando Diaz, and Michèle Finck. 2020. Operationalizing the legal principle of data minimization for personalization. InProceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval. 399–408

  7. [15]

    Arlene S Bierman, Nicole Lurie, Karen Scott Collins, and John M Eisenberg. 2002. Addressing racial and ethnic barriers to effective health care: the need for better data.Health Affairs21, 3 (2002), 91–102

  8. [16]

    Emily Black, Rakshit Naidu, Rayid Ghani, Kit Rodolfa, Daniel Ho, and Hoda Heidari. 2023. Toward Operationalizing Pipeline-aware ML Fairness: A Research Agenda for Developing Practical Guidelines and Tools. InProceedings of the 3rd ACM Conference on Equity and Access in Algorit...

  9. [17]

    D Keith Branham, Kenneth Finegold, Lucy Chen, Melony Sorbero, Roald Euller, Marc N Elliott, and Benjamin D Sommers. 2022. Trends in missing race and eth- nicity information after imputation in HealthCare. gov marketplace enrollment data, 2015-2021.JAMA Network Open5, 6 (2022),...

  10. [18]

    Karla L Caballero Barajas and Ram Akella. 2015. Dynamically modeling patient’s health state from electronic medical records: A time series approach. InProceed- ings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 69–78

  11. [19]

    César Caraballo, Chima D Ndumele, Brita Roy, Yuan Lu, Carley Riley, Jeph Herrin, and Harlan M Krumholz. 2022. Trends in racial and ethnic disparities in barriers to timely medical care among adults in the US, 1999 to 2018. InJAMA Health Forum, Vol. 3. American Medical Associat...

  12. [20]

    2008.Electronic health records: a guide for clinicians and admin- istrators

    Jerome H Carter. 2008.Electronic health records: a guide for clinicians and admin- istrators. ACP Press

  13. [21]

    Jiahao Chen, Nathan Kallus, Xiaojie Mao, Geoffry Svacha, and Madeleine Udell

  14. [22]

    Lingwei Cheng, Isabel O Gallegos, Derek Ouyang, Jacob Goldin, and Dan Ho

  15. [23]

    Sergey Chernenko and David S Scharfstein. 2023. The Limits of Algorithmic Measures of Race in Studies of Outcome Disparities.A vailable at SSRN 4426161 (2023)

  16. [24]

    Isabel Chien, Nina Deliu, Richard Turner, Adrian Weller, Sofia Villar, and Niki Kilbertus. 2022. Multi-disciplinary fairness considerations in machine learn- ing for clinical trials. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 906–924

  17. [25]

    CMS. 2018. MEASURES MANAGEMENT SYSTEM. https://www.cms. gov/Medicare/Quality-Initiatives-Patient-Assessment-instruments/MMS/ Downloads/A-Brief-Overview-of-Qualified-Clinical-Data-Registries.pdf

  18. [26]

    CMS. 2025. CMS Framework for Health Equity. https://www.cms.gov/priorities/ health-equity/minority-health/equity-programs/framework

  19. [27]

    National Research Council et al. 2004. Eliminating health disparities: Measure- ment and data needs. (2004)

  20. [28]

    Alexander D’Amour, Hansa Srinivasan, James Atwood, Pallavi Baljekar, David Sculley, and Yoni Halpern. 2020. Fairness is not static: deeper understanding of FAccT ’25, June 23–26, 2025, Athens, Greece Gosciak et al. long term fairness via simulation studies. InProceedings of th...

  21. [29]

    Jacob W Dembosky, Amelia M Haviland, Ann Haas, Katrin Hambarsoomian, Robert Weech-Maldonado, Shondelle M Wilson-Frederick, Sarah Gaillot, and Marc N Elliott. 2019. Indirect estimation of race/ethnicity for survey respondents who do not report race/ethnicity.Medical care57, 5 (...

  22. [30]

    Stephen F Derose, Richard Contreras, Karen J Coleman, Corinna Koebnick, and Steven J Jacobsen. 2013. Race and ethnicity data quality and imputation using US Census data in an integrated health system: the Kaiser Permanente Southern California experience.Medical Care Research a...

  23. [31]

    Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. 2012. Fairness through awareness. InProceedings of the 3rd Innovations in Theoretical Computer Science Conference. 214–226

  24. [32]

    Laura Dwyer-Lindgren, Parkes Kendrick, Yekaterina O Kelly, Dillon O Sylte, Chris Schmidt, Brigette F Blacker, Farah Daoud, Amal A Abdi, Mathew Baumann, Farah Mouhanna, et al. 2022. Life expectancy by county, race, and ethnicity in the USA, 2000–19: a systematic analysis of hea...

  25. [33]

    Marc N Elliott, Peter A Morrison, Allen Fremont, Daniel F McCaffrey, Philip Pantoja, and Nicole Lurie. 2009. Using the Census Bureau’s surname list to improve estimates of race/ethnicity and associated disparities.Health Services and Outcomes Research Methodology9 (2009), 69–83

  26. [34]

    Hadi Elzayn, Evelyn Smith, Thomas Hertz, Cameron Guage, Arun Ramesh, Robin Fisher, Daniel E Ho, and Jacob Goldin. 2024. Measuring and mitigating racial disparities in tax audits.The Quarterly Journal of Economics(2024), qjae027

  27. [35]

    Andrea Esuli, Alessandro Fabris, Alejandro Moreo, and Fabrizio Sebastiani. 2023. Learning to quantify. Springer Nature

  28. [36]

    Martínez-Plumed Fernando, Ferri Cèsar, Nieves David, and Hernández-Orallo José. 2021. Missing the missing values: The ugly duckling of fairness in machine learning.International Journal of Intelligent Systems36, 7 (2021), 3217–3258

  29. [37]

    Agency for Healthcare Research and Quality. 2024. 2023 National Healthcare Quality and Disparities Report. https://www.ahrq.gov/research/findings/nhqrdr/ nhqdr23/index.html

  30. [38]

    Allen Fremont, Joel S Weissman, Emily Hoch, and Marc N Elliott. 2016. When race/ethnicity data are lacking: using advanced indirect estimation methods to measure disparities.RAND Health Quarterly6, 1 (2016)

  31. [39]

    Yingqiang Ge, Shuchang Liu, Ruoyuan Gao, Yikun Xian, Yunqi Li, Xiangyu Zhao, Changhua Pei, Fei Sun, Junfeng Ge, Wenwu Ou, et al . 2021. Towards long- term fairness in recommendation. InProceedings of the 14th ACM International Conference on Web Search and Data Mining. 445–453

  32. [40]

    Marzyeh Ghassemi, Marco Pimentel, Tristan Naumann, Thomas Brennan, David Clifton, Peter Szolovits, and Mengling Feng. 2015. A multivariate timeseries modeling approach to severity of illness assessment and forecasting in ICU with sparse, heterogeneous clinical data. InProceedi...

  33. [41]

    Carmen R Green, Karen O Anderson, Tamara A Baker, Lisa C Campbell, Sheila Decker, Roger B Fillingim, Donna A Kaloukalani, Kathyrn E Lasch, Cynthia Myers, Raymond C Tait, et al. 2003. The unequal burden of pain: confronting racial and ethnic disparities in pain.Pain Medicine4, ...

  34. [42]

    Nikolitsa Grigoropoulou and Mario L Small. 2022. The data revolution in social science needs qualitative research.Nature Human Behaviour6, 7 (2022), 904–906

  35. [43]

    Hyeouk Chris Hahm, Benjamin Le Cook, Andrea Ault-Brutus, and Margarita Alegría. 2015. Intersection of race-ethnicity and gender in depression care: screening, access, and minimally adequate treatment.Psychiatric Services66, 3 (2015), 258–264

  36. [44]

    Fern R Hauck, Kawai O Tanabe, and Rachel Y Moon. 2011. Racial and ethnic disparities in infant mortality. InSeminars in perinatology, Vol. 35. Elsevier, 209– 220

  37. [45]

    J Sonya Haw, Megha Shah, Sara Turbow, Michelle Egeolu, and Guillermo Umpier- rez. 2021. Diabetes complications in racial and ethnic minority populations in the USA.Current Diabetes Reports21 (2021), 1–8

  38. [46]

    Margaret M Heckler. 1985. Report of the Secretary’s Task Force Report on Black and Minority Health: The Heckler Report

  39. [47]

    Peter Hepburn, Renee Louis, and Matthew Desmond. 2020. Racial and gender disparities among evicted Americans.Sociological Science7 (2020), 649–662

  40. [48]

    Latoya Hill, Nambi Ndugga, Samantha Artiga, and Anthony Damico. 2025. Health Coverage by Race and Ethnicity, 2010-2023. https://www.kff.org/racial-equity- and-health-policy/issue-brief/health-coverage-by-race-and-ethnicity/

  41. [49]

    Daniel E Ho and Alice Xiang. 2020. Affirmative algorithms: The legal grounds for fairness as awareness.U. Chi. L. Rev. Online(2020), 134

  42. [50]

    Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudik, and Hanna Wallach. 2019. Improving fairness in machine learning systems: What do industry practitioners need?. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–16

  43. [51]

    Rick Hong, Brigitte M Baumann, and Edwin D Boudreaux. 2007. The emergency department for routine healthcare: race/ethnicity, socioeconomic status, and perceptual factors.The Journal of Emergency Medicine32, 2 (2007), 149–158

  44. [52]

    Elizabeth A Howell. 2018. Reducing disparities in severe maternal morbidity and mortality.Clinical Obstetrics and Gynecology61, 2 (2018), 387–399

  45. [53]

    Kosuke Imai and Kabir Khanna. 2016. Improving ecological inference by predict- ing individual ethnicity from voter registration records.Political Analysis24, 2 (2016), 263–272

  46. [54]

    2003.Unequal Treatment: Confronting Racial and Ethnic Disparities in Health Care

    Institute of Medicine (US) Committee on Understanding and Eliminating Racial and Ethnic Disparities in Health Care. 2003.Unequal Treatment: Confronting Racial and Ethnic Disparities in Health Care. National Academies Press (US), Washington (DC). http://www.ncbi.nlm.nih.gov/boo...

  47. [55]

    Cara V James, Jennifer M Haley, Eva H Allen, and Taylor Nelson. 2023. Using Race and Ethnicity Data to Advance Health Equity. https://www.urban.org/ research/publication/using-race-and-ethnicity-data-advance-health-equity

  48. [56]

    Vincent Jeanselme, Maria De-Arteaga, Zhe Zhang, Jessica Barrett, and Brian Tom. 2022. Imputation strategies under clinical presence: Impact on algorithmic fairness. InMachine Learning for Health. PMLR, 12–34

  49. [57]

    Falaah Arif Khan, Denys Herasymuk, Nazar Protsiv, and Julia Stoyanovich. 2024. Still More Shades of Null: A Benchmark for Responsible Missing Value Imputation. arXiv preprint arXiv:2409.07510(2024)

  50. [58]

    Jennifer King, Daniel Ho, Arushi Gupta, Victor Wu, and Helen Webley-Brown

  51. [59]

    Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. 2021. Wilds: A benchmark of in-the-wild distribution shifts. InInternational Conference on Machine Learni...

  52. [60]

    Michael Anne Kyle and Austin B Frakt. 2021. Patient administrative burden in the US health care system.Health Services Research56, 5 (2021), 755–765

  53. [61]

    Khoa Lam, Benjamin Lange, Borhane Blili-Hamelin, Jovana Davidovic, Shea Brown, and Ali Hasan. 2024. A framework for assurance audits of algorithmic systems. InThe 2024 ACM Conference on Fairness, Accountability, and Trans- parency. 1078–1092

  54. [62]

    InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency

    The Privacy-Bias Tradeoff: Data Minimization and Racial Disparity Assess- ments in US Government. InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 492–505

  55. [63]

    Dingwen Li, Patrick Lyons, Jeff Klaus, Brian Gage, Marin Kollef, and Chenyang Lu. 2021. Integrating static and time-series data in deep recurrent models for oncology early warning systems. InProceedings of the 30th ACM International Conference on Information & Knowledge Manage...

  56. [64]

    Roderick JA Little and Donald B Rubin. 1989. The analysis of social science data with missing values.Sociological Methods & Research18, 2-3 (1989), 292–326

  57. [65]

    2019.Statistical analysis with missing data

    Roderick JA Little and Donald B Rubin. 2019.Statistical analysis with missing data. Vol. 793. John Wiley & Sons

  58. [66]

    Michelle Seng Ah Lee and Jat Singh. 2021. The landscape and gaps in open source fairness toolkits. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–13

  59. [67]

    David Machledt. 2021. Addressing health equity in Medicaid managed care. National Health Law Program(2021)

  60. [68]

    Susan E Manning, Antonia M Blinn, Sabrina C Selk, Christine F Silva, Katie Stetler, Sarah L Stone, Mahsa M Yazdy, and Monica Bharel. 2022. The Massachusetts racial equity data road map: data as a tool toward ending structural racism.Journal of Public Health Management and Prac...

  61. [69]

    Shira Mitchell, Eric Potash, Solon Barocas, Alexander D’Amour, and Kristian Lum. 2021. Algorithmic fairness: Choices, assumptions, and definitions.Annual Review of Statistics and Its Application8, 1 (2021), 141–163

  62. [70]

    Marian F MacDorman, Marie Thoma, Eugene Declcerq, and Elizabeth A Howell

  63. [71]

    The National Artificial Intelligence Advisory Committee (NAIAC)

  64. [72]

    2017.Communities in Action: Pathways to Health Equity

    Engineering National Academies of Sciences and Medicine. 2017.Communities in Action: Pathways to Health Equity. The National Academies Press, Washington, DC. doi:10.17226/24624

  65. [73]

    National Association of Community Health Centers. 2023. Closing the Primary Care Gap

  66. [74]

    California Department of Health Care Access and Information. 2022. Hospi- tal Equity Measures Reporting Program. https://hcai.ca.gov/data/healthcare- quality/hospital-equity-measures-reporting-program/

  67. [75]

    Marco Morik, Ashudeep Singh, Jessica Hong, and Thorsten Joachims. 2020. Con- trolling fairness and bias in dynamic learning-to-rank. InProceedings of the 43rd international ACM SIGIR Conference on Research and Development in Information Retrieval. 429–438

  68. [76]

    Fernanda CG Polubriaginof, Patrick Ryan, Hojjat Salmasian, Andrea Wells Shapiro, Adler Perotte, Monika M Safford, George Hripcsak, Shaun Smith, Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments FAccT ’25, June 23–26, 2025, Athens, G...

  69. [77]

    Inioluwa Deborah Raji, Andrew Smart, Rebecca N White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020. Closing the AI accountability gap: Defining an end-to-end frame- work for internal algorithmic auditing. InProceedi...

  70. [78]

    Nithya Sambasivan, Erin Arnesen, Ben Hutchinson, Tulsee Doshi, and Vinodku- mar Prabhakaran. 2021. Re-imagining algorithmic fairness in india and beyond. InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Trans- parency. 315–328

  71. [79]

    Keith R Spangler, Jonathan I Levy, M Patricia Fabian, Beth M Haley, Fei Carnes, Prasad Patil, Koen Tieskens, R Monina Klevens, Elizabeth A Erdman, T Scott Troppy, et al. 2023. Missing race and ethnicity data among COVID-19 cases in Massachusetts.Journal of Racial and Ethnic He...

  72. [80]

    Center for Population Health Sciences Stanford Medicine. [n. d.]. American Family Cohort (AFC)

  73. [81]

    Office of Management and Budget

    U.S. Office of Management and Budget. 2024. 1997 Standards for Maintain- ing, Collecting, and Presenting Federal Data on Race and Ethnicity. https: //spd15revision.gov/content/spd15revision/en/history/1997-standards.html

  74. [82]

    Health Research & Educational Trust. 2013. Reducing Health Care Disparities: Collection and Use of Race, Ethnicity and Language Data. https://www.aha.org/ system/files/hpoe/Reports-HPOE/equity-care-report-august2013.PDF

  75. [83]

    Michael Veale, Max Van Kleek, and Reuben Binns. 2018. Fairness and account- ability design needs for algorithmic support in high-stakes public sector decision- making. InProceedings of the 2018 CHI Conference on Human Factors in Computing Systems. 1–14

  76. [84]

    Ioan Voicu. 2018. Using first name information to improve race and ethnicity classification.Statistics and Public Policy5, 1 (2018), 1–13

  77. [85]

    Sidney D Watson. 2024. Community Engagement, Public Reporting, and Financial Incentives: Lessons from Michigan on Tackling Racial and Ethnic Disparities in Medicaid Managed Care.Houston Journal of Health Law & Policy23, 1 (2024), 111–144

  78. [86]

    Qian Yang, John Zimmerman, Aaron Steinfeld, Lisa Carey, and James F Antaki

  79. [87]

    Harini Suresh and John Guttag. 2021. A framework for understanding sources of harm throughout the machine learning life cycle. InProceedings of the 1st ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. 1–9

  80. [88]

    Yiliang Zhang and Qi Long. 2021. Assessing fairness in the presence of missing data.Advances in Neural Information Processing Systems34 (2021), 16007–16019

  81. [89]

    Senile de- mentia with depression

    James Zou, Judy Wawira Gichoya, Daniel E Ho, and Ziad Obermeyer. 2023. Implications of predicting race variables from medical images.Science381, 6654 (2023), 149–150. FAccT ’25, June 23–26, 2025, Athens, Greece Gosciak et al. A Detecting when Race and Ethnicity is Unknown or D...

  82. [93]

    InProceedings of the 2016 CHI Conference on Human Factors in Computing Systems

    Investigating the heart pump implant decision process: opportunities for decision support tools to help. InProceedings of the 2016 CHI Conference on Human Factors in Computing Systems. 4477–4488

  83. [94]

    Yan Zhang. 2018. Assessing fair lending risks using race/ethnicity proxies.Man- agement Science64, 1 (2018), 178–197

  84. [2016]

    Race, medical mistrust, and segregation in primary care as usual source of care: findings from the exploring health disparities in integrated communities study.Journal of Urban Health93 (2016), 456–467

  85. [2019]

    InProceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19)

    Fairness Under Unawareness: Assessing Disparity When Protected Class Is Unobserved. InProceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19). Association for Computing Machinery, New York, NY, USA, 339–348. doi:10.1145/3287560.3287594

  86. [2021]

    Racial and ethnic disparities in maternal mortality in the United States using enhanced vital records, 2016–2017.American Journal of Public Health111, 9 (2021), 1673–1681

  87. [2023]

    InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency

    How redundant are redundant encodings? blindness in the wild and racial disparity when race is unobserved. InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 667–686

  88. [2024]

    https://ai.gov/wp- content/uploads/2024/06/RECOMMENDATION_Data-Challenges-and- Privacy-Protections-for-Safeguarding-Civil-Rights-in-Government.pdf

    RECOMMENDATION: Data Challenges and Privacy Protec- tions for Safeguarding Civil Rights in Government. https://ai.gov/wp- content/uploads/2024/06/RECOMMENDATION_Data-Challenges-and- Privacy-Protections-for-Safeguarding-Civil-Rights-in-Government.pdf

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.