{"id":"65ee42f1-b96f-4e12-a608-02590709c0d0","arxiv_id":"2412.07712","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Healthcare access barriers are linked to less reliable EHR data and lower sensitivity of a Type 2 diabetes risk model, and adding self-reported conditions reduces the gap.","lead":"This study uses All of Us survey and electronic health record data from 134,513 participants to show that people facing cost or delay barriers to care have less complete EHR records and lower sensitivity in a Type 2 diabetes risk model. It finds that adding patient self-reported conditions to the model narrows the performance gap.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 9.4-point sensitivity gap likely reflects differential outcome ascertainment in EHR labels, not a true model performance disparity; the mitigation result may simply recover prevalent diabetes that the EHR missed at baseline.","rationale":"I agree with the reader's conditional assessment and the identification of outcome-label under-ascertainment as the key threat. My stress test sharpens this: the bias is not just noise; it is aligned with the access groups that define the headline comparison, so it can create a spurious sensitivity gap even when model quality is identical. The reported confidence intervals around sensitivity are also implausibly narrow for the event counts implied by a 0.6% incidence rate, which independently suggests the uncertainty quantification needs correction. The reliability analysis (missing EHR diagnosis rates for self-reported conditions) does not depend on the diabetes outcome and remains a credible descriptive finding. But the abstract's claim about model performance and the mitigation experiment should be interpreted as hypothesis-generating until outcome labels are validated against self-report, labs, or chart review. No code or data are provided, so these checks cannot be done from the manuscript alone. The verdict remains CONDITIONAL: the descriptive reliability result is likely real, but the quantitative performance-gap and mitigation claims require stronger validation.","tokens_in":746,"tokens_out":926,"duration_ms":61740,"concrete_test":"Re-run the diabetes prediction experiment on the restricted cohort of participants who (i) had at least one EHR encounter in the 2-year follow-up window and (ii) did not self-report Type 2 diabetes at baseline. Re-estimate the sensitivity gap between standard-care and low-access groups. If the 9.4-point gap collapses or becomes non-significant, the headline claim is driven by differential outcome ascertainment rather than by genuine differences in predictive performance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing issue is the outcome definition in the diabetes prediction task. The Methods define the outcome as an indicator for whether someone had a record of Type II diabetes in the two years following the index date, and the cohort excludes only participants with an EHR record of Type 2 diabetes at baseline. For low-access patients, diabetes that is already present but undocumented at baseline is not excluded; if it later appears in the EHR it is labeled an incident event, and if it does not appear within two years the patient is labeled a non-event. Consequently, the outcome label is differentially misclassified by access group. A model trained on these labels can show lower sensitivity for low-access groups even when underlying risk discrimination is identical, because the positive labels it is asked to detect are partly a function of healthcare contact rather than true disease onset. The mitigation result is vulnerable to the same artifact: adding self-reported conditions as features raises sensitivity by 11.2 points for cost-constrained patients, but if those self-reports capture baseline diabetes that the EHR missed, the model is detecting prevalent disease, not improving risk prediction. This does not invalidate the descriptive EHR-reliability comparison, but it directly undermines the quantitative claim of a 9.4-point sensitivity gap and the interpretation of the mitigation as improving equity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses All of Us survey and EHR data (N=134,513) to examine whether healthcare access barriers affect EHR data completeness and the performance of a machine learning model for 2-year Type 2 diabetes incidence. The authors report that low-access groups (cost-constrained or delayed care) have higher rates of self-reported conditions missing from the EHR (for 29 of 37 conditions), and in the diabetes prediction task (N=52,046) show lower balanced accuracy (3.6 percentage points) and sensitivity (9.4 percentage points) than standard-care patients. They propose two mitigations: adding an access indicator and adding self-reported condition features; the latter is reported to increase sensitivity for cost-constrained patients by 11.2 percentage points. The paper concludes that access barriers propagate through the machine learning pipeline and that patient-reported data can help close the gap.","tokens_in":17745,"tokens_out":7621,"duration_ms":68254,"significance":"If the main claims hold, this would be a valuable large-scale empirical demonstration that healthcare access affects both EHR data quality and clinical prediction performance, with direct implications for algorithmic fairness and for the design of data collection in learning health systems. The study leverages a large, diverse national cohort with linked survey and EHR data, and the authors are transparent about several limitations. However, the prediction-performance claim rests on an outcome label that is itself an EHR record, and the mitigation analysis likely includes the self-report of the outcome condition. These issues make the central quantitative claims uncertain and require additional analyses or reframing. The descriptive EHR-reliability comparisons are simpler and more robust, though they depend on the assumption that self-report is ground truth.","major_comments":[{"comment":"The outcome is an indicator of an EHR record of Type 2 diabetes within two years after the index date, and the baseline exclusion removes only participants with an EHR-recorded Type 2 diabetes diagnosis. For low-access patients, diabetes that is already present but undocumented at baseline is not excluded, and incident diabetes may go unrecorded if the patient does not seek care within the follow-up window. The 9.4-percentage-point sensitivity gap could therefore reflect differential outcome ascertainment by access group rather than a true difference in model discrimination. The authors acknowledge in the Limitations that the outcome labels 'may be imprecise especially for individuals with lower access to care,' but they continue to interpret the gap as a performance disparity. I recommend reframing the task as predicting EHR-recorded T2D incidence and adding robustness checks, such as restricting the sample to participants with at least one EHR encounter in the follow-up period, adjusting for visit frequency, or using self-reported T2D as a secondary outcome.","section":"Algorithmic predictive performance is lower for patients without access to care; Methods: Implications for Clinical…"},{"comment":"The second mitigation adds self-reported condition features to the prediction model. Because the cohort excludes only EHR-recorded T2D at baseline, participants with self-reported T2D at baseline are included. If the self-reported Type 2 diabetes variable is among the added features, the model can trivially identify patients who already know they have diabetes and who are likely to receive an EHR diagnosis within two years. This would inflate the reported 11.2-percentage-point sensitivity increase and does not represent a genuine improvement in incident risk prediction. Please report the results excluding the self-report of the outcome condition, and specify exactly which self-reported conditions were included in the mitigation model.","section":"Potential Solutions; Methods: Implications for Clinical Prediction Models"},{"comment":"The reported 95% confidence intervals for sensitivity are numerically inconsistent with the event counts. For example, the standard-care group has roughly 142 positive events (0.6% of 23,705), but the reported sensitivity of 66.4% with 95% CI 65.9–67.0 implies a standard error of about 0.0028, which would require an effective sample size in the tens of thousands rather than the number of true positives. The intervals appear to have been computed with the full cohort size in the denominator rather than the number of positive events. This error affects the statistical significance claims for the performance gaps and should be corrected by recomputing the intervals with the appropriate denominator.","section":"Algorithmic predictive performance is lower for patients without access to care; Material and Methods, final paragraph"},{"comment":"The EHR-reliability analysis treats the participant's self-report as ground truth, stated in the Methods as 'assuming that the patient is the source of truth, and that any missing EHR record of a condition is an error.' This assumption is acknowledged but not examined. If low-access participants are more likely to self-report conditions that are not yet diagnosed (e.g., because they are experiencing untreated symptoms), the observed missingness gap would be inflated relative to true documentation errors. I recommend reframing the result as 'discordance between self-report and EHR' rather than 'EHR reliability,' and adding a sensitivity analysis that validates self-report against an objective source such as medication records or lab values for a subset of conditions.","section":"Patients with lower access to care have lower EHR reliability; Comparing Self-Reported Conditions to EHR Conditions"}],"minor_comments":[{"comment":"The cost-constrained and delayed-care groups are not mutually exclusive (69.1% of the delayed-care group also report affordability concerns), but the paper does not explain how participants in both groups are assigned in the analyses or in Figure 3. Please clarify whether the groups are treated as overlapping or whether participants in both are assigned to a single 'low access' category.","section":"Sample; Table 1"},{"comment":"The text states that balanced accuracy drops significantly for the cost-constrained and delayed-care groups but does not report the numerical gap (the abstract reports 3.6 percentage points) or its confidence interval. Please provide these values alongside the sensitivity results.","section":"Algorithmic predictive performance is lower for patients without access to care"},{"comment":"The caption says the figure shows 'all conditions with a statistically significant difference,' while the text reports 29 of 37 conditions were significant. Please clarify whether non-significant conditions are omitted and how the FDR correction was applied in the figure.","section":"Figure 2 caption"},{"comment":"The caption states 'top ten most common self-reported conditions' but the table lists 15 conditions; please correct the caption.","section":"Table S2 caption"},{"comment":"The list of clinical measures includes 'diabetes' alongside lab-based measurements, which is potentially confusing given that the outcome is Type 2 diabetes. Please clarify that this refers to a glucose-derived indicator (as in Table S3) and distinguish it from the outcome definition.","section":"Methods: Implications for Clinical Prediction Models"},{"comment":"There is a typo on page 2: 'patients with either cost-constrained care or delayed case' should read 'delayed care.'","section":"Sample section"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for cs.CY as an empirical study of data inequality in clinical ML. The central descriptive finding about self-report/EHR discordance is publishable, but the prediction-performance claim is currently overstated because the outcome label is likely differentially misclassified by access, and the mitigation result appears to be confounded by including the self-reported outcome itself. The confidence-interval error is a concrete statistical flaw that should be caught in revision. I recommend major revision with a request for sensitivity analyses around the outcome definition and careful statistical re-evaluation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a paper worth reading for its descriptive reliability analysis, but the prediction claim is shakier than the abstract suggests. The authors show, using 134k All of Us participants, that people who report cost-constrained or delayed care have significantly higher rates of self-reported conditions missing from the EHR for 29 of 37 conditions. That finding is straightforward, plausibly real, and useful—it confirms with direct survey data what earlier work only proxied with race or insurance. The age-adjusted prevalence differences and the regression are fine.\n\nThe soft spot is the diabetes prediction task. The outcome is an incident EHR record of Type 2 diabetes within two years, and the baseline cohort only excludes people with an EHR record of diabetes at baseline. Anyone with self-reported diabetes but no EHR record is still in the cohort, and when that condition later shows up in the EHR it's counted as an incident event. Low-access participants report diabetes more often, so the model is partly learning to detect prevalent, undocumented disease. That alone could produce most of the 9.4-point sensitivity gap. The mitigation result—adding self-reported conditions as features improves sensitivity by 11.2 points—is the same artifact: the model recovers diabetes that the EHR missed at baseline. This doesn't invalidate the missingness-descriptive part, but it directly undermines the \"performance gap\" interpretation.\n\nThere's also a statistical problem. Sensitivity of 57% with a 95% CI of 56.3–57.7 for a group with maybe 100 events is impossible. The CI is off by an order of magnitude; they must have misapplied the CLT. The paper even admits the labels can't be validated, which is honest, but it doesn't fix the core issue.\n\nI'd send this to peer review—the topic matters and the descriptive core is solid—but I'd require the authors to redo the prediction analysis excluding self-reported diabetes at baseline, re-report CIs correctly, and either drop or reframe the mitigation as recovering undocumented prevalence.","headline":"Strong descriptive evidence on EHR missingness by access, but the headline sensitivity gap is likely inflated by an outcome-label artifact and the confidence intervals don't survive scrutiny.","tokens_in":18269,"tokens_out":3282,"would_cite":true,"duration_ms":31402,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Low access to care weakens EHR data and diabetes risk scores.","keywords":["healthcare access","EHR reliability","clinical risk prediction","algorithmic fairness","Type 2 diabetes","All of Us","patient-reported outcomes","missing data"],"falsifier":"Conduct a study where a sample of All of Us participants receives an independent clinical assessment (physical exam, lab panel, or adjudicated chart review) that does not rely on either the survey or the routine EHR, then compare missing-EHR rates for low-access versus standard-care groups against that gold standard. If the gap shrinks or disappears once self-report error is controlled, the reliability and prediction findings are largely an artifact of differential self-reporting. A complementary check: re-run the diabetes model using chart-reviewed or self-reported diabetes as the outcome instead of EHR codes; if sensitivity still differs by access group, the gap reflects true model failure, not just label noise.","tokens_in":17322,"feed_emoji":"🩺","tokens_out":9127,"duration_ms":73584,"temperature":0.7,"pith_summary":"This paper asks whether uneven access to healthcare distorts the electronic health records (EHRs) that machine-learning models learn from, and whether the distortion harms predictions for the patients who see doctors least. Using survey and EHR data from 134,513 All of Us participants, the authors show that people who delayed care or could not afford it had higher rates of conditions they reported themselves but that never appeared in their EHRs — significantly higher for 29 of 37 conditions. In a model predicting 2-year Type 2 diabetes incidence, those same patients were flagged as high-risk far less often: sensitivity fell by 9.4 percentage points and balanced accuracy by 3.6 points relative to standard-care patients. Adding the patients' self-reported conditions as model features raised sensitivity for the cost-constrained group by 11.2 percentage points, largely closing the gap. The paper's core claim is that access barriers propagate through the data pipeline, so making models fair requires fixing both data collection and algorithms.","feed_headline":"Low access to care weakens EHR data and diabetes risk scores","feed_subtitle":"Patients who delay or can't afford care have sparser records, and models miss more of their diabetes cases.","key_machinery":"The central measurement device is the missing EHR diagnosis rate: the share of participants with a self-reported condition who have no record of that condition in the EHR, computed under the stated assumption that the patient is the source of truth. The prediction mechanism is a 2-year Type 2 diabetes incidence task built from a 2-year EHR lookback window of conditions, labs, vitals, procedures, medications, and demographics, trained with a LASSO logistic regression and evaluated at the Youden-J threshold stratified by access group. The load-bearing comparison is sensitivity at that threshold, which drops sharply for low-access patients, and the mitigation experiment adds self-reported conditions as features to test whether missing EHR data causes the gap.","core_discovery":"The paper's central discovery is a two-part empirical regularity in the All of Us cohort. First, EHR reliability — measured as the rate at which a patient's self-reported conditions are missing from their record — is systematically worse for patients with cost-constrained or delayed care: for 29 of 37 conditions the missing-diagnosis rate is significantly higher, and low-access participants average 2.0 missing EHR diagnoses versus 1.7 for standard-care participants. Second, this data degradation carries into clinical prediction: a logistic-regression model with L1 regularization for 2-year Type 2 diabetes incidence achieves comparable AUC across access groups but significantly lower balanced accuracy and sensitivity for low-access patients (sensitivity 57–58% versus 66.4% for standard care), meaning the model misses more true diabetes cases among people who already face access barriers. Including self-reported conditions as extra features lifts sensitivity for the cost-constrained group by 11.2 percentage points and largely erases the performance gap, indicating that missing data — not just algorithm design — drives the disparity.","pith_inferences":["If the paper is right, the reliability gap may be overstated to the extent that untreated symptoms make low-access patients more likely to report a condition; a validation study using biomarkers or adjudicated diagnoses would sharpen the estimate.","The same missing-data mechanism likely affects other EHR-based tasks such as readmission prediction or comorbidity adjustment, a generalization the authors leave implicit.","Because the outcome label is an EHR record of diabetes, the sensitivity gap could partly reflect under-ascertainment of the outcome rather than model failure; chart-review or self-reported outcomes would separate label noise from model error.","The paper notes the access coefficient doubles when only the prior year of EHR is used; this suggests shorter lookback windows may amplify the equity gap, a testable extension."],"forward_implications":["EHR-based clinical risk scores will systematically under-flag Type 2 diabetes risk in patients who delay or cannot afford care, directing fewer preventive interventions to the group with the highest self-reported burden.","Quality measures and epidemiological estimates that treat the EHR as ground truth will overstate the health of low-access populations; access-adjusted comparisons are needed.","Collecting patient-reported conditions together with EHR data closes most of the sensitivity gap in diabetes incidence models, a concrete data-collection priority for health systems.","Adding a simple access-group label as a model feature does not fix the gap; the improvement comes from supplying the missing clinical information, not from the label itself.","Because the missingness pattern spans most examined conditions, similar access-driven degradation should be expected in other chronic-disease risk models built from EHR data."],"supporting_citations":[{"why":"Supplies the All of Us cohort used for every estimate, linking healthcare-access survey responses to standardized OMOP EHR data.","marker":"12"},{"why":"Prior evidence that AI diagnostic performance is lower for underserved groups, framing the paper's access-based hypothesis.","marker":"7"},{"why":"Shows an algorithm trained on utilization encodes racial disparities, motivating the shift from protected attributes to access as the mechanism.","marker":"8"},{"why":"Documents discordance between survey-collected and EHR health data, the methodological precedent for the reliability analysis.","marker":"21"},{"why":"Provides a trial-based concordance estimate between patient-reported and electronic health data, the closest benchmark for the paper's missingness rates.","marker":"22"},{"why":"Shows patient-initiated issues are often omitted from primary care notes, supporting the assumption that a missing EHR record is an error.","marker":"23"},{"why":"Supplies the Benjamini-Yekutieli false-discovery correction used to declare 29 of 37 condition differences significant.","marker":"26"},{"why":"Establishes that adjusting clinical algorithms for a patient characteristic can correct disparities in data quality, framing the mitigation test.","marker":"27"},{"why":"Argues for embedding patient-reported outcomes in AI health technologies, the conceptual basis for the self-report feature intervention.","marker":"29"}],"fun_headline_variants":["Care access gaps skew EHR data and diabetes predictions","When care is delayed, diabetes risk models miss more cases","EHR reliability drops for patients with limited care access","Diabetes models underperform for patients who skip or delay care","Missing data from care barriers weakens diabetes risk scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a participant's self-report on the Personal and Family Health History survey is the ground truth for whether they truly have a condition, so a condition missing from the EHR counts as an error; if self-reports are inaccurate in ways that differ by access group, the measured reliability gap — and the prediction gap built on it — is overstated.","fun_headline_variants_meta":{"raw":{"variants":["Care access gaps skew EHR data and diabetes predictions","When care is delayed, diabetes risk models miss more cases","EHR reliability drops for patients with limited care access","Diabetes models underperform for patients who skip or delay care","Missing data from care barriers weakens diabetes risk scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1410,"prompt_tokens":1022,"completion_tokens":388,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":638,"completion_tokens_details":{"reasoning_tokens":311}},"tokens_in":638,"tokens_out":388,"duration_ms":4108,"temperature":1.0,"reasoning_tokens":311,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:33:50.562398+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct a study where a sample of All of Us participants receives an independent clinical assessment (physical exam, lab panel, or adjudicated chart review) that does not rely on either the survey or the routine EHR, then compare missing-EHR rates for low-access versus standard-care groups against that gold standard. If the gap shrinks or disappears once self-report error is controlled, the reliability and prediction findings are largely an artifact of differential self-reporting. A complementary check: re-run the diabetes model using chart-reviewed or self-reported diabetes as the outcome instead of EHR codes; if sensitivity still differs by access group, the gap reflects true model failure, not just label noise.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a trial-based concordance estimate between patient-reported and electronic health data, the closest benchmark for the paper's missingness rates."}],"review_version":1}