{"id":"d3d4f9fb-8b1b-4e2a-a887-8e87293543d2","arxiv_id":"2607.22504","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Four eGFR equations in 1.9M adults show near-identical prediction of kidney failure (AUROC 0.846-0.862) but large differences in CKD prevalence (8.5%-10.8%) by equation and region of birth.","lead":"Using electronic health records from 1.9 million Israeli adults, this study compared four standard formulas that estimate kidney function from a blood test. It found the formulas differ little in predicting kidney failure, but change how many people are labeled with chronic kidney disease, from 8.5% to 10.8%.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Figure 3's AUROC ranking rests on a sicker UACR-tested, 5-year-survivor subpopulation; without a sensitivity analysis in the full CKD cohort, the 'little effect on prediction' claim is not externally anchored.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing point: the predictive comparison is performed in a UACR-available, 5-year-survivor subpopulation that differs markedly from the full cohort. I could not find a more decisive internal inconsistency that would overturn the main conclusion. The paper's strengths—enormous sample size, use of published equations without recalibration, and an outcome definition with long follow-up—make the central prevalence/staging message likely robust. However, the 'little effect on prediction' claim is only as strong as the AUROC comparison in the selected subpopulation. The manuscript's own limitation statement concedes selection bias, but no empirical check is offered. A full-cohort eGFR-only AUROC analysis would directly test whether the ranking and the small effect size are generalizable. I also noted the text's EKFC Q values appear sex-labeled counterintuitively relative to the observed medians (0.62 for men vs 0.80 for women); this should be clarified, but it is likely a reporting typo rather than a structural flaw in the argument, and the reader's recommendation to release analytic code would settle it. Overall, the reader's CONDITIONAL verdict remains appropriate.","tokens_in":9460,"tokens_out":12291,"duration_ms":130376,"concrete_test":"Recompute the eGFR-only AUROC for 5-year kidney failure in the full CKD cohort (n=284,399) without requiring UACR, treating death as a competing risk and censoring at disenrollment rather than excluding deaths, and compare to Figure 3A. If the MDRD vs EKFC AUROC difference (0.016 in the subpopulation) changes by more than ~0.005, or if the ordering is not preserved, the reported ranking is an artifact of UACR/survival selection. If the ordering and magnitude are preserved, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim's first half—that eGFR equation choice has little effect on incident kidney outcome prediction—rests on the AUROC/KFRE comparisons in Figure 3, calculated in 182,034 patients with CKD, available UACR, and no death or disenrollment within 5 years. Table 1 shows this subpopulation is far sicker than the full cohort (diabetes 65.9% vs 19.0%; dialysis during follow-up 4.81% vs 0.77%; a high proportion aged 75+). If UACR availability or 5-year survival is correlated with disease severity and with the eGFR values produced by each equation, the relative discriminative ranking (MDRD 0.862 vs EKFC 0.846) may not generalize to the full 1.9M population or to primary-care settings. The manuscript acknowledges selection bias only qualitatively as a limitation; no inverse-probability weighting, multiple imputation, or full-cohort sensitivity analysis is provided. Because the 'little effect on prediction' conclusion depends on these numeric differences and their ordering, the assumption that selection is ignorable is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This retrospective cohort study uses Clalit Health Services data on approximately 1.9 million adults to compare four creatinine-based eGFR equations—2006 MDRD, 2009 CKD-EPI, 2021 CKD-EPI, and 2021 EKFC—in terms of calibration, discrimination of 5-year kidney failure, and age-adjusted CKD prevalence. The authors report that 2021 EKFC yields lower GFR estimates and 2021 CKD-EPI higher GFR estimates than earlier equations; that 2006 MDRD has the highest AUROC (0.862) and 2021 EKFC the lowest (0.846) for predicting kidney failure; that KFRE improves prediction substantially over eGFR alone; and that age-adjusted CKD prevalence varies by equation (8.5% to 10.8%) and by region of birth (8.9% to 15.3%). The central conclusion is that equation choice has little effect on incident kidney outcome prediction but large effects on CKD staging, prevalence, and clinical classification.","tokens_in":9672,"tokens_out":6263,"duration_ms":65998,"significance":"If the findings are robust, this is a valuable head-to-head comparison in a large, multiethnic, population-based cohort with long follow-up. The study's design avoids circularity: the equations are external, and the authors deliberately use default EKFC Q values rather than recalibrating to the Clalit population. The use of Fine-Gray subdistribution hazards for stage-specific outcomes and the inclusion of KFRE alongside eGFR alone are strengths. The prevalence differences across equations and birth regions have direct implications for health-system decisions and health equity. The main limitation is that the predictive comparisons in Figure 3 are restricted to a UACR-available, 5-year-survivor subpopulation that is substantially sicker than the full cohort, so the generalizability of the AUROC ranking and the 'little effect on prediction' claim needs additional support.","major_comments":[{"comment":"The AUROC/KFRE comparisons that support the 'little effect on prediction' claim are computed in 182,034 patients with available UACR and with at least 5 years of follow-up without death or disenrollment. Table 1 shows this subpopulation is much sicker than the full cohort: CKD 100% vs 15.1%, diabetes 65.9% vs 19.0%, dialysis during follow-up 4.81% vs 0.77%, and death 46.5% vs 13.9%. If UACR availability or 5-year survival is informative, the relative ranking (MDRD highest, EKFC lowest) may not generalize to the full cohort or to primary-care populations. The Discussion acknowledges selection bias only qualitatively (second limitation). Because eGFR-only discrimination does not require UACR, an AUROC analysis in the full CKD cohort (n=284,399) and an inverse-probability-weighted or multiple-imputation sensitivity analysis for UACR availability are needed to support the general claim.","section":"Figure 3 and Table 1"},{"comment":"The handling of death is inconsistent across analyses. Figures 1 and 2 use the Fine-Gray subdistribution hazard with death as a competing risk, but Figure 3 excludes deaths and disenrollments within 5 years entirely. This may bias discrimination estimates, especially in a subpopulation with 46.5% mortality. A competing-risk-aware discrimination measure (e.g., inverse-probability-of-censoring-weighted AUC) or a sensitivity analysis treating death as a composite outcome would clarify whether the small AUROC differences (≤0.016) are robust to how the competing event is handled.","section":"Figure 3 and Methods (Statistical analysis)"},{"comment":"The prevalence analyses assume that individuals without UACR data have no albuminuria: 'Individuals without data on either urinary albumin or urinary creatinine were assumed to not have albuminuria.' Only 27.6% of the full cohort has UACR recorded, and 7.5% have UACR>30 mg/g. If missingness is informative with respect to region of birth or eGFR level, the age-adjusted prevalence differences across equations and regions (Figure 4) will be affected. A sensitivity analysis using an eGFR-only definition of CKD or imputation of missing UACR would strengthen the prevalence claim.","section":"Methods, Outcome definitions; Figure 4"}],"minor_comments":[{"comment":"The text states the study population comprised 1,908,042 adults, while the Abstract and Table 1 report 1,909,042. Please correct the inconsistency.","section":"Results, first paragraph"},{"comment":"The caption gives n=182,034, while Table 1 reports 191,370 for 'CKD and Recorded UACR.' The additional exclusions (5-year follow-up, no death/disenrollment, complete covariate data) should be stated explicitly so readers can reconcile the numbers.","section":"Figure 3 caption"},{"comment":"The CKD definition requires 'two or more consecutive measures spanning at least three months' for eGFR, but for UACR the text says 'any time frame prior to the index date' without specifying a second measurement. Clarify whether a single UACR≥30 suffices and how this interacts with the 'assumed no albuminuria' rule for missing UACR.","section":"Methods, Outcome definitions"},{"comment":"The sentence 'the age-adjusted prevalence of kidney disease in the adult Clalit population was comparable to the global all-age crude prevalence of 9.1%' compares an age-adjusted adult prevalence to a crude all-age global prevalence. This comparison is not age-standardized and should be rephrased or removed.","section":"Discussion"},{"comment":"The sentence describing stage-specific risk as 'divergent for 2006 MDRD and 2021 EKFC (higher and lower rates of kidney failure respectively)' is ambiguous. Please specify which equation corresponds to higher versus lower rates within each CKD stage.","section":"Figure 2"},{"comment":"The cross-tabulations in Figure 1 are based on 18,202 participants with measured creatinine clearance, a subgroup with high dialysis (8.9%) and death (40.7%) rates. Reporting cell counts and confidence intervals for the 5-year risk estimates would improve transparency.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful, well-designed comparison of eGFR equations with a large and diverse population. The main barrier is the selection-bias issue in Figure 3; if the authors can provide a full-cohort eGFR-only AUROC analysis and a sensitivity analysis addressing informative UACR missingness, I would support publication. The prevalence claim also needs a sensitivity analysis for the missing-UACR assumption."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the largest head-to-head comparison of 2006 MDRD, 2009/2021 CKD-EPI, and 2021 EKFC in one cohort, and it gives the cleanest evidence yet that equation choice changes CKD prevalence and stage distributions far more than it changes outcome prediction. The prevalence numbers are probably right. The predictive ranking is less certain: it rests on a sicker UACR-tested subpopulation conditioned on five-year survival, and the authors don't do enough to show that selection is ignorable.\n\nWhat's new: prior comparisons only pitted 2009 against 2021 CKD-EPI in Swedish and US cohorts. Adding MDRD and EKFC in ~1.9M adults with region-of-birth stratification is a real contribution. The decision to use default EKFC Q values rather than refitting to Clalit is the right call for fairness. Fine-Gray models, direct age standardization, and a clear tabulation of the subcohorts are all solid. The acknowledgment that UACR-based analysis may be selected is present, but it's only qualitative.\n\nWhere it's soft: Figure 3's AUROC/KFRE comparisons use 182,034 patients with CKD, UACR, and no death/disenrollment for 5 years. Table 1 shows that group has 65.9% diabetes and 4.81% dialysis initiation vs 19.0% and 0.77% in the full cohort. If UACR testing or survival is correlated with severity and with the eGFR values produced by each equation, the observed ordering—MDRD best, EKFC worst—may not hold in the full CKD population or in primary care. The authors never run a sensitivity analysis in the full CKD cohort, and the AUC differences are small (≤0.016), so the ranking is not robustly established. This doesn't kill the central claim—\"little effect on prediction\" still has support—but it weakens the stronger version that MDRD is genuinely more discriminative. Minor: prevalence estimates are reported without CIs; the abstract says 1,909,042 while the text says 1,908,042; no code or data are available.\n\nBottom line: This deserves a serious referee. The prevalence and calibration results are likely to hold, and the subpopulation concern is addressable. I'd ask for a sensitivity analysis in the full CKD cohort or an explicit selection model, CIs for prevalence, and a fix for the count typo. I'd bring it to reading group; I'd cite the prevalence numbers.","headline":"Largest head-to-head of four eGFR equations to date; the prevalence and calibration results are likely robust, but the predictive ranking rests on a sicker, survival-conditioned UACR subpopulation and needs a sensitivity analysis.","tokens_in":10285,"tokens_out":2763,"would_cite":true,"duration_ms":28076,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The choice of eGFR equation has little effect on predicting who develops kidney failure, but it substantially changes how many people are labeled with chronic kidney disease and at what stage.","keywords":["eGFR","chronic kidney disease","kidney failure prediction","MDRD","CKD-EPI","EKFC","race-neutral equations","CKD prevalence"],"falsifier":"A prospective study in a general primary-care population that systematically measures UACR in everyone, follows all participants for 5 years without informative dropout, and compares the four equations' AUROC for kidney failure would settle whether the paper's ranking (2006 MDRD best, 2021 EKFC worst) is generalizable; if the AUROC differences shrink to zero or reverse, the central predictive claim would be falsified.","tokens_in":9305,"feed_emoji":"🩺","tokens_out":2477,"duration_ms":28993,"temperature":0.7,"pith_summary":"This paper asks whether the switch to race-neutral equations for estimating kidney function changes clinical predictions or only reshuffles labels. Using 1.9 million adults followed for up to 10 years, it finds that the oldest equation (2006 MDRD) best discriminates 5-year kidney failure, while the newest European equation (2021 EKFC) performs worst, yet the absolute differences are small. The bigger effect is on classification: equations disagree enough to move people between CKD stages and to shift age-adjusted CKD prevalence from 8.5% to 10.8% depending on the equation. Across regions of birth, prevalence varies even more, from 8.9% to 15.3%. The paper concludes that equation choice matters more for epidemiology and clinical thresholds than for predicting which individual will progress to kidney failure.","feed_headline":"Kidney equation swap shifts CKD prevalence, not prognosis","feed_subtitle":"Across 1.9 million adults, the oldest GFR equation best predicted kidney failure, while newer equations differ mainly in how many people are","key_machinery":"The central objects are the four eGFR equations: 2006 MDRD, 2009 CKD-EPI, 2021 CKD-EPI, and 2021 EKFC. These are compared using three mechanisms: cross-tabulation of CKD stage reclassification against measured creatinine clearance, survival curves with death as a competing risk, and AUROC/AUPRC for predicting 5-year kidney failure (initiation of dialysis or transplant), both as standalone eGFR and as an input to the Kidney Failure Risk Equation (KFRE). The work also uses age-standardized prevalence to compare CKD burden across regions of birth. The KFRE, which adds age, sex, and urinary albumin-to-creatinine ratio, improves AUROC by roughly 4.4 percentage points on average, far more than the","core_discovery":"The paper claims that among four creatinine-based eGFR equations—2006 MDRD, 2009 CKD-EPI, 2021 CKD-EPI, and 2021 EKFC—the choice of equation barely changes the prediction of incident kidney failure, but it significantly changes the correspondence between CKD stage and failure risk, and it changes the measured prevalence of CKD across populations. Concretely, 2006 MDRD yields the highest AUROC for 5-year kidney failure (0.862) and 2021 EKFC the lowest (0.846), a difference of only 0.016. Meanwhile, 2021 EKFC produces lower GFR estimates and 2021 CKD-EPI higher ones relative to prior equations, causing net reclassification to more severe stages with EKFC and less severe stages with 2021 CKD-EP","pith_inferences":["If the predictive plateau is real, future gains in kidney failure risk prediction will likely come from biomarkers such as cystatin C or from directly incorporating albuminuria, not from further tweaks to creatinine-based equations.","The equal predictive performance of race-stratified and race-neutral equations in this large multiethnic population suggests that removing race from eGFR equations does not sacrifice prognostic accuracy, a point the paper's data support but which is not the paper's central claim.","The UACR-available subpopulation used for discrimination analyses is markedly sicker than the full cohort, so the ranking of equations (MDRD highest, EKFC lowest) may not generalize to general primary-care populations if UACR missingness is informative; the paper acknowledges this as selection bias.","Prevalence differences across regions of birth of similar magnitude to equation differences imply that regional kidney disease burden estimates are sensitive to equation choice, which has implications for global health resource allocation."],"forward_implications":["If the central claim holds, health systems can choose between race-neutral eGFR equations without meaningfully changing which patients progress to kidney failure; prediction accuracy is nearly equation-independent.","The choice of equation will change CKD prevalence estimates by up to 2.3 percentage points nationally, and this difference is comparable to or larger than differences between regions of birth.","Stage reclassification from equation choice could affect clinical actions tied to GFR thresholds, including specialist referral, transplant listing, and dosing of renally cleared medications.","The KFRE's larger gains suggest that adding albuminuria and other predictors matters more than refining the creatinine-based equation itself.","The plateau in discrimination among eGFR equations points to serum creatinine as a limiting factor, motivating further study of alternative filtration markers."],"fun_headline_variants":["eGFR equation choice skews CKD prevalence, not risk prediction","In 1.9 million adults, GFR equation barely changes kidney failure prediction","Equation choice shifts CKD prevalence more than kidney failure risk","Multiethnic study: eGFR equation changes prevalence, not prognosis","Oldest eGFR equation edges out newer ones in kidney failure prediction"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The ranking of eGFR equations by discrimination is based on a subpopulation that has both UACR measurements and complete follow-up, and this subpopulation is much sicker than the full 1.9 million-person cohort, so the ranking may not hold if missing UACR or dropout is not random.","fun_headline_variants_meta":{"raw":{"variants":["eGFR equation choice skews CKD prevalence, not risk prediction","In 1.9 million adults, GFR equation barely changes kidney failure prediction","Equation choice shifts CKD prevalence more than kidney failure risk","Multiethnic study: eGFR equation changes prevalence, not prognosis","Oldest eGFR equation edges out newer ones in kidney failure prediction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000917,"raw_usage":{"total_tokens":3882,"prompt_tokens":960,"completion_tokens":2922,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":704,"completion_tokens_details":{"reasoning_tokens":2833}},"tokens_in":704,"tokens_out":2922,"duration_ms":18300,"temperature":1.0,"reasoning_tokens":2833,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T04:31:24.597908+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A prospective study in a general primary-care population that systematically measures UACR in everyone, follows all participants for 5 years without informative dropout, and compares the four equations' AUROC for kidney failure would settle whether the paper's ranking (2006 MDRD best, 2021 EKFC worst) is generalizable; if the AUROC differences shrink to zero or reverse, the central predictive claim would be falsified.","supporting_citations":[],"review_version":1}