REVIEW 2 major objections 6 minor 178 references
When Does Survey-Aware Cross-Validation Matter? The ICC, Not the Design Effect
T0 review · 2 major / 6 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read Whether survey-aware cross-validation matters is answered before modeling by two cheap ICC diagnostics, not the design effect.
desk verdict Cheap dual-ICC pre-check closes Wieczorek’s open question; three national surveys sit in a true null regime and DEFF is the wrong trigger. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two pre-analysis diagnostics: the one-way ANOVA ICC of the binary outcome and of the linear predictor from a single full-sample weighted LASSO, both computed within the survey’s primary sampling units. They measure how much cluster-specific signal a naive fold split could leak; a calibrated simulation shows naive-CV optimism rising with this ICC while cluster-level folds remain honest.
What would settle it
A real survey or a simulation cell in which the linear-predictor ICC is low (well below 0.1) yet naive CV is still optimistically biased for new-cluster performance by more than 0.01 AUC relative to cluster-level folds, or the reverse: high ICC with no practical scheme difference.
Extended reading notes
Core claim
The within-cluster ICC of the outcome and of a preliminary linear predictor—not the design effect—correctly anticipates whether naive versus design-respecting cross-validation will differ by a practical amount. In three national health surveys sitting at linear-predictor ICCs of 0.02–0.04, every paired |ΔAUC| was well under 0.01, and the only statistically nonzero difference was pessimistic, opposite to the optimism that cluster leakage produces.
Load-bearing premise
That a single in-sample linear-predictor ICC (plus the outcome ICC) is a good enough early warning of leakage risk even for nonlinear learners such as random forests, and that the simulation-derived threshold of roughly 0.1 transfers as a decision rule outside the stylized positive-control grid.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper answers an open question left by Wieczorek et al. (2022): when does design-respecting (survey-aware) K-fold CV change conclusions relative to naive random folds? It proposes two pre-analysis diagnostics—the within-PSU ICC of the outcome and of a preliminary weighted-LASSO linear predictor—and argues that these, not the design effect (DEFF), flag the cluster-leakage optimism that survey CV guards against. A simulated positive control (and a factorial extension across learners, weights, stratification, imbalance, and decoupled clustering channels) shows naive-CV optimism rising with the linear-predictor ICC while cluster-level folds stay honest. Paired naive-versus-design-respecting evaluations on three national health surveys (NHIS chronic pain reanalysis; NHANES diabetes; YRBS suicidality with a prespecified ICC rule) all sit at ICC_lp 0.02–0.04 and yield practically null scheme differences (|ΔAUC| well under 0.01), with the only zero-excluding interval in the pessimistic direction. The paper also supplies a fallback hierarchy when stratified Survey CV is infeasible in public-use designs and documents weight-handling plug-in errors that produced order-of-magnitude artifacts.
Significance. If the result holds, the paper converts an open methodological worry into an operational pre-analysis check that can be run before any model is validated. That is practically valuable: design-respecting folds are already available (surveyCV), but applied work still defaults to random folds, and the paper shows both when that default is safe and when it is not. Strengths that should count in the assessment include (i) a calibrated positive-control simulation with a factorial extension covering LASSO, random forest, unequal weights, stratification, and decoupled clustering channels; (ii) three real surveys spanning reanalysis, prospective application, and a prespecified decision rule; (iii) an explicit feasibility fallback hierarchy; (iv) documentation of weight-handling errors that dwarf fold-scheme effects; and (v) seeded reproducible code. The contribution is incremental rather than foundational, but it is the right increment for applied survey prediction and is carefully scoped.
major comments (2)
- [§2.1 Design diagnostics] §2.1 states that a single weighted-LASSO linear predictor supplies the diagnostic ICC for all learners, including random forest, while acknowledging that an additive predictor cannot capture nonlinear cluster interactions a tree could leak. The backstops (learner-independent outcome ICC, which is even lower here, and the direct RF contrast in Table 1) are appropriate for the three null surveys, and Web Appendix D shows RF optimism still tracks the linear-predictor ICC ruler. For the diagnostic to be used as a general pre-analysis tool, however, the operational rule for when those backstops are required needs to be sharper: the text says only that a 'large excess' of outcome ICC over linear-predictor ICC would prompt a learner-native check. Please quantify or at least illustrate what excess would trigger that check, and state explicitly whether the ICC_lp ≳ 0.1 rule is intended to apply u
- [§3.1–3.2 and Discussion] The paper uses two related but distinct decision quantities: the ICC_lp ≳ 0.1 trigger for insisting on design-respecting folds (calibrated where naive bias is still negligible in Figure 1 / Web Appendix D) and the |ΔAUC| < 0.01 practical-equivalence margin used to interpret Table 1. Both are disclosed as simulation-calibrated heuristics rather than formal cutoffs, which is fine, but their relationship for a practitioner is not fully spelled out. In particular, Table 1’s nulls are so far below both thresholds that the empirical claim does not depend on the exact numbers; still, §3.1–3.2 and the Discussion should state more clearly that the 0.1 trigger is a conservative screen for when to bother with design-respecting folds, whereas the 0.01 margin is a post-comparison equivalence criterion, and that neither is claimed to transfer unchanged to other metrics (Brier, calibration slope) or to
minor comments (6)
- [Table 1] Table 1 packs design, DEFF decomposition, ICCs, ΔAUC CIs, and max stratified K into one dense row structure; a short note in the caption that the DEFF split is illustrative (as already stated in §2.1) would reduce the risk that readers treat the product as exact.
- [§2.3 Three surveys] §2.3 NHANES: the argument that pooling cycles would not relax the two-PSUs-per-stratum limit is important for the fallback hierarchy; consider elevating one sentence of that reasoning into the main text rather than leaving it mostly in Web Appendix B.
- [§3.3 / Abstract] §3.3 weight-handling errors are among the most useful practical contributions; a one-line pointer in the abstract or keywords to 'weight-handling pitfalls' would help applied readers find them.
- [Figures 1–2] Figure 1 shaded band (ICC 0.02–0.04) is effective; ensure the Web Appendix D tables that underlie the factorial curves in Figure 2 are cross-referenced by scenario name in the main caption so readers can locate exact numbers without hunting.
- [Conflict of interest] Conflict of interest is properly disclosed (co-authorship of the NHIS reanalysis source). No change needed; noted only for completeness.
- [§2.1] Minor prose: 'DEFFcluster+stratum' and similar inline products are hard to parse; a short notation list or consistent subscripting would help.
Circularity Check
No significant circularity: ICC diagnostic is calibrated on independent simulations with known new-cluster AUC, then applied prospectively/prespecified to real surveys; reanalysis of authors' prior NHIS study is disclosed data reuse, not a load-bearing self-justification.
full rationale
The paper's central claim (within-cluster ICC of outcome and of a preliminary linear predictor, not DEFF, anticipates whether naive vs design-respecting CV differs by a practical amount) is supported by an independent positive-control simulation (finite populations with known true new-cluster AUC; naive optimism rises with ICC_lp while cluster folds stay honest; factorial extension across learners/designs) plus three empirical applications whose scheme differences are measured by paired AUC contrasts, not forced by the diagnostic. The linear-predictor ICC is an in-sample heuristic explicitly calibrated against the simulation ruler (trigger ICC_lp ≳ 0.1 and |ΔAUC| < 0.01 equivalence margin are simulation-derived heuristics, disclosed as such). YRBS used a prespecified decision rule fixed before any computation. The NHIS reanalysis reuses the authors' prior study (disclosed COI) only as a data source; the scheme-null result is not entailed by that paper's conclusions, and NHANES/YRBS are independent. No equation reduces a claimed prediction to a fitted input by construction, no uniqueness theorem is imported from self-citation, and no ansatz is smuggled. Self-citation of the prior NHIS paper and of surveyCV is ordinary data/method reuse, not load-bearing circularity. The derivation chain is therefore self-contained against external benchmarks (simulation truth + public survey data).
Assumptions & free parameters
free parameters (4)
- ICC_lp trigger threshold =
≈ 0.1
- Practical equivalence margin |ΔAUC| =
0.01
- Simulation mixing parameter κ =
grid in {0, 0.1, 0.2, 0.4, 0.6}
- Nested LASSO λ selection rule =
min-CV-error λ
assumptions (6)
- domain assumption Standard K-fold CV assumes exchangeable observations; stratification, clustering, and unequal weights of complex surveys violate that assumption.
- domain assumption The target of internal validation here is generalization to the rest of the finite population beyond sampled PSUs (new-cluster performance).
- domain assumption One-way ANOVA ICC within globally unique strata-by-PSU clusters is an adequate measure of within-cluster homogeneity for the diagnostic.
- ad hoc to paper A single weighted-LASSO linear predictor supplies the diagnostic ICC for all learners, including random forest.
- domain assumption Survey weights enter fitting and evaluation in every scheme; ‘naive’ refers only to fold assignment ignoring design.
- standard math Paired differences with t-based / Monte Carlo CIs and Rubin pooling for MI are valid for scheme comparison; overlapping marginal intervals are not.
invented entities (2)
-
Dual pre-analysis ICC diagnostic (outcome ICC + preliminary linear-predictor ICC) as trigger for survey-aware CV
independent evidence
-
Fallback hierarchy for infeasible stratified Survey CV (max stratified K → PSU-level folds → report not estimable)
independent evidence
Cite this review
Pith. "Pith review of When Does Survey-Aware Cross-Validation Matter? The ICC, Not the Design Effect." pith.science (2026). https://pith.science/paper/JKODKD5F
@misc{pith2026260710634,
author = {Pith},
title = {Pith review of: When Does Survey-Aware Cross-Validation Matter? The ICC, Not the Design Effect},
year = {2026},
howpublished = {\url{https://pith.science/paper/JKODKD5F}},
note = {Machine review of arXiv:2607.10634}
}
read the original abstract
K-fold cross-validation assumes exchangeable observations, violated by the stratification, clustering, and unequal weights of complex sample surveys. Design-respecting "survey CV" exists, but the question of when the extra care changes any conclusion has remained open. We answer it before validating any model, with two inexpensive diagnostics: the within-cluster intraclass correlation (ICC) of the outcome and of a preliminary linear predictor. A simulated positive control demonstrates their sensitivity, with naive cross-validation growing optimistic about new-cluster performance as the ICC rises while cluster-level folds stay honest. We then evaluate paired naive-versus-design-respecting cross-validation in three national health surveys (chronic-pain, diabetes, and adolescent-suicidality prediction; penalized and random-forest learners, plus an unpenalized comparator) - one reanalysis, one prospective application, and one prespecified screen. In all three the diagnostics correctly anticipated the outcome: no scheme difference of practical size, and the only interval excluding zero showed pessimism, not the optimism that cluster leakage produces - even where the design effect was large (which, unlike the ICC, is not the right trigger). We also show the stratified recipe is often infeasible in public-use designs and give a fallback hierarchy, and we document weight-handling errors whose order-of-magnitude artifacts dwarfed any fold-scheme effect. Reproducible code accompanies the paper.
Figures
Reference graph
Works this paper leans on
-
[1]
The Journal of Pain , volume=
The Problem of Pain in the United States: A Population-Based Characterization of Biopsychosocial Correlates of High Impact Chronic Pain Using the National Health Interview Survey , author=. The Journal of Pain , volume=. 2023 , publisher=
2023
-
[2]
Stat , volume=
K-fold cross-validation for complex sample surveys , author=. Stat , volume=. 2022 , publisher=
2022
-
[3]
2022 , note =
surveyCV: Cross validation based on survey design , author =. 2022 , note =
2022
-
[4]
Statistical methods in medical research , volume=
Collaborative-controlled LASSO for constructing propensity score-based estimators in high-dimensional data , author=. Statistical methods in medical research , volume=. 2019 , publisher=
2019
-
[5]
arXiv preprint arXiv:2301.09397 , year=
ddml: Double/debiased machine learning in Stata , author=. arXiv preprint arXiv:2301.09397 , year=
-
[6]
Publicly Available Code , author=
-
[7]
Epidemiology , volume=
Machine learning for causal inference: on the use of cross-fit estimators , author=. Epidemiology , volume=. 2021 , DOI=
2021
-
[8]
arXiv preprint arXiv:1801.09138 , year=
Cross-fitting and fast remainder rates for semiparametric estimation , author=. arXiv preprint arXiv:1801.09138 , year=
Show all 178 references
-
[9]
2018 , volume =
Double/debiased machine learning for treatment and structural parameters , author=. 2018 , volume =. doi:10.1111/ectj.12097 , journal=
2018 doi
-
[10]
Journal of the American Statistical Association , volume=
Semiparametric efficiency in multivariate regression models with missing data , author=. Journal of the American Statistical Association , volume=. 1995 , publisher=
1995
-
[11]
noncollapsibility
Defining, quantifying, and interpreting “noncollapsibility” in epidemiologic studies of measures of “effect” , author=. American Journal of Epidemiology , volume=. 2021 , publisher=
2021
-
[12]
American Journal of Epidemiology , volume=
AIPW: an r package for augmented inverse probability--weighted estimation of average causal effects , author=. American Journal of Epidemiology , volume=. 2021 , publisher=
2021
-
[13]
Nature reviews Disease primers , volume=
Gestational diabetes mellitus , author=. Nature reviews Disease primers , volume=. 2019 , publisher=
2019
-
[14]
Medical care , volume=
Why summary comorbidity measures such as the Charlson comorbidity index and Elixhauser score work , author=. Medical care , volume=. 2015 , publisher=
2015
-
[15]
Metabolic response to stress, can we control it? , author=. Nutrici
-
[16]
Annals of Pharmacotherapy , volume=
Medications associated with weight gain , author=. Annals of Pharmacotherapy , volume=. 2005 , publisher=
2005
-
[17]
Nature Reviews Endocrinology , volume=
Diabetes and its comorbidities—where East meets West , author=. Nature Reviews Endocrinology , volume=. 2013 , publisher=
2013
-
[18]
Current opinion in endocrinology, diabetes, and obesity , volume=
What causes the insulin resistance underlying obesity? , author=. Current opinion in endocrinology, diabetes, and obesity , volume=. 2012 , publisher=
2012
-
[19]
Plos one , volume=
Generating and evaluating a propensity model using textual features from electronic medical records , author=. Plos one , volume=. 2019 , publisher=
2019
-
[20]
Computational statistics & data analysis , volume=
Plasmode simulation for the evaluation of pharmacoepidemiologic methods in complex healthcare databases , author=. Computational statistics & data analysis , volume=. 2014 , publisher=
2014
-
[21]
Biometrika , volume=
Bias reduction of maximum likelihood estimates , author=. Biometrika , volume=. 1993 , publisher=
1993
-
[22]
Biometrics , pages=
Matching using estimated propensity scores: relating theory to practice , author=. Biometrics , pages=. 1996 , publisher=
1996
-
[23]
Annals of internal medicine , volume=
Estimating causal effects from large data sets using propensity scores , author=. Annals of internal medicine , volume=. 1997 , publisher=
1997
-
[24]
American journal of epidemiology , volume=
Variable selection for propensity score models , author=. American journal of epidemiology , volume=. 2006 , publisher=
2006
-
[25]
Journal of clinical epidemiology , volume=
Validating recommendations for coronary angiography following acute myocardial infarction in the elderly: a matched analysis using propensity scores , author=. Journal of clinical epidemiology , volume=. 2001 , publisher=
2001
-
[26]
Statistics in medicine , volume=
Moving towards best practice when using inverse probability of treatment weighting (IPTW) using the propensity score to estimate causal treatment effects in observational studies , author=. Statistics in medicine , volume=. 2015 , publisher=
2015
-
[27]
2013 , pages =
Automated Use of Electronic Health Record Text Data To Improve Validity in Pharmacoepidemiology Studies , booktitle =. 2013 , pages =
2013
-
[28]
Pharmacoepidemiology and Drug Safety , volume=
A comparison of confounder selection and adjustment methods for estimating causal effects using large healthcare databases , author=. Pharmacoepidemiology and Drug Safety , volume=. 2022 , publisher=
2022
-
[29]
International journal of epidemiology , volume=
Evaluating large-scale propensity score performance through real-world and synthetic data experiments , author=. International journal of epidemiology , volume=. 2018 , publisher=
2018
-
[30]
American journal of epidemiology , volume=
Regularized regression versus the high-dimensional propensity score for confounding adjustment in secondary database analyses , author=. American journal of epidemiology , volume=. 2015 , publisher=
2015
-
[31]
Epidemiology (Cambridge, Mass.) , volume=
Targeted maximum likelihood estimation for pharmacoepidemiologic research , author=. Epidemiology (Cambridge, Mass.) , volume=. 2016 , publisher=
2016
-
[32]
Statistical methods in medical research , volume=
Diagnosing and responding to violations in the positivity assumption , author=. Statistical methods in medical research , volume=. 2012 , publisher=
2012
-
[33]
Journal of the Royal Statistical Society Series A: Statistics in Society , volume=
Prediction of default probability by using statistical models for rare events , author=. Journal of the Royal Statistical Society Series A: Statistics in Society , volume=. 2019 , publisher=
2019
-
[34]
Pharmacoepidemiology and Drug Safety , volume=
Implementing high-dimensional propensity score principles to improve confounder adjustment in UK electronic health records , author=. Pharmacoepidemiology and Drug Safety , volume=. 2020 , publisher=
2020
-
[35]
American Journal of Neuroradiology , volume=
ICD-10: history and context , author=. American Journal of Neuroradiology , volume=. 2016 , publisher=
2016
-
[36]
Pharmacoepidemiology and Drug Safety , volume=
Novel methods for pregnancy drug safety surveillance in the FDA Sentinel System , author=. Pharmacoepidemiology and Drug Safety , volume=. 2023 , publisher=
2023
-
[37]
Expert Review of Clinical Pharmacology , volume=
Abiraterone acetate versus docetaxel for metastatic castration-resistant prostate cancer: a cohort study within the French nationwide claims database , author=. Expert Review of Clinical Pharmacology , volume=. 2022 , publisher=
2022
-
[38]
The international journal of biostatistics , volume=
Effect estimation in point-exposure studies with binary outcomes and high-dimensional covariate data--a comparison of targeted maximum likelihood estimation and inverse probability of treatment weighting , author=. The international journal of biostatistics , volume=. 2016 , p...
2016
-
[39]
American journal of hypertension , volume=
Abdominal obesity, body mass index, and hypertension in US adults: NHANES 2007--2010 , author=. American journal of hypertension , volume=. 2012 , publisher=
2007
-
[40]
Diabetes care , volume=
Joint effects of obesity and vitamin D insufficiency on insulin resistance and type 2 diabetes: results from the NHANES 2001--2006 , author=. Diabetes care , volume=. 2012 , publisher=
2001
-
[41]
International journal of endocrinology , volume=
The association of sleep disorder, obesity status, and diabetes mellitus among US adults—The NHANES 2009-2010 survey results , author=. International journal of endocrinology , volume=. 2013 , publisher=
2009
-
[42]
Obesity , volume=
Trends in cardiovascular disease risk factors by obesity level in adults in the United States, NHANES 1999-2010 , author=. Obesity , volume=. 2014 , publisher=
1999
-
[43]
Journal of Applied Statistics , volume=
Propensity score prediction for electronic healthcare databases using super learner and high-dimensional propensity score methods , author=. Journal of Applied Statistics , volume=. 2019 , publisher=
2019
-
[44]
Statistical methods in medical research , volume=
Scalable collaborative targeted learning for high-dimensional data , author=. Statistical methods in medical research , volume=. 2019 , publisher=
2019
-
[45]
Statistics in medicine , volume=
High-dimensional propensity score algorithm in comparative effectiveness research with time-varying interventions , author=. Statistics in medicine , volume=. 2015 , publisher=
2015
-
[46]
Statistics in medicine , volume=
Variable selection and raking in propensity scoring , author=. Statistics in medicine , volume=. 2007 , publisher=
2007
-
[47]
Journal of clinical epidemiology , volume=
Propensity score model overfitting led to inflated variance of estimated odds ratios , author=. Journal of clinical epidemiology , volume=. 2016 , publisher=
2016
-
[48]
Journal of Neurology, Neurosurgery, and Psychiatry , volume=
Comparison of fingolimod, dimethyl fumarate and teriflunomide for multiple sclerosis: when methodology does not hold the promise , author=. Journal of Neurology, Neurosurgery, and Psychiatry , volume=. 2019 , url=
2019
-
[49]
Political analysis , volume=
Matching as nonparametric preprocessing for reducing model dependence in parametric causal inference , author=. Political analysis , volume=. 2007 , publisher=
2007
-
[50]
Health Services and Outcomes Research Methodology , volume=
Using propensity scores to help design observational studies: application to the tobacco litigation , author=. Health Services and Outcomes Research Methodology , volume=. 2001 , publisher=
2001
-
[51]
Pharmacoepidemiology and Drug Safety , volume=
Transparency of high-dimensional propensity score analyses: guidance for diagnostics and reporting , author=. Pharmacoepidemiology and Drug Safety , volume=. 2022 , publisher=
2022
-
[52]
Pharmacoepidemiology and drug safety , volume=
Instrumental variable methods in comparative safety and effectiveness research , author=. Pharmacoepidemiology and drug safety , volume=. 2010 , publisher=
2010
-
[53]
Pharmacoepidemiology and drug safety , volume=
The implications of propensity score variable selection strategies in pharmacoepidemiology: an empirical illustration , author=. Pharmacoepidemiology and drug safety , volume=. 2011 , publisher=
2011
-
[54]
American journal of epidemiology , volume=
Effects of adjusting for instrumental variables on bias and precision of effect estimates , author=. American journal of epidemiology , volume=. 2011 , publisher=
2011
-
[55]
Pharmacoepidemiology and Drug Safety , volume=
High-dimensional propensity scores for empirical covariate selection in secondary database studies: Planning, implementation, and reporting , author=. Pharmacoepidemiology and Drug Safety , volume=. 2023 , publisher=
2023
-
[56]
Journal of statistical software , volume=
Regularization paths for generalized linear models via coordinate descent , author=. Journal of statistical software , volume=. 2010 , publisher=
2010
-
[57]
Journal of Statistical Software , year =
Susan Gruber and Mark J. Journal of Statistical Software , year =
-
[58]
The Stata Journal , volume=
hdps: A suite of commands for applying high-dimensional propensity-score approaches , author=. The Stata Journal , volume=. 2023 , publisher=
2023
-
[59]
Epidemiology , volume=
Using super learner prediction modeling to improve high-dimensional propensity score estimation , author=. Epidemiology , volume=. 2018 , publisher=
2018
-
[60]
American Journal of Epidemiology , volume=
Demystifying statistical inference when using machine learning in causal research , author=. American Journal of Epidemiology , volume=. 2023 , publisher=
2023
-
[61]
The annals of statistics , volume=
Multivariate adaptive regression splines , author=. The annals of statistics , volume=. 1991 , publisher=
1991
-
[62]
American Journal of Epidemiology , volume=
Challenges in obtaining valid causal effect estimates with machine learning algorithms , author=. American Journal of Epidemiology , volume=. 2023 , publisher=
2023
-
[63]
American journal of epidemiology , volume=
Mortality risk score prediction in an elderly population using machine learning , author=. American journal of epidemiology , volume=. 2013 , publisher=
2013
-
[64]
International Journal of Epidemiology , volume=
Practical considerations for specifying a super learner , author=. International Journal of Epidemiology , volume=. 2023 , publisher=
2023
-
[65]
American journal of epidemiology , volume=
Improving propensity score estimators' robustness to model misspecification using super learner , author=. American journal of epidemiology , volume=. 2015 , publisher=
2015
-
[66]
Pharmacoepidemiology and Drug Safety , volume=
Machine learning for improving high-dimensional proxy confounder adjustment in healthcare database studies: An overview of the current literature , author=. Pharmacoepidemiology and Drug Safety , volume=. 2022 , publisher=
2022
-
[68]
Studies in Nonlinear Dynamics & Econometrics , volume=
Revisiting the statistical specification of near-multicollinearity in the logistic regression model , author=. Studies in Nonlinear Dynamics & Econometrics , volume=. 2016 , publisher=
2016
-
[69]
2023 , howpublished=
Crossfit: An R package to apply sample splitting (cross-fit) to AIPW and TMLE in causal inference , author=. 2023 , howpublished=
2023
-
[70]
2023 , month =
Karim, ME , title =. 2023 , month =. doi:10.5281/zenodo.7877767 , howpublished=
2023 doi
-
[71]
Epidemiology , volume=
Can we train machine learning methods to outperform the high-dimensional propensity score algorithm? , author=. Epidemiology , volume=. 2018 , publisher=
2018
-
[72]
Cell metabolism , volume=
Why does obesity cause diabetes? , author=. Cell metabolism , volume=. 2022 , publisher=
2022
-
[73]
Journal of clinical epidemiology , volume=
A chronic disease score from automated pharmacy data , author=. Journal of clinical epidemiology , volume=. 1992 , publisher=
1992
-
[74]
Medical care , pages=
Comorbidity measures for use with administrative data , author=. Medical care , pages=. 1998 , publisher=
1998
-
[75]
Journal of chronic diseases , volume=
A new method of classifying prognostic comorbidity in longitudinal studies: development and validation , author=. Journal of chronic diseases , volume=. 1987 , publisher=
1987
-
[76]
American Journal of Managed Care , volume=
A comparison of comorbidity measurements to predict healthcare expenditures , author=. American Journal of Managed Care , volume=. 2006 , publisher=
2006
-
[77]
Diabetologia , volume=
Risk factors for diabetes mellitus by age and sex: results of the National Population Health Survey , author=. Diabetologia , volume=. 2001 , publisher=
2001
-
[78]
Epidemiology (Cambridge, Mass.) , volume=
High-dimensional propensity score adjustment in studies of treatment effects using health care claims data , author=. Epidemiology (Cambridge, Mass.) , volume=. 2009 , publisher=
2009
-
[79]
Epidemiology , volume=
Erratum: high-dimensional propensity score adjustment in studies of treatment effects using health care claims data , author=. Epidemiology , volume=. 2018 , publisher=
2018
-
[80]
Pharmacoepidemiology and drug safety , volume=
Sensitivity analysis and external adjustment for unmeasured confounders in epidemiologic database studies of therapeutics , author=. Pharmacoepidemiology and drug safety , volume=. 2006 , publisher=
2006
-
[81]
Journal of clinical epidemiology , volume=
Prognostic score--based balance measures can be a useful diagnostic for propensity score methods in comparative effectiveness research , author=. Journal of clinical epidemiology , volume=. 2013 , publisher=
2013
-
[82]
Journal of chronic diseases , volume=
Spurious effects from an extraneous variable , author=. Journal of chronic diseases , volume=. 1966 , publisher=
1966
-
[83]
Health services research , volume=
The potential of high-dimensional propensity scores in health services research: an exemplary study on the quality of Care for Elective Percutaneous Coronary Interventions , author=. Health services research , volume=. 2018 , publisher=
2018
-
[84]
Journal of Statistical Software , volume=
Various versatile variances: an object-oriented implementation of clustered covariances in R , author=. Journal of Statistical Software , volume=
-
[85]
Causal Inference: What If , publisher =
Hern. Causal Inference: What If , publisher =. 2023 , edition =. doi:10.1201/9781315374932 , subject =
2023 doi
-
[86]
International journal of obesity , volume=
Does obesity shorten life? The importance of well-defined interventions to answer causal questions , author=. International journal of obesity , volume=. 2008 , publisher=
2008
-
[87]
2022 , howpublished=
WeightIt: Weighting for Covariate Balance in Observational Studies , author =. 2022 , howpublished=
2022
-
[88]
2023 , note =
Pharmacoepidemiology Toolbox , author =. 2023 , note =
2023
-
[89]
, howpublished=
Karim, ME. , howpublished=. 2023 , note =
2023
-
[90]
2023 , howpublished=
SuperLearner: Super Learner Prediction , author =. 2023 , howpublished=
2023
-
[91]
2020 , howpublished=
autoCovariateSelection: Automatic Covariate Selection , author =. 2020 , howpublished=
2020
-
[92]
Pharmacoepidemiology and drug safety , volume=
On the role of marginal confounder prevalence--implications for the high-dimensional propensity score algorithm , author=. Pharmacoepidemiology and drug safety , volume=. 2015 , publisher=
2015
-
[93]
Pharmacoepidemiology and drug safety , volume=
Quantifying bias reduction with fixed-duration versus all-available covariate assessment periods , author=. Pharmacoepidemiology and drug safety , volume=. 2019 , publisher=
2019
-
[94]
Journal of comparative effectiveness research , volume=
Comparing high-dimensional confounder control methods for rapid cohort studies from electronic health records , author=. Journal of comparative effectiveness research , volume=. 2016 , publisher=
2016
-
[95]
Epidemiology , volume=
Deep learning-based propensity scores for confounding control in comparative effectiveness research: A large-scale, real-world data study , author=. Epidemiology , volume=. 2021 , publisher=
2021
-
[96]
American journal of epidemiology , volume=
Estimating risk ratios and risk differences using regression , author=. American journal of epidemiology , volume=. 2020 , publisher=
2020
-
[97]
Multiple Sclerosis Journal , volume=
Recommendations for the use of propensity score methods in multiple sclerosis research , author=. Multiple Sclerosis Journal , volume=. 2022 , publisher=
2022
-
[98]
Multiple Sclerosis Journal , volume=
The use and quality of reporting of propensity score methods in multiple sclerosis literature: A review , author=. Multiple Sclerosis Journal , volume=. 2022 , publisher=
2022
-
[99]
European journal of epidemiology , volume=
Stacked generalization: an introduction to super learning , author=. European journal of epidemiology , volume=. 2018 , publisher=
2018
-
[100]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
Regularization and variable selection via the elastic net , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2005 , publisher=
2005
-
[101]
Epidemiology , volume=
Variable selection for confounding adjustment in high-dimensional covariate spaces when analyzing healthcare databases , author=. Epidemiology , volume=. 2017 , publisher=
2017
-
[102]
Biometrics , volume=
Outcome-adaptive lasso: variable selection for causal inference , author=. Biometrics , volume=. 2017 , publisher=
2017
-
[103]
Biometrics , volume=
Robust inference on the average treatment effect using the outcome highly adaptive lasso , author=. Biometrics , volume=. 2020 , publisher=
2020
-
[104]
Clinical epidemiology , pages=
Automated data-adaptive analytics for electronic healthcare data to study causal treatment effects , author=. Clinical epidemiology , pages=. 2018 , publisher=
2018
-
[105]
Biometrics , volume=
Estimating average treatment effects with a double-index propensity score , author=. Biometrics , volume=. 2020 , publisher=
2020
-
[106]
International journal of epidemiology , volume=
Use of comorbidity scores for control of confounding in studies using administrative databases , author=. International journal of epidemiology , volume=. 2000 , publisher=
2000
-
[107]
European journal of epidemiology , volume=
Principles of confounder selection , author=. European journal of epidemiology , volume=. 2019 , publisher=
2019
-
[108]
BMC Health Services Research , volume=
Predictive performance of comorbidity measures in administrative databases for diabetes cohorts , author=. BMC Health Services Research , volume=. 2013 , publisher=
2013
-
[109]
Osteoporosis international , volume=
Performance of comorbidity measures for predicting outcomes in population-based osteoporosis cohorts , author=. Osteoporosis international , volume=. 2011 , publisher=
2011
-
[110]
Epidemiology , pages=
Causal diagrams for epidemiologic research , author=. Epidemiology , pages=. 1999 , publisher=
1999
-
[111]
National Center for Health Statistics , year=
National Health and Nutrition Examination Survey (NHANES) , author=. National Center for Health Statistics , year=
-
[112]
Statistics in medicine , volume=
Propensity score weighting with multilevel data , author=. Statistics in medicine , volume=. 2013 , publisher=
2013
-
[113]
Annals of Epidemiology , volume=
The association between diabetes and excessive daytime sleepiness among American adults aged 20--79 years: findings from the 2015--2018 National Health and Nutrition Examination Surveys , author=. Annals of Epidemiology , volume=. 2022 , publisher=
2015
-
[114]
Health services research , volume=
Generalizing observational study results: applying propensity score methods to complex surveys , author=. Health services research , volume=. 2014 , publisher=
2014
-
[115]
Biometrika , volume=
The central role of the propensity score in observational studies for causal effects , author=. Biometrika , volume=. 1983 , publisher=
1983
-
[116]
Statistical methods in medical research , volume=
Propensity score matching and complex surveys , author=. Statistical methods in medical research , volume=. 2018 , publisher=
2018
-
[117]
Statistics in medicine , volume=
Propensity score weighting for covariate adjustment in randomized clinical trials , author=. Statistics in medicine , volume=. 2021 , publisher=
2021
-
[118]
Statistical methods in medical research , volume=
Subgroup balancing propensity score , author=. Statistical methods in medical research , volume=. 2020 , publisher=
2020
-
[119]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
Sensitivity analysis for inverse probability weighting estimators via the percentile bootstrap , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2019 , publisher=
2019
-
[120]
Journal of the American Statistical Association , volume=
Sharp sensitivity analysis for inverse propensity weighting via quantile balancing , author=. Journal of the American Statistical Association , volume=. 2023 , publisher=
2023
-
[121]
Kish, Leslie , title =
-
[122]
Journal of Official Statistics , volume =
Kish, Leslie , title =. Journal of Official Statistics , volume =
-
[123]
and McNeil, Barbara J
Hanley, James A. and McNeil, Barbara J. , title =. Radiology , volume =
-
[124]
, title =
Rubin, Donald B. , title =
-
[125]
, title =
Brier, Glenn W. , title =. Monthly Weather Review , volume =
-
[126]
Journal of Statistical Software , volume =
Lumley, Thomas , title =. Journal of Statistical Software , volume =
-
[127]
Lumley, Thomas , title =
-
[128]
Stat , volume =
Iparragirre, Amaia and Lumley, Thomas and Barrio, Irantzu and Arostegui, Inmaculada , title =. Stat , volume =
-
[129]
Stat , volume =
Iparragirre, Amaia and Barrio, Irantzu and Arostegui, Inmaculada , title =. Stat , volume =
-
[130]
Canadian Journal of Statistics , volume =
Holbrook, Andrew and Lumley, Thomas and Gillen, Daniel , title =. Canadian Journal of Statistics , volume =
-
[131]
Journal of Survey Statistics and Methodology , volume =
Lumley, Thomas and Scott, Alastair , title =. Journal of Survey Statistics and Methodology , volume =
-
[132]
and Kording, Konrad P
Saeb, Sohrab and Lonini, Luca and Jayaraman, Arun and Mohr, David C. and Kording, Konrad P. , title =. GigaScience , volume =
-
[133]
ACM Transactions on Knowledge Discovery from Data , volume =
Kaufman, Shachar and Rosset, Saharon and Perlich, Claudia and Stitelman, Ori , title =. ACM Transactions on Knowledge Discovery from Data , volume =
-
[134]
Journal of the American Statistical Association , volume =
Rabinowicz, Assaf and Rosset, Saharon , title =. Journal of the American Statistical Association , volume =
-
[135]
and van Smeden, Maarten and Wynants, Laure and Steyerberg, Ewout W
Van Calster, Ben and McLernon, David J. and van Smeden, Maarten and Wynants, Laure and Steyerberg, Ewout W. , title =. BMC Medicine , volume =
-
[136]
and Vickers, Andrew J
Steyerberg, Ewout W. and Vickers, Andrew J. and Cook, Nancy R. and Gerds, Thomas and Gonen, Mithat and Obuchowski, Nancy and Pencina, Michael J. and Kattan, Michael W. , title =. Epidemiology , volume =
-
[137]
and Reitsma, Johannes B
Collins, Gary S. and Reitsma, Johannes B. and Altman, Douglas G. and Moons, Karel G. M. , title =. BMC Medicine , volume =
-
[138]
Machine Learning , volume =
Breiman, Leo , title =. Machine Learning , volume =
-
[139]
Journal of the Royal Statistical Society: Series B , volume =
Tibshirani, Robert , title =. Journal of the Royal Statistical Society: Series B , volume =
-
[140]
Journal of Statistical Software , volume =
van Buuren, Stef and Groothuis-Oudshoorn, Karin , title =. Journal of Statistical Software , volume =
-
[141]
and Ziegler, Andreas , title =
Wright, Marvin N. and Ziegler, Andreas , title =. Journal of Statistical Software , volume =
-
[142]
MMWR Morbidity and Mortality Weekly Report , volume =
Dahlhamer, James and Lucas, Jacqueline and Zelaya, Carla and Nahin, Richard and Mackey, Sean and DeBar, Lynn and Kerns, Robert and Von Korff, Michael and Porter, Linda and Helmick, Charles , title =. MMWR Morbidity and Mortality Weekly Report , volume =
-
[143]
, title =
Dinh, An and Miertschin, Stacey and Young, Amber and Mohanty, Somya D. , title =. BMC Medical Informatics and Decision Making , volume =
-
[144]
Statistical Science , volume =
Skinner, Chris and Wakefield, Jon , title =. Statistical Science , volume =
-
[145]
Journal of the American Statistical Association , volume =
Dagdoug, Mehdi and Goga, Camelia and Haziza, David , title =. Journal of the American Statistical Association , volume =
-
[146]
and Hedberg, E
Hedges, Larry V. and Hedberg, E. C. , title =. Educational Evaluation and Policy Analysis , volume =
-
[147]
and Ukoumunne, Obioha C
Gulliford, Martin C. and Ukoumunne, Obioha C. and Chinn, Susan , title =. American Journal of Epidemiology , volume =
-
[148]
and Ukoumunne, Obioha C
Adams, Gail and Gulliford, Martin C. and Ukoumunne, Obioha C. and Eldridge, Sandra and Chinn, Susan and Campbell, Michael J. , title =. Journal of Clinical Epidemiology , volume =
-
[149]
Dwivedi, Laxmi Kant and Mahapatra, Bidhubhusan and Bansal, Anjali and Gupta, Jitendra and Singh, Abhishek and Roy, T. K. , title =. SSM - Population Health , volume =
-
[150]
Journal of Medical Internet Research , volume =
Kim, Hyejun and Son, Yejun and Lee, Hojae and Kang, Jiseung and Hammoodi, Ahmed and Choi, Yujin and Kim, Hyeon Jin and Lee, Hayeon and Fond, Guillaume and Boyer, Laurent and Kwon, Rosie and Woo, Selin and Yon, Dong Keon , title =. Journal of Medical Internet Research , volume =
-
[151]
Clinical Psychopharmacology and Neuroscience , volume =
Lim, Jae Seok and Yang, Chan-Mo and Baek, Ju-Won and Lee, Sang-Yeol and Kim, Bung-Nyun , title =. Clinical Psychopharmacology and Neuroscience , volume =
-
[152]
C., Ukoumunne, O
Adams, G., Gulliford, M. C., Ukoumunne, O. C., Eldridge, S., Chinn, S., and Campbell, M. J. (2004). Patterns of intra-cluster correlation from primary care research to inform study design and analysis. Journal of Clinical Epidemiology , 57(8):785--794
2004
-
[153]
Breiman, L. (2001). Random forests. Machine Learning , 45(1):5--32
2001
-
[154]
S., Reitsma, J
Collins, G. S., Reitsma, J. B., Altman, D. G., and Moons, K. G. M. (2015). Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis ( TRIPOD ): the TRIPOD statement. BMC Medicine , 13:1
2015
-
[155]
Dahlhamer, J., Lucas, J., Zelaya, C., Nahin, R., Mackey, S., DeBar, L., Kerns, R., Von Korff, M., Porter, L., and Helmick, C. (2018). Prevalence of chronic pain and high-impact chronic pain among adults --- U nited S tates, 2016. MMWR Morbidity and Mortality Weekly Report , 67...
2018
-
[156]
Dinh, A., Miertschin, S., Young, A., and Mohanty, S. D. (2019). A data-driven approach to predicting diabetes and cardiovascular disease with machine learning. BMC Medical Informatics and Decision Making , 19:211
2019
-
[157]
K., Mahapatra, B., Bansal, A., Gupta, J., Singh, A., and Roy, T
Dwivedi, L. K., Mahapatra, B., Bansal, A., Gupta, J., Singh, A., and Roy, T. K. (2022). Intra-cluster correlations in socio-demographic variables and their implications: an analysis based on large-scale surveys in I ndia. SSM - Population Health , 21:101317
2022
-
[158]
B., Weber, II, K
Falasinnu, T., Hossain, M. B., Weber, II, K. A., Helmick, C. G., Karim, M. E., and Mackey, S. (2023). The problem of pain in the united states: A population-based characterization of biopsychosocial correlates of high impact chronic pain using the national health interview sur...
2023
-
[159]
Friedman, J., Hastie, T., and Tibshirani, R. (2010). Regularization paths for generalized linear models via coordinate descent. Journal of statistical software , 33(1):1
2010
-
[160]
Guerin, C., McMahon, T., and Wieczorek, J. (2022). surveyCV: Cross validation based on survey design . R package version 0.1.1
2022
-
[161]
C., Ukoumunne, O
Gulliford, M. C., Ukoumunne, O. C., and Chinn, S. (1999). Components of variance and intraclass correlations for the design of community-based surveys and intervention studies: data from the H ealth S urvey for E ngland 1994. American Journal of Epidemiology , 149(9):876--883
1999
-
[162]
Hedges, L. V. and Hedberg, E. C. (2007). Intraclass correlation values for planning group-randomized trials in education. Educational Evaluation and Policy Analysis , 29(1):60--87
2007
-
[163]
Holbrook, A., Lumley, T., and Gillen, D. (2020). Estimating prediction error for complex samples. Canadian Journal of Statistics , 48(2):204--221
2020
-
[164]
Iparragirre, A., Barrio, I., and Arostegui, I. (2023a). Estimation of the ROC curve and the area under it with complex survey data. Stat , 12(1):e635
-
[165]
Iparragirre, A., Lumley, T., Barrio, I., and Arostegui, I. (2023b). Variable selection with LASSO regression for complex survey data. Stat , 12(1):e578
-
[166]
Kaufman, S., Rosset, S., Perlich, C., and Stitelman, O. (2012). Leakage in data mining: formulation, detection, and avoidance. ACM Transactions on Knowledge Discovery from Data , 6(4):1--21
2012
-
[167]
J., Lee, H., Fond, G., Boyer, L., Kwon, R., Woo, S., and Yon, D
Kim, H., Son, Y., Lee, H., Kang, J., Hammoodi, A., Choi, Y., Kim, H. J., Lee, H., Fond, G., Boyer, L., Kwon, R., Woo, S., and Yon, D. K. (2024). Machine learning--based prediction of suicidal thinking in adolescents by derivation and validation in 3 independent worldwide cohor...
2024
-
[168]
Kish, L. (1965). Survey Sampling . John Wiley & Sons, New York
1965
-
[169]
Kish, L. (1992). Weighting for unequal P_i . Journal of Official Statistics , 8(2):183--200
1992
-
[170]
S., Yang, C.-M., Baek, J.-W., Lee, S.-Y., and Kim, B.-N
Lim, J. S., Yang, C.-M., Baek, J.-W., Lee, S.-Y., and Kim, B.-N. (2022). Prediction models for suicide attempts among adolescents using machine learning techniques. Clinical Psychopharmacology and Neuroscience , 20(4):609--620
2022
-
[171]
and Scott, A
Lumley, T. and Scott, A. (2015). AIC and BIC for modeling with complex survey data. Journal of Survey Statistics and Methodology , 3(1):1--18
2015
-
[172]
and Rosset, S
Rabinowicz, A. and Rosset, S. (2022). Cross-validation for correlated data. Journal of the American Statistical Association , 117(538):718--731
2022
-
[173]
Rubin, D. B. (1987). Multiple Imputation for Nonresponse in Surveys . John Wiley & Sons, New York
1987
-
[174]
C., and Kording, K
Saeb, S., Lonini, L., Jayaraman, A., Mohr, D. C., and Kording, K. P. (2017). The need to approximate the use-case in clinical machine learning. GigaScience , 6(5):gix019
2017
-
[175]
and Wakefield, J
Skinner, C. and Wakefield, J. (2017). Introduction to the design and analysis of complex survey data. Statistical Science , 32(2):165--175
2017
-
[176]
Tibshirani, R. (1996). Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society: Series B , 58(1):267--288
1996
-
[177]
and Groothuis-Oudshoorn, K
van Buuren, S. and Groothuis-Oudshoorn, K. (2011). mice: Multivariate imputation by chained equations in R . Journal of Statistical Software , 45(3):1--67
2011
-
[178]
Wieczorek, J., Guerin, C., and McMahon, T. (2022). K-fold cross-validation for complex sample surveys. Stat , 11(1):e454
2022
-
[179]
Wright, M. N. and Ziegler, A. (2017). ranger: a fast implementation of random forests for high dimensional data in C++ and R . Journal of Statistical Software , 77(1):1--17
2017
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.