{"id":"74b4add7-cce4-4eae-8329-1269e0b20c9c","arxiv_id":"2607.07065","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":5,"one_line_summary":"Under a null-effect plasmode simulation, outcome-adaptive LASSO (IPTW), GLiDeR, and HAL-TMLE were best calibrated across frequent, rare-exposure, and rare-outcome scenarios, while LASSO-IPTW was biased under rare exposure and HAL G-computation over-covered.","lead":"This paper compares ten statistical pipelines for adjusting confounding in health-data studies with many proxy variables, finding that outcome-aware selection and doubly robust estimation give the best-calibrated results when exposures or outcomes are rare. A smart generalist might read it to learn which modern variable-selection methods to trust—and at what computational cost—when analyzing large observational healthcare databases.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The coverage ranking identifying OAL and GLiDeR as 'best calibrated' is confounded by inference method: bootstrap-based methods systematically over-perform influence-curve-based methods on coverage, so the calibration gap may reflect error-source coverage rather than selection strategy.","rationale":"The reader correctly identifies PS misspecification as a conditioning limitation, but that concern affects OAL (IPTW) specifically while leaving the doubly robust methods (HAL-TMLE, GLiDeR) less affected. The inference-method confounding I identify is more load-bearing because it systematically affects the entire coverage comparison: every bootstrap-based method achieves near-nominal coverage, every influence-curve TMLE method under-covers. This means the 'best calibrated' ranking may be an artifact of which uncertainty sources each SE captures rather than a property of the selection strategy. However, this does not change the verdict from CONDITIONAL for three reasons. First, the authors explicitly acknowledge the confounding in §4.3, so it is a known limitation rather than an overlooked flaw. Second, the bias and robustness findings — LASSO-IPTW's rare-exposure bias and TMLE's correction of it — are independent of the inference method and remain solid. Third, HAL (TMLE) demonstrates that influence-curve methods can achieve near-nominal coverage, suggesting the selection strategy does matter even if inference method is a strong modifier. The CONDITIONAL verdict already captures the idea that conclusions are bounded by design limitations; adding the inference-method confounding as another acknowledged limitation does not move the verdict further. The concrete test (bootstrap LASSO-TMLE) is feasible with the released code and would directly determine whether the calibration ranking is robust to the inference method or is an artifact of it.","tokens_in":13776,"tokens_out":6680,"duration_ms":247661,"concrete_test":"Re-run LASSO (TMLE, dev) in the rare-exposure scenario with 200-replicate percentile bootstrap inference instead of the influence-curve SE. If coverage rises from 91.8% to near 95%, the calibration ranking is driven by inference method rather than selection strategy, and Table 5 should foreground inference choice alongside selection strategy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that OAL (IPTW) and GLiDeR are 'best calibrated' (coverage 94–95%) while LASSO-TMLE, C-TMLE, and Standard TMLE 'under-cover modestly' (91–93%) is confounded by the inference method. OAL and GLiDeR use 200-replicate bootstrap SEs that propagate selection uncertainty through the full pipeline, while every under-covering TMLE method uses efficient influence-curve SEs that condition on the cross-validated nuisance model (§4.3). The authors acknowledge this ('the calibration and runtime comparisons are therefore partly confounded by which error sources each SE captures'), but Table 5's practical recommendation treats calibration as a property of the selection/estimation strategy rather than the inference method. The systematic pattern is striking: all bootstrap-based methods (OAL, GLiDeR, HAL G-Comp) achieve near-nominal or conservative coverage, while all influence-curve TMLE methods under-cover. HAL (TMLE) — which uses influence curves yet achieves near-nominal coverage — partially mitigates this concern, but it is the exception, not the rule. If LASSO-TMLE with bootstrap inference also achieves coverage near 95%, the practical recommendation shifts from 'use OAL/GLiDeR selection' to 'use bootstrap inference with any regularized selection,' which would substantially weaken the selection-strategy recommendation. The reader's PS-misspecification concern is valid but affects only OAL (not the doubly robust methods); the inference-method confounding affects the entire coverage ranking and is therefore more load-bearing for the 'best calibrated' claim.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This manuscript compares ten regularized confounding-adjustment pipelines—spanning traditional LASSO, collaborative-controlled LASSO, outcome-adaptive LASSO (OAL), group LASSO/GLiDeR, and highly adaptive LASSO (HAL), each paired with IPTW and/or TMLE—under a plasmode simulation anchored on NHANES 2013–2018 data. Three scenarios (frequent, rare-exposure, rare-outcome) are evaluated under a known null (RD = 0), with bias, coverage, relative error, and runtime reported. The central finding is that outcome-aware selection (OAL, GLiDeR) and doubly robust estimation (HAL-TMLE) are best calibrated, while LASSO-IPTW shows large bias under rare exposure and HAL G-computation over-covers with large relative error. The design is well-structured (ADEMP framework, 1000 replicates, common-replicate comparison, null truth anchor), and the runtime accounting is a valuable practical contribution.","tokens_in":14565,"tokens_out":1020,"duration_ms":154875,"significance":"The paper addresses a genuine gap: no prior single-design study has assembled this breadth of regularized selection families crossed with IPTW versus doubly robust estimation while varying exposure prevalence and reporting compute cost. The null-anchored plasmode design is a strength—it isolates residual-confounding bias cleanly. The runtime table (Table 4) and practical guidance table (Table 5) are useful for practitioners. The reproducible code repository and bundled anchor data are commendable. The finding that replacing IPTW with TMLE removes rare-exposure bias and instability is practically important and well-demonstrated.","major_comments":[{"comment":"§4.3 and Table 5: The calibration ranking identifying OAL (IPTW) and GLiDeR as 'best calibrated' (coverage 94–95%) while LASSO-TMLE, C-TMLE, and Standard TMLE 'under-cover modestly' (91–93%) is confounded by the inference method. OAL and GLiDeR use 200-replicate bootstrap SEs that propagate selection uncertainty through the full pipeline, while every under-covering TMLE method uses efficient influence-curve SEs that condition on the cross-validated nuisance model. The authors acknowledge this in §4.3 ('the calibration and runtime comparisons are therefore partly confounded by which error sources each SE captures'), but Table 5's practical recommendation treats calibration as a property of the selection/estimation strategy rather than the inference method. The systematic pattern is striking: all bootstrap-based methods (OAL, GLiDeR, HAL G-Comp) achieve near-nominal or conservative, while ","section":null}],"minor_comments":[{"comment":"§2.5, Table 2: The realized exposure prevalence in the rare-exposure scenario is 9.4%, which is moderate rather than truly rare by pharmacoepidemiology standards. The authors should clarify how this maps to its intended severity.","section":null},{"comment":"§2.3, Table 1: The traditional LASSO pipelines enforce inclusion of all investigator-specified covariates, whereas the other methods do not. This difference is mentioned only in the limitations (§4.5) and should be noted in the methods.","section":null},{"comment":"§3.1, Figure 1: The caption notes that interval widths are not strictly comparable across methods due to different inference types. This is important context that should be stated more prominently.","section":null},{"comment":"§4.5: The statement that higher HAL degrees were 'computationally infeasible' could quantify the attempted max_degree values and approximate runtime.","section":null},{"comment":"Table 3: The Monte Carlo error is stated as ~0.7 pp for coverage near 95%, but the bolding threshold is |Δ| ≤ 1.5 pp. Clarify why 1.5 rather than ~1.4 pp was chosen.","section":null},{"comment":"§4.2: The comparison with prior work is thorough but could note that the Franklin et al. (2017) plasmode also used a null, making the design choice less novel than implied.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The inference-method confounding is the most substantive concern. The skeptic's point is valid: if LASSO-TMLE with bootstrap inference also achieves near-95% coverage, the practical recommendation shifts from 'use OAL/GLiDeR selection' to 'use bootstrap inference with any regularized selection,' which would substantially weaken the selection-strategy recommendation. This is addressable either by running LASSO-TMLE with bootstrap (computationally expensive but feasible for a subset of replicates) or by reframing Table 5 to separate selection-strategy effects from inference-method effects. The PS-misspecification concern is secondary—it affects the generalizability of OAL/GLiDeR calibration claims but not the doubly robust methods, and the authors acknowledge it in §4.5."},"author_rebuttal":null,"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper assembles five regularized selection families (LASSO, collaborative-controlled LASSO, OAL, GLiDeR, HAL) crossed with IPTW/TMLE under one plasmode design with a known null, a rare-exposure arm, and runtime accounting. No prior study holds all these axes at once. That breadth plus the honest compute reporting is the real contribution here, and the public code is a plus. The finding that LASSO-IPTW breaks down under rare exposure while TMLE fixes it is clean and useful. The runtime table—spanning sub-second to 16+ hours—is genuinely informative for practitioners. The plasmode design is well-structured: ADEMP-compliant, 1000 replicates, common-replicate comparison, engineered confounding under a null so any nonzero estimate is measurable bias. The real soft spot is the coverage ranking. The stress-test note lands: OAL and GLiDeR use bootstrap SEs that propagate selection uncertainty, while the under-covering TMLE methods use influence-curve SEs that condition on the cross-validated nuisance model. The authors acknowledge this confounding in Section 4.3, but Table 5's practical recommendations still treat calibration as a property of the selection strategy rather than the inference method. The systematic pattern is striking—every bootstrap-based method achieves near-nominal or conservative coverage, every influence-curve TMLE method under-covers, and HAL-TMLE is the exception that proves the rule. If LASSO-TMLE with bootstrap inference also hits 95%, the recommendation shifts from 'use OAL/GLiDeR' to 'use bootstrap inference with any regularized selection,' which would substantially weaken the selection-strategy claim. This is the load-bearing issue for the 'best calibrated' conclusion. The reader's PS-misspecification concern is valid but secondary—it bounds generalizability, while the inference-method confounding affects the internal validity of the ranking itself. The null-only design is a genuine limitation but the authors are transparent about it. This is a solid, well-executed benchmark that deserves a serious referee. The core LASSO-IPTW failure under rare exposure and the runtime accounting will hold up regardless of the calibration confounding issue. Recommend accept for review; the referee should push the authors to either run bootstrap inference for at least one TMLE pipeline to disentangle the confounding, or reframe Table 5's recommendations to explicitly separate selection-strategy effects from inference-method effects.","headline":"Useful benchmark of regularized confounding-adjustment pipelines, but the coverage ranking is confounded by inference method","tokens_in":14598,"tokens_out":576,"would_cite":false,"duration_ms":80197,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Rare exposures break standard confounding adjustment; targeted selection fixes them","keywords":["propensity score","doubly robust estimation","TMLE","variable selection","LASSO","confounding adjustment","plasmode simulation","rare exposure"],"falsifier":"If outcome-aware selection (OAL, GLiDeR) or HAL-TMLE failed to maintain near-nominal coverage under rare exposure or rare outcome in an independent plasmode with a different anchor dataset and a misspecified propensity model, the central recommendation would not hold.","tokens_in":14020,"feed_emoji":"⚖️","tokens_out":1270,"duration_ms":193602,"temperature":0.7,"pith_summary":"The paper sets up a head-to-head contest among ten statistical pipelines used to remove confounding bias in observational health studies where hundreds of proxy variables compete for a few outcome events. Using a plasmode simulation—real covariate data from a national health survey with a simulated null effect injected so that any nonzero estimate is pure residual bias—the authors show that when the exposure is rare, the most common approach (LASSO-selected propensity scores combined with inverse-probability weighting) produces the largest bias and unstable estimates. Switching the downstream estimator from weighting to a doubly robust targeted maximum likelihood step removes that bias and instability, at the cost of slightly under-covering confidence intervals. Methods that select variables based on their association with the outcome rather than the exposure—outcome-adaptive LASSO and group LASSO (GLiDeR)—maintain near-nominal 95% coverage across all rarity settings. The highly adaptive LASSO used as pure G-computation achieves near-zero bias but produces absurdly narrow intervals that cover the truth nearly 100% of the time, inflating relative error to 106–186%. The best-calibrated methods are also the slowest, ranging from under one second to over sixteen hours per analysis on a single core.","feed_headline":"Rare exposures break standard confounding adjustment; targeted selection fixes them","feed_subtitle":"Ten statistical pipelines tested under a known null show outcome-aware selection and doubly robust estimation best calibrated when events or","key_machinery":"The comparison hinges on two axes: (1) five regularized variable-selection families—traditional LASSO, collaborative-controlled LASSO, outcome-adaptive LASSO (OAL), group LASSO/GLiDeR, and highly adaptive LASSO (HAL)—that determine which proxies enter the adjustment set, and (2) two downstream estimators—IPTW (weighting by inverse treatment probability) and TMLE (doubly robust targeted updating of an outcome regression). The plasmode simulation design preserves real covariate correlations from NHANES data while imposing a known null risk difference, so any deviation from zero is measurable confounding bias. Performance is assessed via bias, coverage of 95% intervals, relative error of model–","core_discovery":"The decisive stressor is not the number of variables or the outcome rarity alone but the combination of rare exposure with weighting-based estimation. Under rare exposure (9.4% prevalence), LASSO-IPTW bias reaches −5.2×10⁻³ with inflated standard errors and conservative over-coverage near 98%, while replacing IPTW with TMLE for the same nuisance models cuts bias to near zero and reduces empirical SE, though coverage drops to 91–93%. Outcome-aware selection strategies (OAL, GLiDeR) and HAL-TMLE maintain 94–95% coverage across all three scenarios, making the choice between selection strategy and downstream estimator the central lever for calibration under rarity.","pith_inferences":["If the true propensity model were misspecified—as it likely is in real claims data where exposure drivers are poorly captured—the near-nominal calibration of OAL and GLiDeR might degrade, since the simulation's correctly specified PS gives all PS-based pipelines a common-mode advantage that would not hold in practice.","The mild undercoverage of TMLE pipelines under rarity (91–93%) may reflect a general tension: influence-curve standard errors condition on the selected nuisance model and thus ignore selection uncertainty, whereas bootstrap-based methods propagate it—at the cost of hours of computation.","The prescription-derived proxies in this anchor dataset are predominantly outcome predictors with weak treatment association, so the simulation may under-test scenarios where proxies are strong confounders of treatment rather than outcome.","Extending the comparison to non-null effects would test whether the calibration ranking at the null transfers to power and bias under realistic effect sizes, which the authors flag as a one-argument change to their released pipeline."],"forward_implications":["Researchers using healthcare databases with rare exposures should avoid LASSO-IPTW and either adopt outcome-adaptive selection or pair their selection with a doubly robust TMLE step to avoid the bias and instability documented here.","The six-order-of-magnitude runtime spread (sub-second to 16+ hours) means that method choice in applied pharmacoepidemiology is effectively a compute-budget decision as much as a statistical one, with HAL-TMLE offering the best calibration-per-compute trade-off.","The finding that HAL G-computation produces near-unity coverage with 106–186% relative error suggests that bootstrap-based inference for highly flexible estimators can systematically misrepresent uncertainty even when point estimates are accurate.","The rare-exposure arm exposed the largest methodological gaps, suggesting that future proxy-adjustment benchmarks should default to testing under rarity rather than balanced conditions."],"fun_headline_variants":["Rare exposure breaks LASSO-IPTW calibration; TMLE removes the bias","Outcome-aware selection and doubly robust estimation best under rarity","HAL-TMLE and GLiDeR hold 94–95% coverage across rarity scenarios","Rare exposure plus weighting is the decisive stressor in plasmode study","Swapping IPTW for TMLE cuts bias near zero but drops coverage to 91–93%"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The simulation generates exposure from a correctly specified main-terms logistic model, so the true propensity score lies within the parametric class that all propensity-based pipelines fit. This removes propensity-model misspecification as a bias source, meaning the strong calibration of outcome-aware methods is conditional on a correctly specified exposure model—a condition unlikely to hold in real healthcare database studies.","fun_headline_variants_meta":{"raw":{"variants":["Rare exposure breaks LASSO-IPTW calibration; TMLE removes the bias","Outcome-aware selection and doubly robust estimation best under rarity","HAL-TMLE and GLiDeR hold 94–95% coverage across rarity scenarios","Rare exposure plus weighting is the decisive stressor in plasmode study","Swapping IPTW for TMLE cuts bias near zero but drops coverage to 91–93%"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":1343,"prompt_tokens":1257,"completion_tokens":86,"prompt_tokens_details":null},"tokens_in":1257,"tokens_out":86,"duration_ms":14573,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T20:41:32.928078+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If outcome-aware selection (OAL, GLiDeR) or HAL-TMLE failed to maintain near-nominal coverage under rare exposure or rare outcome in an independent plasmode with a different anchor dataset and a misspecified propensity model, the central recommendation would not hold.","supporting_citations":[],"review_version":1}