{"id":"1129c794-4f1b-4f11-b6e5-27620fb0a079","arxiv_id":"2507.14689","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Penalized weighted GEE Buckley-James estimation gives variable selection with the oracle property for AFT models under stratified sampling with clustered failure times.","lead":"This paper develops a variable-selection method for survival studies that measure only a stratified sample of patients, when patients contribute two correlated teeth and the proportional-hazards assumption fails. The method reweights sampled clusters to undo sampling bias, uses generalized estimating equations for within-patient correlation, and applies a SCAD penalty to select risk factors, with theory and simulations supporting its use.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2's covariance matrix Bω in (13) omits within-cluster dependence, so the claimed oracle normality and 'reliable inference' are not established; the simulation's severe SE undercoverage at 90% censoring is consistent with this gap.","rationale":"The reader's weakest assumption concerned consistency of the pooled weighted Kaplan-Meier estimator when cluster members have different marginal distributions. That is a real limitation, but the more immediate and checkable flaw is the asserted asymptotic covariance in Theorem 2(3). Even granting the common marginal distribution F and the consistency of the KM imputation, the proof of Theorem 2(3) adopts the variance formula (13) from the independent-data Lai-Ying theory and does not derive the covariance contributions between members of the same cluster. Condition (C9) and formulas (12)-(13) contain only marginal probabilities, so Bω cannot represent the variance of a cluster sum unless the cluster members are independent. Since the paper's central selling point is reliable inference for clustered failure times under stratified sampling, an incorrect or unproven covariance matrix is the single most load-bearing concern. The simulation's large SEa/SEe gaps at 90% censoring provide empirical corroboration, although they could also stem from the multiplier bootstrap or selection effects; the proposed test isolates the variance formula directly. I therefore keep the reader's CONDITIONAL verdict but require a corrected derivation of Bω, or an explicit statement that the theory covers only working-independence structures with independent cluster members, plus a simulation check of the variance formula. The selection-consistency portion of the oracle property (β̂2 = 0) is less affected, which is why the paper has salvageable value, but the inferential claim cannot stand as written.","tokens_in":33604,"tokens_out":10475,"duration_ms":144240,"concrete_test":"Simulate the unpenalized weighted GEE-BJ estimating function at the true β0 under the exact DGP of Section 5: K=3, exchangeable Clayton errors with τ=0.6, 90% censoring, four-stratum sampling weights. Over 1000 replications, compute the empirical covariance matrix of n^{-1/2}U_{n,ω}(β0) and compare its diagonal to formula (13), evaluated with the empirical F̂ and the estimated censoring survivor functions. If the empirical diagonal exceeds the formula's diagonal by more than 2/sqrt(1000) times its Monte Carlo standard error, the formula misses the within-cluster covariance. A complementary diagnostic: repeat with K=1 under otherwise identical settings; the formula should match for K=1 but fail for K=3, isolating within-cluster dependence rather than the KM imputation or the sampling weights as the source of the discrepancy.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The load-bearing weakness is in the inferential half of the oracle property. Theorem 2(3) asserts that a centered and scaled penalized estimator converges to N(0, Bω_11), with Bω defined row-wise in (13). But (13) is a marginal variance formula: its integrand involves only F, its hazard λ, and the marginal censoring-survivor limits Γ0, Γ1, Γω,ℓ_1, Γω,ℓ_2 from condition (C9). The estimating function U_{n,ω}(β0) = Σ_i ω_i (X_i−X̄ω)^T Ω^{-1}{α(β0)}(Ŷ_i(β0)−X_iβ0) is a sum over clusters of K-dimensional imputed residual vectors, so its covariance necessarily includes cross-member terms Cov(U_{ik}, U_{ik′}) driven by the joint law of (ε_i1,...,ε_iK, C_i1,...,C_iK). C9 defines Γ0, Γ1, Γω,ℓ_1, Γω,ℓ_2 as limits of sums of marginal probabilities Pr(C_ij−X_ijβ0 ≥ t); no joint or cross-member probability appears anywhere in (13). Therefore Bω_11 in (14) cannot equal the true asymptotic covariance unless the K members of a cluster are independent. The proof of Theorem 2(3) simply states 'we follow Theorem 2 of Lai and Ying [29]' and does not derive the clustered covariance; transplanting an independent-data variance formula while changing the weights does not account for within-cluster dependence, which is exactly the feature the GEE framework is intended to address. The simulation results are consistent with this omission: at 90% censoring, the resampling SEa is 30–45% below SEe for the weighted estimators (e.g., SEa=7.78 vs SEe=11.28 for weighted-EX SCAD1 in Table 4), and coverage is only near nominal because the bias is small. This directly undermines the abstract's claim of a 'reliable inference procedure' and means the central theoretical result, as stated, is not supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a penalized weighted generalized estimating equation (GEE) approach to Buckley–James estimation for the semiparametric accelerated failure time model with clustered failure times under stratified sampling. The estimator solves the penalized estimating equation (6) via a two-layer iterative algorithm: an inner Newton–Raphson loop for the penalized GEE at fixed imputed responses, and an outer loop that updates the imputed responses through a weighted pooled Kaplan–Meier estimator. The authors state consistency and oracle properties in Theorems 1 and 2, including asymptotic normality of the active coefficients in (14). They use stratified cross-validation for tuning and SCAD penalties, and estimate standard errors by perturbation resampling after refitting the selected model. Simulation studies with 80–90% censoring compare weighted and unweighted estimators under three error distributions and working independence/exchangeable correlation, and a dental case-cohort application illustrates the methods.","tokens_in":33872,"tokens_out":12692,"duration_ms":161462,"significance":"If the theoretical results held, this would be a valuable contribution: it fills a real gap by combining variable selection, inverse-probability-of-sampling weights, and GEE-type within-cluster correlation in a semiparametric AFT model, with an algorithm that appears to converge in the reported settings. The simulation study is reasonably extensive and the structured cross-validation suggestion for stratified designs is sensible. However, the inference half of Theorem 2 is not established as written: the covariance formula (13) ignores within-cluster dependence, and the proof transplants an independent-data theorem. The selection-consistency and linearization parts are plausible, but the paper needs to supply the missing cluster-level covariance argument or revise the theorem. The manuscript does not appear to ship code or a reproducibility supplement, and several proof steps defer to 'similar arguments' in Lai and Ying [28,29].","major_comments":[{"comment":"The matrix Bω in Eq. (13) is a marginal variance formula built from the limits Γ0, Γ1, Γω,ℓ_1, and Γω,ℓ_2 in condition (C9), all of which are sums of marginal probabilities Pr(C_ij − X_ijβ0 ≥ t). The estimating function U_{n,ω}(β0) in Eq. (5) is a sum over clusters of K-dimensional residual vectors, so its asymptotic variance necessarily includes cross-member terms Cov(U_{ik}, U_{ik′}) coming from the joint law of (ε_i, C_i). The proof of Theorem 2(3) simply states 'we follow Theorem 2 of Lai and Ying [29]' and transplants an independent-data covariance; no argument shows that the within-cluster cross terms vanish. This is load-bearing because the abstract's 'reliable inference' claim rests on the normality statement (14). Since the simulation standard errors are obtained by perturbation resampling (Section 4.1) rather than from (13), Tables 2–4 do not validate the formula; the authors should derive the correct cluster-level covariance or state Theorem 2(3) with a covariance matrix that is estimable and prove the corresponding convergence.","section":"Section 3.3, Eq. (13) and Theorem 2(3)"},{"comment":"Lemma 2 is stated with an 'almost surely' remainder, but the proof establishes only Op bounds (20)–(21) for each marginal counting process and then sums the per-member expansions U^{(k)}_{n,ω}. Summing over k does not yield a joint central limit theorem or a correct covariance for clustered data; the decomposition in (36)–(37) can give the slope matrix Aω but not the variance Bω without a joint-cluster analysis of the estimating function. Please state the exact mode of convergence and supply the joint representation used for inference.","section":"Section 3.3, Lemma 2 and its proof"},{"comment":"The limits in (C9) involve only marginal survivor probabilities for each cluster member. Even if the errors have common marginal F, the asymptotic covariance of a cluster-level estimating equation depends on second-order joint features such as Pr(C_ij − X_ijβ0 ≥ s, C_ij′ − X_ij′β0 ≥ t) or the joint influence functions of the pooled Kaplan–Meier estimator. Unless such quantities are included in the conditions, the asserted covariance matrix in (13) is not identified from the assumptions; condition (C9) should be expanded or the covariance should be left generic.","section":"Section 3.3, Condition (C9)"}],"minor_comments":[{"comment":"Theorem 2(1) states 'with probability tending to 1, lim_{n→∞} Pr(...)' which is redundant; it should simply read Pr(β̂_{2,ω}=0)→1.","section":"Theorem 2(1)"},{"comment":"The displayed statement of (16) omits the constraint ∥β−β′∥≤n^{-r} that is used in the proof; the statement and proof should be aligned.","section":"Lemma 1"},{"comment":"Line 8 computes α̂(β̂^{(s)}_ω) but line 9 uses H_{n,ω}(α̂(b)); please clarify which value of α enters the Hessian in the update.","section":"Algorithm 1, lines 8–9"},{"comment":"There are several typos: 'standard Gumble' should be 'standard Gumbel', 'donated' should be 'denoted', and 'tunning' should be 'tuning'.","section":"Section 5"},{"comment":"At 90% censoring, the averaged resampling standard errors in the SCAD rows are 30–45% below the empirical standard errors (e.g., 7.78 vs 11.28 for weighted-EX SCAD1); the text comments only on the 50% censoring rows and should acknowledge this discrepancy.","section":"Table 4"},{"comment":"Equation (2) and the surrounding text contain minor typographical issues: 'n ˆYi(β)' lacks a multiplication sign and 'replaced' should be 'replaces'.","section":"Section 2, Eq. (2)"}],"recommendation":"major_revision","confidential_remarks":"The central theoretical claim needs a substantive repair, but the paper's algorithm and simulations are useful. I recommend major revision. If the authors cannot derive a correct covariance for the clustered estimating equation, they should either remove the claimed Bω formula and replace Theorem 2(3) with a result that leaves the covariance unspecified while proving convergence of a consistent estimator (e.g., the perturbation resampling), or explicitly weaken the 'reliable inference' claim. I would not reject because the gap is localized to the covariance formula and its proof and appears fixable within the paper's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Paul,\n\nQuick take: this is the first paper I've seen that does regularized variable selection for AFT models under stratified sampling with clustered failure times. That gap is real, and the simulation evidence that the weighted version beats the unweighted version is genuinely useful. But the paper overreaches in its inferential claim. The oracle normality in Theorem 2 rests on a variance matrix Bω that is built from marginal probabilities only. The estimating function is a sum of cluster-level score vectors, so its true asymptotic covariance has to include within-cluster cross terms. Those terms are nowhere in (13). The proof of Theorem 2(3) just cites Lai and Ying's independence-based result; transplanting that formula while rerouting the weights doesn't account for the dependence the GEE is supposed to model. So the 'reliable inference' advertised in the abstract is not supported by the stated theory. The severe undercoverage of the resampling SE at 90% censoring (SEa about 30-45% below SEe) is consistent with this gap, though the reported CPs are closer to nominal than that ratio alone would predict, which makes me suspect the perturbation method is also missing something.\n\nWhat's good: the setup is sensible, the simulation design is thorough (1000 reps, three error laws, three dependence levels), and the weighted correction clearly reduces bias and false positives. The iterative algorithm is described in enough detail to reproduce, though no code or data is shipped. The literature review is honest; the authors correctly say no one has done penalized selection under this design.\n\nMinor issues: the WOLS initial value violates Condition C7, and the paper only offers simulations to bridge that gap. The CV fold construction for tuning with imputed responses is described ambiguously. The appendix defers several steps to 'similar arguments' in Lai and Ying, which matters here because clustered dependence is exactly what those arguments do not cover.\n\nMy bottom line: the selection half of the paper is credible and worth refereeing. The inference half, as stated, is not. The right fix is to derive the correct cluster-robust sandwich variance and then revisit the resampling procedure. If that's done, the paper could be a solid contribution. As it stands, I'd require major revision, but I would not desk-reject.","headline":"First penalized selection for stratified clustered AFT data with convincing selection simulations, but the oracle inference claim is not supported because the variance matrix omits within-cluster dependence.","tokens_in":34662,"tokens_out":4335,"would_cite":false,"duration_ms":49926,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62N01","62N02","62J07"],"pacs":[],"model":"deepseek-v4-flash","headline":"A penalized weighted GEE achieves the oracle property for clustered survival data under stratified sampling.","keywords":["accelerated failure time model","Buckley–James estimator","generalized estimating equation","inverse probability weights","Kaplan–Meier estimator","oracle property","stratified case-cohort sampling","within-cluster dependence"],"falsifier":"Simulate clustered AFT data under the same stratified design but give different members of a cluster different marginal error distributions (for example, one member standard normal and another heavy-tailed), keeping conditions C1–C9 otherwise intact; if the coverage of the 95% Wald intervals for active coefficients drops substantially below nominal or the correct-model selection rate fails to improve with sample size, the consistency claim for the pooled weighted Kaplan–Meier imputation is contradicted.","tokens_in":1969,"feed_emoji":"📊","tokens_out":3774,"duration_ms":619513,"temperature":0.7,"pith_summary":"The paper builds a variable-selection and inference tool for semiparametric accelerated failure time (AFT) models when data come from stratified sampling designs, case-cohort studies being the main example, and failure times are clustered. It claims that a weighted penalized estimating equation, built from Buckley–James imputation of censored responses through a weighted pooled Kaplan–Meier estimator, recovers the true set of predictors and yields asymptotically normal coefficient estimates, meaning it has the oracle property. This matters because existing penalized survival methods either assume a full cohort, ignore the sampling weights or the within-cluster correlation, or rely on the proportional hazards assumption. If the theory holds, applied analysts gain a principled way to choose risk factors and construct confidence intervals when the AFT scale is appropriate and only a biased subsample was measured.","feed_headline":"Penalized estimator gains oracle property for clustered survival data","feed_subtitle":"A weighted GEE with Buckley–James imputation corrects stratified-sampling bias while selecting true risk factors.","key_machinery":"The load-bearing object is the penalized weighted estimating equation (6), whose censored responses are imputed by the Buckley–James mechanism: each censored log-failure time is replaced by the conditional expectation computed under the pooled weighted Kaplan–Meier estimator $\\hat{F}_{n,\\omega}^\\beta$ of the common marginal error distribution $F$ (equations (3)–(4)). Inverse-probability weights $\\omega_i$ correct the stratified sampling bias, the working covariance matrix $\\Omega(\\alpha)$ absorbs within-cluster dependence (its estimate is asymptotically negligible under condition C8), and the SCAD penalty drives sparsity. The inner–outer iterative algorithm (Algorithm 1) linearizes the discontinuous Buckley–James estimating function by fixing imputed responses at a current estimate $b$, solving a penalized weighted GEE by minorization–maximization plus Newton–Raphson, then updating $b$, following the iterative scheme of [18] and the penalized GEE updates of [43]. The limiting distribution is governed by two slope matrices: $A_\\omega$, the asymptotic slope of the imputed estimating function, and $D_\\omega$, the weighted least-squares curvature; their difference feeds the bias term in the oracle expansion (14).","core_discovery":"The paper's central claim is Theorem 2: under regularity conditions C1–C9, with tuning parameter $\\lambda_n \\to 0$ and $\\sqrt{n}\\,\\lambda_n \\to \\infty$, the approximate solution $\\hat{\\beta}_{\\omega}$ of the penalized weighted GEE (equation (6)) is an oracle estimator. Zero coefficients are estimated as exactly zero with probability tending to one, and the active coefficients satisfy an asymptotic normality expansion (equation (14)) centered on $\\beta_{10}$ with covariance $B_\\omega$, after subtraction of the penalty-induced bias $q_{01}$. The authors position this as the first formal treatment of penalized variable selection under stratified sampling in the semiparametric AFT framework with clustered failure times, carried by inverse-probability sampling weights, a pooled weighted Kaplan–Meier imputation for censored responses, and a working covariance matrix that absorbs within-cluster dependence.","pith_inferences":["The same inverse-probability weighting scheme should transfer to nested case–control and length-biased sampling designs; a decisive test would be whether the oracle expansion (14) survives those mechanisms, since the paper only sketches them as future work.","Because the common-marginal assumption is load-bearing, a cheap diagnostic for practitioners would compare member-specific Kaplan–Meier curves within clusters before pooling, since the theory gives no guidance on how much disagreement is tolerable.","For $p$ larger than the number of clusters, the Newton–Raphson inner loop is expected to break down; the boosting-style estimating-equation updates the paper points to are the plausible route, but their oracle behavior is untested.","The multiplier-resampling variance estimator assumes that perturbation of the weighted Kaplan–Meier imputation propagates correctly; a small-sample comparison against a strata-respecting bootstrap would show when the Wald intervals can be relied on."],"forward_implications":["When the AFT scale is appropriate, the estimator gives a stratification-aware alternative to penalized Cox regression: sampling weights remove the bias that unweighted variable selection incurs under unequal selection probabilities.","Confidence intervals obtained from refitting the selected model plus the multiplier-resampling variance step reach near-nominal coverage in simulations once weights are used, in contrast to unweighted approaches.","The method is robust to misspecification of the within-cluster correlation structure: an exchangeable working covariance helps at low to moderate censoring, while an independence working structure is a safe fallback at high censoring.","Stratified cross-validation, which splits within strata and aggregates folds, preserves case–control balance and is needed for reliable tuning-parameter selection in heavy-censoring designs.","The oracle property holds at each outer-layer iteration, so the theoretical guarantees apply even if the iterative algorithm has not fully converged."],"supporting_citations":[{"why":"Origin of the least-squares Buckley–James estimator for censored linear regression that the method's imputed responses build on.","marker":"[2]"},{"why":"The marginal semiparametric multivariate AFT estimator with GEE that this paper extends to stratified sampling and penalization.","marker":"[8]"},{"why":"Defines the SCAD penalty and the oracle-property framework that Theorem 2 transposes to the weighted GEE setting.","marker":"[14]"},{"why":"The iterative least-squares scheme for Buckley–James estimation that Algorithm 1 adapts into its two-layer procedure.","marker":"[18]"},{"why":"The penalized estimating-function theory this paper extends; supplies the approximate zero-crossing definition used for the penalized solution.","marker":"[21]"},{"why":"Empirical-process results, including Theorem 3, used in the proof of Lemma 1 for the weighted pooled Kaplan–Meier estimator.","marker":"[28]"},{"why":"Large-sample theory of the modified Buckley–James estimator, including tail modification and conditions C1–C4, the backbone of the asymptotics.","marker":"[29]"},{"why":"The GEE framework and moment-based correlation estimation used for the working covariance structure.","marker":"[30]"},{"why":"The penalized GEE with minorization–maximization and Newton–Raphson updates that the inner layer of the algorithm uses.","marker":"[43]"},{"why":"Generalization of the product-limit estimator used in Lemma 1's proof for the weighted empirical processes and tail stability condition C5.","marker":"[50]"}],"fun_headline_variants":["Oracle property for clustered AFT under stratified sampling","Penalized GEE with Buckley–James selects true risk factors","Clustered survival: penalized estimator beats sampling bias","Oracle variable selection for clustered failure times","Weighted GEE imputation: oracle property under stratification"],"cache_read_input_tokens":36352,"weakest_assumption_plain":"The pooled weighted Kaplan–Meier estimator must consistently estimate one common marginal error distribution shared by every member of every cluster — a premise the authors themselves flag in Section 7 as open — so if cluster members' error distributions differ, or censoring depends on unmeasured cluster-level factors, the imputation is inconsistent and the oracle theory collapses.","fun_headline_variants_meta":{"raw":{"variants":["Oracle property for clustered AFT under stratified sampling","Penalized GEE with Buckley–James selects true risk factors","Clustered survival: penalized estimator beats sampling bias","Oracle variable selection for clustered failure times","Weighted GEE imputation: oracle property under stratification"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00064,"raw_usage":{"total_tokens":2946,"prompt_tokens":945,"completion_tokens":2001,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":1924}},"tokens_in":561,"tokens_out":2001,"duration_ms":15940,"temperature":1.0,"reasoning_tokens":1924,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:51:36.167918+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate clustered AFT data under the same stratified design but give different members of a cluster different marginal error distributions (for example, one member standard normal and another heavy-tailed), keeping conditions C1–C9 otherwise intact; if the coverage of the 95% Wald intervals for active coefficients drops substantially below nominal or the correct-model selection rate fails to improve with sample size, the consistency claim for the pooled weighted Kaplan–Meier imputation is contradicted.","supporting_citations":[{"cited_title":"Buckley, I","cited_arxiv_id":null,"evidence_quote":"Origin of the least-squares Buckley–James estimator for censored linear regression that the method's imputed responses build on."},{"cited_title":"Chiou, S","cited_arxiv_id":null,"evidence_quote":"The marginal semiparametric multivariate AFT estimator with GEE that this paper extends to stratified sampling and penalization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the SCAD penalty and the oracle-property framework that Theorem 2 transposes to the weighted GEE setting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The iterative least-squares scheme for Buckley–James estimation that Algorithm 1 adapts into its two-layer procedure."},{"cited_title":"Johnson, D","cited_arxiv_id":null,"evidence_quote":"The penalized estimating-function theory this paper extends; supplies the approximate zero-crossing definition used for the penalized solution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Empirical-process results, including Theorem 3, used in the proof of Lemma 1 for the weighted pooled Kaplan–Meier estimator."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Large-sample theory of the modified Buckley–James estimator, including tail modification and conditions C1–C4, the backbone of the asymptotics."},{"cited_title":"Liang, S","cited_arxiv_id":null,"evidence_quote":"The GEE framework and moment-based correlation estimation used for the working covariance structure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The penalized GEE with minorization–maximization and Newton–Raphson updates that the inner layer of the algorithm uses."},{"cited_title":"Yang, A generalization of the product-limit estimator with an application to censored regression, The Annals of Statistics 25 (1997) 1088–1108","cited_arxiv_id":null,"evidence_quote":"Generalization of the product-limit estimator used in Lemma 1's proof for the weighted empirical processes and tail stability condition C5."}],"review_version":1}