{"id":"a6a01ec0-6b91-4e1d-8570-4a5a784003fa","arxiv_id":"1908.03968","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A split-sample bootstrap test for stacked estimating equations gives correct level in small samples and power close to the optimal F-test in simulations.","lead":"This paper proposes a bootstrap hypothesis test for two-stage estimation problems where one parameter estimate feeds into a second estimating equation. It uses repeated random sample splits and a simulated null distribution to keep false positives near the nominal level when samples are small.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The finite-sample exactness claim is only specified for null simulations that ignore covariates; Section 6's logistic model has covariates Z under H0, and no null simulation preserving Z is given, so the reported p=0.003 is unsupported.","rationale":"The reader's weakest assumption is exactly the unstated null simulation for the Section 6 analysis with covariates Z, and my stress-test concurs. The paper's exactness proof and simulation code are only given for the linear setting where the null makes Y independent of X (Supplement S.1, Section 5.3). In the real-data logistic model, the null does not make Y independent of the covariates Z, and the first-stage theta is estimated from a different outcome W; the null distribution of the averaged p-value must reflect both features. The paper is silent on this, so the headline p=0.003 is not reproducible and the claimed finite-sample level does not demonstrably extend to the application. This does not refute the core bootstrap idea for settings where the null simulation is correctly specified, but it is a load-bearing gap for the paper's motivating data example and for its general claim of exactness. The reader's CONDITIONAL verdict is appropriate: the authors should either supply the missing null simulation (and code) or explicitly restrict the exactness claim to the no-nuisance linear case. No change to the verdict is needed beyond what the reader already recommended.","tokens_in":11501,"tokens_out":10014,"duration_ms":112890,"concrete_test":"For the Section 6 model, specify and run a null simulation that generates Y* from the fitted null logistic model pr(Y=1|Z)=H(Z^T beta_hat), or an equivalent binary-outcome residual/permutation bootstrap, while keeping X and W as observed and preserving the B-split estimation of theta from W; recompute p*. If the recomputed p* remains below 0.05, the Section 6 conclusion is robust; if it exceeds 0.05, the reported significance is an artifact of the unspecified null simulation. The authors should also report N, B, and the null distribution of pH1 for this simulation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the bootstrap test is exact or nearly exact in finite samples. Section 5.2 establishes exactness only for a null under which Y is independent of X and there are no unknown nuisance parameters. The simulation protocol in Supplement S.1.1 generates Y* ~ N(0,1) independent of X, and S.1.2 uses residual resampling from the linear model Y = X^T theta + epsilon. The Section 6 data analysis, however, tests H0: beta0 = 0 in model (12), pr(Y=1|X,Z)=H{beta0 f(X;theta) + Z^T beta}. Under this null, Y still depends on the covariates Z through beta, and the first-stage score f(X;theta) is estimated from a separate outcome W on each split. The paper nowhere specifies how the null distribution of pH1 was simulated for this case: whether Z was included in the null model, how beta was estimated, whether theta was re-estimated from W or held fixed, or how residual resampling was adapted to a binary outcome. Without this specification, the p* = 0.003 reported in Section 6 cannot be checked, and the finite-sample level guarantee from Section 5.3 does not automatically transfer. The paper's own Section 5.2 concedes that with estimated nuisance parameters the test is not exact, so the abstract's claim that the procedure 'does not rely on asymptotic approximations' overstates the realistic scope.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a bootstrap-based hypothesis test for parameters in stacked estimating equations. The procedure repeatedly splits the sample, estimates the first-stage parameter θ on the training portion, then estimates the second-stage parameter β on the test portion, and averages the per-split p-values. The null distribution of this averaged p-value is obtained by simulating responses under the null and repeating the splitting procedure. The authors claim the test is exact up to simulation error when the null imposes no nuisance parameters and the error distribution is known, and they present simulations showing correct level and power close to the F-test. A data analysis builds a physical activity score and tests its association with mortality.","tokens_in":11808,"tokens_out":11419,"duration_ms":105378,"significance":"The idea of combining repeated sample splitting with a bootstrap null distribution is useful for small-sample inference in estimating-equation settings where Wald tests fail. The exactness argument for the strong null with known errors is logically sound, and the level/power simulations in Tables 2–5 are extensive and largely support the procedure's practical performance in the linear setting. However, the paper’s central claims are weakened by the absence of a derivation of the limiting distribution (deferred to a companion paper), a numerical mismatch in the reported power difference, and the lack of a specification of the null simulation for the real-data logistic regression application.","major_comments":[{"comment":"The abstract states that 'we derive the limiting distribution of sample splitting estimator,' but Section 4.3 only summarizes results from Kravitz et al. (2019) and contains no derivation. The manuscript should either provide the derivation or revise the abstract to credit the companion paper.","section":"Abstract and §4.3"},{"comment":"Section 5.4 reports that 'the maximum difference between the two tests is 0.015,' but the displayed tables contradict this. In Table 4, for β0=0.6 and B=10, the bootstrap power is 0.91 while the F-test power is 0.95, a difference of 0.04; Table 5 shows the same discrepancy. The reported 0.015 figure should be corrected, and the power comparison should be stated consistently with the tables.","section":"§5.4, Tables 4–5"},{"comment":"Section 6 reports p*=0.003 for testing β0=0 in the logistic model (12), but the paper nowhere specifies how the null distribution of pH1 was simulated in this setting. Supplement S.1.1 and S.1.2 describe null simulations only for linear models; they do not state how covariates Z are handled under the null (in particular, the null model still depends on Z through β), how θ is re-estimated in null replications, or how residual resampling would be adapted to a binary outcome. Without this information, the reported p-value cannot be reproduced, and the finite-sample level properties from Section 5.3 do not automatically transfer.","section":"§6 and Supplement S.1"},{"comment":"Model (8) uses the same outcome Y in both estimating equations, so under the alternative β0≠0 the two equations cannot both be unbiased: if Y = β0 X^T θ + ε, then E[X(Y−X^T θ)] = (β0−1)E[X X^T θ] ≠ 0. Thus the first-stage estimator is inconsistent under the alternative, and the power simulations in Tables 4–5 are not simulations of a correctly specified stacked estimating equations model as defined in Section 4. The authors should either use two distinct outcomes (W and Y) as in Section 4, or explicitly justify the working-model interpretation and its implications for the test.","section":"§5.3, model (8)"},{"comment":"Section 5.1 (step 2) and Supplement S.1 leave unspecified how the per-split p-value p_b is computed from the second-stage estimating equation. Since pH1 is defined as the mean of the p_b, this is a central ingredient of the test statistic. In particular, it is not stated whether a Wald, t, or bootstrap p-value is used, and whether the second-stage regressor X^T θ̂_b is treated as fixed or whether the uncertainty from estimating θ̂_b is accounted for. The algorithm is therefore not fully defined as written.","section":"§5.1 and Supplement S.1"}],"minor_comments":[{"comment":"The notation β̂_b := B^{-1} Σ_b β̂_b is self-referential; use a distinct symbol (e.g., β̂_avg) for the averaged estimate.","section":"§5.1, step 3"},{"comment":"The expression I(~p_n) should be I(~p_i) or I(~p_b) to index the null p-values, not the sample size n.","section":"§5.1, step 5"},{"comment":"The text says simulations use B=50 and β0 from −1 to 1, but Table 4 shows B=10,...,250 and β0 from 0 to 1; please align the text with the tables.","section":"§5.4"},{"comment":"The distributions of X used in the level and power simulations are not stated; please specify whether X is fixed or random and from which distribution.","section":"§5.3"},{"comment":"‘uniformally’ should be ‘uniformly’, and ‘The variance assumed to be known’ should be ‘The error variance is assumed known’.","section":"Table 4 caption"},{"comment":"‘calucaltions’ should be ‘calculations’.","section":"§5.5"},{"comment":"The citation to Efron (1977) for residual resampling is likely incorrect; residual resampling is usually attributed to Efron (1979) or Efron and Tibshirani (1994).","section":"References"},{"comment":"‘The p-values from the from splits’ has a duplicated phrase.","section":"§6.3"}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily on a companion paper (Kravitz et al., 2019) for its asymptotic theory, which is only cited as a reference; the authors should clarify the relationship and availability of this companion paper. The numerical mismatch in the power claim and the underspecified null simulation in Section 6 raise questions about the rigor of the reported results, but the core idea may be salvageable with a careful revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Kravitz, Carroll, and Ruppert propose a bootstrap-based hypothesis test for stacked estimating equations using repeated sample splitting and averaging of p-values. The new piece is applying the multi-split idea of Meinshausen et al. to this two-stage M-estimation setting and simulating the null distribution of the averaged p-value rather than using a conservative quantile correction. That is a genuine, if modest, extension. The simulation work is done carefully: level stays near nominal across error distributions, power tracks the UMP F-test within 0.015, and residual resampling works when the error distribution is unknown. The variance calculation for choosing B is a nice touch. For the strong null where the error distribution is known and there are no nuisance parameters, the test is exact up to simulation error, and the paper says this clearly.\n\nThe soft spots are real but not fatal. First, the abstract says 'we derive the limiting distribution' and that the procedure 'does not rely on asymptotic approximations.' The derivation is not in this paper; it is deferred to a companion paper, and the bootstrap test is only exact in the narrow no-nuisance case. Those are overstatements. Second, the data analysis in Section 6 tests beta0 = 0 in a logistic model with covariates Z, but the paper never specifies how the null distribution of pH1 was simulated for that model. The supplement only shows the null simulation for Y independent of X (or residual resampling from a linear first-stage). Without knowing whether Z was included and how theta and beta were handled under the null, the reported p = 0.003 cannot be checked. This is a load-bearing gap in the application section, not just a stylistic complaint.\n\nThe references are appropriate. The self-citation to the companion paper is fine as a pointer, but the abstract should not claim derivations that live elsewhere. The authors are honest that with estimated nuisance parameters the test is approximate.\n\nWho is this for? Statisticians working with small-sample M-estimation and two-stage estimators. It deserves a serious referee, but the revision should fix the abstract, either include the asymptotics or explicitly defer them, and document the Section 6 null simulation. My verdict: conditional accept.","headline":"Solid bootstrap test for stacked estimating equations, but the exactness claim only holds in a narrow case and the real-data analysis is missing key null-simulation details.","tokens_in":12308,"tokens_out":2316,"would_cite":false,"duration_ms":23688,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62F40","62G09","62J12"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces a bootstrap hypothesis test for stacked estimating equations that uses repeated sample splitting and averaged p-values to achieve exact or nearly exact finite-sample level, with power close to the uniformly most…","keywords":["stacked estimating equations","bootstrap hypothesis test","sample splitting","p-value lottery","finite-sample level","M-estimation","residual resampling","physical activity score"],"falsifier":"Simulate null data from the real-data logistic model: $Y$ is a binary outcome whose probability depends on covariates $Z$ only, with $\\beta_0=0$ and with the activity predictors $X$ correlated with $Z$; run the paper's bootstrap algorithm exactly as described, regenerating $Y^*$ without $Z$. If the rejection rate over many null datasets exceeds the nominal level by more than simulation error, the claimed level fails for this null model.","tokens_in":11303,"feed_emoji":"📊","tokens_out":8083,"duration_ms":81684,"temperature":0.7,"pith_summary":"Stacked estimating equations solve two parameters jointly when the second estimating equation depends on the first parameter; Wald tests based on the asymptotic covariance matrix can have unreliable level in small samples. The paper proposes a bootstrap test: split the sample repeatedly, estimate the first parameter on one half, plug it into the second equation on the other half, take the mean of the resulting p-values, and compare that mean to an empirical null distribution obtained by simulating responses under the null. When the null response distribution is known, the test is exact up to simulation error; with residual resampling it has approximately the nominal level. In Gaussian linear simulations with $n=100$, the maximum power difference from the uniformly most powerful F-test is 0.015. The paper also cites a companion result that the sample-splitting estimator is asymptotically equivalent to the stacked estimating equation estimator in parametric models.","feed_headline":"Split-sample bootstrap test holds its level in small samples","feed_subtitle":"Averaging p-values over 50 splits matches the F-test's power to within 0.015.","key_machinery":"The load-bearing object is the averaged split p-value $p_{H_1}=B^{-1}\\sum_{b=1}^B p_b$, where each $p_b$ comes from a fresh Bernoulli(1/2) sample split. Its null distribution is not uniform even though each individual $p_b$ is, so the paper simulates the null distribution by bootstrap: draw responses $Y^*$ independent of the predictors, or residual-resample with replacement, then rerun the whole splitting and estimation procedure many times and set $p^*=\\hat F(p_{H_1})$, the empirical fraction of null $p_{H_1}$ values below the observed value. This empirical CDF of null p-values carries the inference, replacing the asymptotic covariance matrix of the stacked estimator. For parametric models, the paper invokes the companion asymptotic expansion showing that the split estimator has the same limiting distribution as the stacked estimating equation estimator, which justifies the method while the bootstrap handles finite-sample calibration.","core_discovery":"The central claim is that finite-sample inference for stacked estimating equations can be made exact, or nearly exact, by a bootstrap that operates on split-sample p-values rather than on asymptotic variance estimates. The method estimates $\\theta$ from the training half using $\\sum_i \\delta_{ib}\\Psi(W_i,\\theta)=0$, then estimates $\\beta$ from the test half using $\\sum_i (1-\\delta_{ib})K(Y_i,\\hat\\theta_b,\\beta)=0$, producing an individual p-value $p_b$; the test statistic is $p_{H_1}=B^{-1}\\sum_b p_b$. Under the null, responses are regenerated either from the known error distribution, which gives an exact test up to simulation error, or by residual resampling, which gives an approximately valid level. The paper reports simulated level near nominal for eight error distributions and power within 0.015 of the F-test, and it applies the procedure to test whether a physical-activity score predicts mortality.","pith_inferences":["Editorial inference: if the null simulation is replaced by one that preserves the covariate structure $Z$ used in the real-data logistic model, the claimed level should be checked separately; the paper describes a null simulation without $Z$, so the reported $p=0.003$ would be on firmer ground with a $Z$-preserving null bootstrap.","Editorial inference: the same averaged-p-value bootstrap should extend to generalized estimating equations and other two-stage estimators where asymptotic variance estimates are unreliable, provided a credible null response distribution can be simulated; this is a natural next step the paper does not pursue.","Editorial inference: one can benchmark the method's calibration directly by comparing the empirical distribution of $p_{H_1}$ under a known null with the bootstrap's estimate $\\hat F$; any gap would show exactly where finiteness bites."],"forward_implications":["At samples as small as $n=100$, the bootstrap test keeps rejection rates near the nominal level across normal, $t$, and Laplace errors, so practitioners no longer need to rely on asymptotic standard errors in this setting.","The test's power tracks the uniformly most powerful F-test; the largest simulated gap is 0.015, and residual resampling costs only a small amount of power compared with knowing the error distribution.","Because separate p-values from single splits can range from 0 to 1, averaging over $B\\approx 50$ splits removes the p-value lottery while retaining proper level.","For parametric models, the sample-splitting estimator and the stacked estimating equation estimator have the same limiting distribution, so the bootstrap output can be interpreted as inference for the stacked solution itself."],"supporting_citations":[{"why":"defines stacked estimating equations as the simultaneous solution of two M-estimating equations, the target class of the test.","marker":"Carroll et al. (2006, Appendix A.6.6)"},{"why":"supplies the M-estimation asymptotics that justify solving the stacked vector equation.","marker":"Huber (1964, 1967)"},{"why":"overview of M-estimation used to frame the estimating-equation setting and its asymptotic covariance.","marker":"Stefanski and Boos (2002)"},{"why":"provides the asymptotic expansion and equivalence result for sample splitting estimators that the method relies on for parametric justification.","marker":"Kravitz et al. (2019)"},{"why":"introduces split-sample testing after selection, the single-split precursor this paper generalizes.","marker":"Wasserman and Roeder (2009)"},{"why":"introduces multi-split p-value aggregation that the paper replaces with a bootstrap average p-value.","marker":"Meinshausen et al. (2009)"},{"why":"documents the p-value lottery that motivates repeated splitting.","marker":"Dezeure et al. (2015)"},{"why":"supplies residual resampling and bootstrap resampling used for the approximate null distribution.","marker":"Efron and Tibshirani (1994)"},{"why":"basis for residual resampling used in the null simulation.","marker":"Efron (1977)"},{"why":"benchmark uniformly most powerful F-test that the power simulations compare against.","marker":"Neyman and Pearson (1928, 1933)"}],"fun_headline_variants":["Bootstrap test for stacked equations works in small samples","Split-sample bootstrap gives exact finite-sample testing","Non-asymptotic hypothesis test for stacked estimating equations","Averaged split p-values match F-test power to 0.015","Exact bootstrap test for stacked equations without asymptotics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the bootstrap null distribution used to calibrate the test matches the real null model under which the data were generated; in the real-data example the null model includes additional adjustment variables, while the paper's described null simulation regenerates responses without those variables, so the stated p-value may not have the claimed exact level unless the null simulation is adapted.","fun_headline_variants_meta":{"raw":{"variants":["Bootstrap test for stacked equations works in small samples","Split-sample bootstrap gives exact finite-sample testing","Non-asymptotic hypothesis test for stacked estimating equations","Averaged split p-values match F-test power to 0.015","Exact bootstrap test for stacked equations without asymptotics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00066,"raw_usage":{"total_tokens":2988,"prompt_tokens":882,"completion_tokens":2106,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":2027}},"tokens_in":498,"tokens_out":2106,"duration_ms":14459,"temperature":1.0,"reasoning_tokens":2027,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:56:01.688804+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate null data from the real-data logistic model: $Y$ is a binary outcome whose probability depends on covariates $Z$ only, with $\\beta_0=0$ and with the activity predictors $X$ correlated with $Z$; run the paper's bootstrap algorithm exactly as described, regenerating $Y^*$ without $Z$. If the rejection rate over many null datasets exceeds the nominal level by more than simulation error, the claimed level fails for this null model.","supporting_citations":[{"cited_title":"J., Ruppert, D., Crainiceanu, C","cited_arxiv_id":null,"evidence_quote":"defines stacked estimating equations as the simultaneous solution of two M-estimating equations, the target class of the test."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the M-estimation asymptotics that justify solving the stacked vector equation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"overview of M-estimation used to frame the estimating-equation setting and its asymptotic covariance."},{"cited_title":"S., Carroll, R","cited_arxiv_id":null,"evidence_quote":"provides the asymptotic expansion and equivalence result for sample splitting estimators that the method relies on for parametric justification."},{"cited_title":"and Roeder, K","cited_arxiv_id":null,"evidence_quote":"introduces split-sample testing after selection, the single-split precursor this paper generalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"introduces multi-split p-value aggregation that the paper replaces with a bootstrap average p-value."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"documents the p-value lottery that motivates repeated splitting."},{"cited_title":"and Tibshirani, R","cited_arxiv_id":null,"evidence_quote":"supplies residual resampling and bootstrap resampling used for the approximate null distribution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"basis for residual resampling used in the null simulation."},{"cited_title":"and Pearson, E","cited_arxiv_id":null,"evidence_quote":"benchmark uniformly most powerful F-test that the power simulations compare against."}],"review_version":1}