{"id":"e9e263bd-d978-4aa5-8d0a-e44a0ebf42cb","arxiv_id":"1908.03967","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Repeated random sample splits in a two-stage model asymptotically produce the same estimates as solving one stacked estimating equation on the full data set.","lead":"This paper studies what happens when researchers split a dataset in two, build a health score on one half, and test that score on the other half, repeating the split many times. It proves that with enough repeated splits, the results match fitting one combined model to all the data, so ordinary confidence intervals can be used.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equivalence theorem depends on an unproved uniform control of per-split remainder terms; the Appendix only cites Carroll et al. (1996) and does not verify it for the Section 6 two-stage logistic model.","rationale":"The reader identified the weakest assumption as the uniform control of per-split remainder terms; this is indeed the load-bearing condition in the proof of Section 4.3. The paper's Appendix A.3 is explicitly a sketch and cites Carroll et al. (1996) for the remainder order without checking that the condition holds for the specific two-stage estimating equations and for the rescaled physical activity score in Section 6.2. The score construction involves non-differentiable operations (minima, maxima, absolute-value shifts), so the differentiability assumption in A.1 is not automatically satisfied for the motivating application. This is a real proof gap, but it is not a demonstrated failure: for smooth M-estimators the remainder bound is a standard result, and the simulations in Section 5 are consistent with the claimed equivalence. Therefore the reader's CONDITIONAL verdict is appropriate. I would not change the verdict; the concern should be resolved by adding a rigorous uniform remainder proof or by verifying the condition for the logistic model before the paper is accepted as a complete theoretical contribution.","tokens_in":12527,"tokens_out":24751,"duration_ms":244932,"concrete_test":"Independently re-derive the remainder bound for the two-stage logistic estimating equations in Section 6 with the rescaled score of Eq. (12). Specifically, show that sup_b |R_{nb}| = O_P(n^{-1/2}) and that B^{-1}Σ_b R_{nb} = o_P(1) under B = o(n^{1/2-a}); if the bound fails, the equivalence theorem does not apply to the motivating example. As a numerical cross-check, compute the full-data stacked estimate on the NIH-AARP data (or a similar simulated cohort) and compare it to the B=50 averaged split estimate reported in Table 5; a discrepancy larger than the reported standard error would indicate the asymptotic equivalence is not realized at this finite sample size.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central theorem of Section 4.3 is proved by averaging per-split expansions (A.2)-(A.3). The step 'the oP(1) terms ... are OP(n^{-1/2+a})' is imported from Carroll et al. (1996) and is asserted, not proved, for the present two-stage setting. Since the averaged estimator is B^{-1}Σ_b(·), the proof needs a uniform-in-b control of the remainder; without it, B^{-1}Σ_b R_{nb} need not vanish at the stated rate. The gap is not merely cosmetic: the motivating score in Section 6.2 is constructed with minima, maxima, and an absolute-value shift, so f(X;θ) in Eq. (12) may fail the differentiability assumption in Appendix A.1, which is exactly the condition used to justify the remainder order. Thus the claimed equivalence, and with it the conclusion that multi-split inference adds nothing beyond stacked estimating equations, rests on an unverified regularity condition.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops asymptotic theory for sample splitting in a two-stage estimation setting. In the first stage, a parameter θ is estimated from a training half of the data; in the second stage, a function f(X; θ) of that estimate is used as a covariate in a model for β estimated on the validation half. The authors provide asymptotic expansions and variance estimators for one split and for a fixed finite number of splits B, and then claim that when B grows with n at rate B = o(n^{1/2-a}), the split-averaged estimator is asymptotically equivalent to the stacked estimating equations estimator computed on the entire data set. The motivating application is a physical activity score for mortality risk in the NIH-AARP cohort; simulations for linear and logistic models examine finite-sample coverage of confidence intervals.","tokens_in":12785,"tokens_out":7746,"duration_ms":82945,"significance":"If the equivalence result holds under the stated conditions, it is a genuinely interesting conceptual contribution: it says that many-split averaging does not provide asymptotic information beyond full-data stacked estimating equations, and it gives practitioners a way to justify confidence intervals in split-sample analyses. The paper also provides useful simulation evidence and a realistic data application, with explicit asymptotic expansions that can be used directly. However, the proof is presented only as a sketch, and the key step that makes the averaging over B splits work is imported from Carroll et al. (1996) without verification for the specific two-stage estimating equations studied here.","major_comments":[{"comment":"The central equivalence theorem depends on the assertion that the oP(1) remainder terms in expansions (A.2) and (A.3) are OP(n^{-1/2+a}) \"as in Carroll et al. (1996)\". When averaging over B splits, the proof needs this remainder bound to hold uniformly in b, because the step from the averaged expansions (A.4)-(A.5) to the limiting expansions (A.6) requires control of B^{-1} sum_{b=1}^B R_{nb}. A pointwise OP(n^{-1/2+a}) statement for each fixed b does not control the maximum or the average of the remainders, and the manuscript does not state or verify the required stochastic equicontinuity or uniform remainder condition. Since this uniformity is exactly what produces the rate restriction B = o(n^{1/2-a}), the main claim in Section 4.3 is not established as written.","section":"Section 4.3 and Appendix A.3"},{"comment":"The variance estimator for a finite fixed number of splits is incorrectly defined. After defining A_i = B^{-1} sum_{b=1}^B A_{ib}, the text defines Â_i = sum_{b=1}^B Â_{ib}, omitting the factor B^{-1}. With this definition, the sample covariance S(Â_i) would be approximately B^2 times the target covariance, so the resulting standard errors for the averaged estimator would be wrong by a factor of B. The definition should be Â_i = B^{-1} sum_{b=1}^B Â_{ib}, and the left-hand sides of the two displays in this section, which read n^{1/2}(θ̂_b - θ) and n^{1/2}(β̂_b - β), should refer to the averaged estimators θ̂ and β̂, respectively.","section":"Section 4.2"},{"comment":"The asymptotic theory assumes that f(X; θ) is differentiable with respect to θ (Appendix A.1). The physical activity score defined in Section 6.2 depends on θ through the location of the minimum of each fitted marginal model, the absolute value of that minimum, and the total T obtained by summing the maxima of the rescaled curves. These operations are not automatically differentiable in θ, and the manuscript does not provide conditions under which the score is smooth, such as uniqueness and interiority of the optimizing points together with an envelope theorem argument. Without such verification, the theorem of Section 4.3 is not directly applicable to the motivating data analysis, and the confidence intervals reported in Table 5 rely on an unverified regularity condition.","section":"Section 6.2 and Appendix A.1"}],"minor_comments":[{"comment":"The displayed expansion for n^{1/2}(β̂ - β) contains K(Y_i, β, θ̂_b), but the subscript b is not defined in the averaged expansion; this should be K(Y_i, β, θ), matching the appendix version in (A.6).","section":"Section 4.3, Eq. (8)"},{"comment":"The assumption that E(θ̂_k) = E(θ̂_l) = θ and E(β̂_k) = E(β̂_l) = β is not generally satisfied by nonlinear M-estimators and is not used in the subsequent expansions; it would be cleaner to replace this with a consistency assumption or remove it.","section":"Appendix A.1"},{"comment":"The appendix is labeled a sketch, but the uniform remainder control is the load-bearing step of the paper's main theorem; a proof or a precise high-level condition should be supplied instead of a citation.","section":"Appendix A.3"},{"comment":"\"Motivating by problems\" should be \"Motivated by problems\".","section":"Abstract"},{"comment":"The phrase \"They create B copies of the data\" is imprecise; the procedure creates B random partitions of the same dataset, not B copies.","section":"Section 1"},{"comment":"The word \"logistc\" should be \"logistic\".","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is appropriate for a statistics methodology journal and the central conceptual message is attractive. The main obstacle is the unproved uniformity condition behind the B-growth result in Section 4.3; if the authors can supply that argument or state the needed regularity conditions precisely, the paper would be publishable. The variance estimator typo in Section 4.2 must also be corrected since it affects the practical finite-B procedure. The self-citation to Kravitz et al. (2019) is not needed for the equivalence claim and does not by itself raise concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The paper's main contribution is the asymptotic equivalence result in Section 4.3: as the number of splits B and n go to infinity, with B growing slower than n^{1/2-a}, the averaged split-sample estimator converges to the full-data stacked estimating equations estimator. That is a useful formalization of something people suspected, and it is new in this two-stage setting where the first-stage estimate generates a covariate for the second stage. The single-split and finite-B expansions are standard M-estimation with random Bernoulli weights, but they are laid out cleanly. The simulations are honest and show the expected Fieller under-coverage at small n. The code is available, and the application to physical activity scoring is sensible.\n\nThe soft spots are real but not fatal to the theory. The proof of the B→∞ equivalence needs uniform control of the per-split remainder terms, and the authors import that from Carroll et al. (1996) rather than proving it for their setting. More importantly, the score f(X;θ) constructed in Section 6.2 uses minima, maxima, and an absolute-value shift. That makes f non-differentiable in θ, which violates the assumption in Appendix A.1 that the theorem relies on. So the motivating application may not actually be covered by the theory. The authors could fix this by smoothing the score or by showing the non-smoothness is benign. Also, there is a typo in Section 4.2: the definition of ˆAi is missing the B^{-1} factor. The appendix is a sketch; it is enough for a first pass but not for the result to be considered fully proved.\n\nOverall, the central theorem is plausible and the simulations support it for smooth f. The paper deserves a serious referee: the equivalence result is worth publishing, but the referee should require the authors to address the uniformity assumption and the non-smoothness of the application's score. I would not desk reject it.","headline":"The paper proves a clean asymptotic equivalence between averaged sample splitting and stacked estimating equations, but the motivating physical activity score likely violates the smoothness assumptions the proof relies on.","tokens_in":13227,"tokens_out":2685,"would_cite":true,"duration_ms":28738,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62J12","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that averaging many sample splits yields the same asymptotic estimator as fitting both stages to the full data, so multi-split inference adds no information beyond the unsplit analysis.","keywords":["sample splitting","M-estimation","estimating equations","stacked estimating equations","two-stage modeling","physical activity score","asymptotic theory","multiple splits"],"falsifier":"Simulate the paper's logistic two-stage model at sample sizes $n = 250$ and $n = 1000$, compute the $B$-split average for $B$ proportional to $n^{0.6}$ and to $n^{1/2}$ with $\\pi = 1/2$, and examine $\\sqrt{n}$ times the difference between the split average and the stacked-estimating-equations estimate; if this difference does not shrink to zero as $n$ grows, the uniform remainder assumption fails.","tokens_in":12347,"feed_emoji":"📊","tokens_out":11378,"duration_ms":106909,"temperature":0.7,"pith_summary":"Sample splitting lets one data set serve two roles—here, building a physical-activity score on half of a cohort and then using that score to predict mortality on the other half—without letting the same observations do both jobs. The paper develops asymptotic theory for the split-sample estimator in a two-stage M-estimation setting and gives formulas for standard errors and confidence intervals. Its central result is that as the number of splits $B$ and the sample size $n$ grow, with $B$ growing more slowly than $\\sqrt{n}$, the average of the split estimates is asymptotically the same as solving the stacked estimating equations on the whole data set. Thus when the split proportion is $\\pi = 1/2$, the multi-split estimator carries no information beyond an ordinary full-data analysis, and confidence intervals can be obtained without splitting at all. The motivating application builds a physical-activity score from the NIH-AARP cohort and estimates its association with mortality: the log-odds coefficient is about $-0.026$, with a confidence interval excluding zero.","feed_headline":"Averaging many sample splits equals one full-data analysis","feed_subtitle":"The paper shows the multi-split average becomes the same as solving the estimating equations on all the data.","key_machinery":"The load-bearing object is the per-split asymptotic linear expansion of each M-estimator. Each split contributes an influence-function term whose average over $B$ splits, combined with the law-of-large-numbers limits $B^{-1}\\sum_b \\delta_{ib} \\to \\pi$ and $B^{-1}\\sum_b (1-\\delta_{ib}) \\to 1-\\pi$, reproduces the influence function of the stacked estimating equations estimator—the estimator obtained by solving the first-stage and second-stage score equations jointly on the entire data set. The technical condition that makes the collapse work is that the $o_P(1)$ remainder terms in the per-split expansions are uniformly $O_P(n^{-1/2+a})$ across all splits, so their average vanishes when $B$ grows slower than $n^{1/2-a}$.","core_discovery":"The paper sets up each sample split as two estimating equations: the first-stage parameter $\\theta$ is estimated from the Bernoulli($\\pi$) training subsample, and the second-stage parameter $\\beta$ is estimated from the complementary subsample using the fitted score $f(X;\\hat{\\theta})$ as a covariate. A single split has a standard M-estimation expansion, and the paper obtains the joint limiting distribution and a covariance estimator for it. For a fixed number of splits $B$, the final estimators are the sample means of the per-split estimates, and their covariance is estimated from averaged influence functions. The main theorem in Section 4.3 lets $B$ grow with $n$ at rate $B = o(n^{1/2-a})$ for any $a>0$. Because the Bernoulli indicators average to their probabilities, $B^{-1}\\sum_{b=1}^{B}\\delta_{ib} \\to \\pi$, and the averaged influence functions collapse to those of the stacked estimating equations estimator that solves both score equations on the entire sample. The paper states the conclusion directly: when $\\pi = 1/2$, the estimates from sample splitting become asymptotically equivalent to estimates from the entire data set without sample splitting.","pith_inferences":["Editorial inference: for estimation and confidence intervals, the theorem implies that once the sample is large, spending computation on many splits is unnecessary; the full-data stacked estimating equations give the same asymptotic answer. The paper itself continues to recommend the exact split-based test when testing $\\beta_0 = 0$.","Editorial inference: this reframes the 'p-value lottery' critique of single splits. Averaging over many splits does not create a new inference procedure; it converges to the p-value of the unsplit full-data analysis, so the lottery disappears by becoming the standard analysis.","Editorial inference: the same equivalence should apply to other two-stage protocols that use a first-stage fit to construct a covariate for a second-stage model, such as propensity scores or polygenic risk scores, whenever the estimating functions meet the smoothness and uniform-remainder conditions.","Editorial inference: the theorem requires $B$ to grow slower than $\\sqrt{n}$. At $B$ comparable to or larger than $\\sqrt{n}$, the averaged remainders may not vanish, so the common practice of taking $B$ very large with moderate $n$ may behave differently from this limit; checking that boundary empirically is an open question."],"forward_implications":["Averaging over many splits gives the same asymptotic estimator as stacking the estimating equations on the full data, so the multi-split average contains no extra information beyond an unsplit fit.","At $\\pi = 1/2$, the split-sample estimator and the full-data stacked estimator share the same influence function, so their confidence intervals are asymptotically identical.","For one split or a fixed number of splits, the covariance formulas support Wald tests and normal-based confidence intervals, with simulation coverage close to 95 percent once $n \\geq 250$.","In the physical-activity application, the score coefficient is about $-0.026$ with a 95 percent confidence interval excluding zero, so a higher score is associated with lower mortality risk even after accounting for score uncertainty.","Treating the score as estimated rather than fixed widens the intervals compared with a naive two-stage analysis that ignores score-building uncertainty."],"supporting_citations":[{"why":"Supplies the uniform control of the per-split expansion remainders as $O_P(n^{-1/2+a})$, the condition that lets averaged remainders vanish.","marker":"Carroll et al. (1996)"},{"why":"Defines the stacked estimating equations estimator whose influence function the split estimator is shown to match.","marker":"Carroll et al. (2006)"},{"why":"Provides the standard M-estimation calculus used for the single-split and fixed-split expansions.","marker":"Stefanski and Boos (2002)"},{"why":"Motivates averaging over many splits to avoid the variability and 'p-value lottery' of a single split.","marker":"Meinshausen et al. (2009)"},{"why":"Supplies the exact test of $\\beta_0 = 0$ that the paper recommends alongside the asymptotic confidence intervals.","marker":"Kravitz et al. (2019)"},{"why":"Describes the NIH-AARP cohort data used in the physical-activity scoring application.","marker":"Schatzkin et al. (2001)"},{"why":"Identifies the ratio-of-normal interval problem whose undercoverage appears in the simulations.","marker":"Fieller (1932)"}],"fun_headline_variants":["Many sample splits converge to the full-data fit","As splits grow, estimates match whole-data analysis","Averaging splits collapses to one estimating equation","Infinite splits? Same as using all the data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof relies on the approximation error from each split's statistical estimate being uniformly small across all B splits, so averaging the estimates does not accumulate error; the paper borrows this uniform bound from earlier work rather than proving it for its own two-stage logistic model.","fun_headline_variants_meta":{"raw":{"variants":["Many sample splits converge to the full-data fit","As splits grow, estimates match whole-data analysis","Averaging splits collapses to one estimating equation","Infinite splits? Same as using all the data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000152,"raw_usage":{"total_tokens":1219,"prompt_tokens":976,"completion_tokens":243,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":183}},"tokens_in":592,"tokens_out":243,"duration_ms":3223,"temperature":1.0,"reasoning_tokens":183,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:56:46.321652+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the paper's logistic two-stage model at sample sizes $n = 250$ and $n = 1000$, compute the $B$-split average for $B$ proportional to $n^{0.6}$ and to $n^{1/2}$ with $\\pi = 1/2$, and examine $\\sqrt{n}$ times the difference between the split average and the stacked-estimating-equations estimate; if this difference does not shrink to zero as $n$ grows, the uniform remainder assumption fails.","supporting_citations":[{"cited_title":"J., K \\\"u chenhoff, H., Lombard, F., and Stefanski, L","cited_arxiv_id":null,"evidence_quote":"Supplies the uniform control of the per-split expansion remainders as $O_P(n^{-1/2+a})$, the condition that lets averaged remainders vanish."},{"cited_title":"S., Carroll, R","cited_arxiv_id":null,"evidence_quote":"Supplies the exact test of $\\beta_0 = 0$ that the paper recommends alongside the asymptotic confidence intervals."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Identifies the ratio-of-normal interval problem whose undercoverage appears in the simulations."}],"review_version":1}