{"id":"2c345d3a-a8e5-48df-b31d-c62087ee3582","arxiv_id":"2506.20972","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A scaled wild bootstrap is proven asymptotically valid for inference on a regression coefficient when the number of covariates is of the same order as the sample size and errors are heteroskedastic.","lead":"This paper modifies the wild bootstrap for linear regressions with many covariates and heteroskedastic errors, adding a scale factor so the bootstrap t-statistic remains valid even when the number of controls is close to the sample size. Simulations show the proposed bootstrap controls rejection rates near the nominal level, unlike normal-based tests that over-reject.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1's proof invokes a multiplier CLT without verifying Lindeberg; the asserted bound max_i sigma*_i^2 = O_p(1) is not implied by Assumptions 1-5, so the main theorem is not established as stated.","rationale":"The reader's weakest-assumption analysis points to Assumption 4 and the Monte Carlo design, and that concern is real: the stated Bernoulli-dummy DGP almost surely produces a zero diagonal entry of the residual-maker matrix, making the procedure undefined. I agree that this must be fixed. However, the most load-bearing threat to the central claim, Theorem 1, is the unverified conditional Lindeberg condition in the bootstrap CLT step. The proof asserts convergence by invoking a multiplier CLT but does not verify its conditions; the intermediate bound max_i σ_i*^2 = O_p(1) is not a consequence of the stated assumptions, since only fourth moments are assumed for the errors. This is not a matter of realism or simulation cleanliness; it is a missing step in the proof of the theorem itself. Because the paper's construction is plausible and may be valid under a strengthened condition, I would not reject or mark the paper as unverified outright, but I would keep the reader's CONDITIONAL verdict: the authors should add and verify the Lindeberg condition (or an assumption that implies it) and correct the simulation DGP so that Assumption 4 is satisfied.","tokens_in":11499,"tokens_out":31932,"duration_ms":349506,"concrete_test":"Re-derive the conditional CLT step for t_n^*: write t_n^* = Σ_i c_i ω_i^* + o_p*(1) with c_i = vhat_i ũ_i(β0) / (Σ_k vhat_k^2 ũ_k^2(β0))^{1/2}, and check whether Assumptions 1-5 imply max_i c_i^2 = o_p(1). If the implication is not proved, construct a sequence satisfying Assumptions 1-5 with heavy-tailed x_i (e.g., Pareto tail index 5), u_i = x_i + ε_i with ε_i ~ N(0,1), and W chosen so the diagonal entries of M_n are bounded away from zero (e.g., a balanced group structure with q/n = 1/2). Simulate the null rejection rate of the proposed bootstrap at n = 200 and n = 2000; if the rejection rate does not approach the nominal level, the missing Lindeberg verification is material. Separately, re-run the Section 5 design with the stated independent Bernoulli dummies and record the fraction of replications where min_i (M_n)_ii = 0, which would show the reported tables are non-reproducible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central step of the proof of Theorem 1 is the claim, in Appendix B, that the bootstrap t-statistic converges conditionally to N(0,1) 'by Lemma 2.9.5 of van der Vaart and Wellner (1996)'. Up to o_p* terms, t_n^* is a weighted sum of the i.i.d. bootstrap weights, with coefficients c_i proportional to vhat_i a_n(β0) ũ_i(β0). After the final scaling, c_i^2 = vhat_i^2 ũ_i^2(β0) / Σ_k vhat_k^2 ũ_k^2(β0). A conditional multiplier CLT requires the Lindeberg condition max_i c_i^2 = o_p(1). No such condition is stated or verified. The proof instead relies on the bound max_i σ_i*^2 = O_p(1), where σ_i*^2 = a_n^2(β0) ũ_i^2(β0). This bound is not implied by Assumptions 1-5: with only fourth moments on ε_i (Assumption 2), max_i |ũ_i| can be of order n^{1/4} in probability, so max_i σ_i*^2 can grow like n^{1/2}. Even combined with max_i vhat_i^2 = o_p(n) (Assumption 3), the ratio max_i vhat_i^2 ũ_i^2 / Σ_k vhat_k^2 ũ_k^2 need not vanish. If that ratio fails to vanish, the bootstrap distribution is driven by a single weight and need not be standard normal, so Theorem 1 does not follow from the stated assumptions. Separately, the Monte Carlo DGP in Section 5 with Bernoulli dummies also violates Assumption 4: singleton dummy cells make (M_n)_ii = 0, so the procedure is undefined; this is a valid but secondary concern.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a modification of the wild bootstrap for inference on a scalar regression coefficient when the number of controls q_n is a non-negligible fraction of the sample size. The modification is an adjustment factor a_n(β0) that multiplies null-imposed residuals, following the construction of Jochmans (2022). The main theoretical result, Theorem 1, claims that under Assumptions 1-5 the conditional bootstrap distribution of the percentile-t statistic approximates the null distribution uniformly. Monte Carlo simulations compare HC0, HCK, HCA and the proposed wild bootstrap with Gaussian and Rademacher weights for n=100 and q_n/n up to 0.9, and report that the bootstrap controls size better than normal-based methods.","tokens_in":11943,"tokens_out":14639,"duration_ms":160889,"significance":"If the theorem is correct, the paper fills a genuine gap: no previously proven-valid bootstrap is available for linear regressions with both q_n/n not tending to zero and heteroskedasticity. The proposed adjustment factor is a deterministic function of the data under the null and involves no fitted parameters, which is a useful feature. The simulation evidence is suggestive, though the reported Monte Carlo design is problematic. The main concern is the proof of the bootstrap central limit theorem, which is asserted rather than verified at the key step.","major_comments":[{"comment":"The conditional multiplier CLT is not established. The proof invokes Lemma 2.9.5 of van der Vaart and Wellner (1996) to conclude that \\bar t_n^* converges conditionally to N(0,1), but it does not verify the Lindeberg condition for the weighted sum. Conditional on the data, \\bar t_n^* is a weighted sum of the independent bootstrap weights with coefficients proportional to \\hat v_i a_n(β0) \\tilde u_i(β0); after normalization the relevant quantity is max_i \\hat v_i^2 \\tilde u_i^2 / \\sum_k \\hat v_k^2 \\tilde u_k^2. The proof instead uses the asserted bound max_i σ_i^{*2}=O_p(1), where σ_i^{*2}=a_n^2(β0)\\tilde u_i^2(β0). That bound is not implied by Assumptions 1-5: with only \\max_i E[ε_i^4|X_n,W_n]=O_p(1) in Assumption 2, max_i |\\tilde u_i| can be of order n^{1/4} in probability, so max_i σ_i^{*2} need not be O_p(1). The earlier variance bound for the first term of Eq. (4), n^{-2}\\sum \\hat v_i^4(a_n^4 E[ω_i^{*4}]\\tilde u_i^4 - a_n^4\\tilde u_i^4)=o_p(1), also relies on \\tilde u_i^4 being controlled at the right rate, which is not stated among the assumptions. Thus Theorem 1 is not established as written; the proof needs either a direct verification of the Lyapunov/Lindeberg condition or an additional assumption that delivers max_i \\tilde u_i^2=o_p(n) and max_i \\hat v_i^2 \\tilde u_i^2/\\hatΣ_n(β0)=o_p(1).","section":"Appendix B, paragraph after Eq. (7)"},{"comment":"The Monte Carlo DGP violates Assumptions 2 and 4 with high probability. With q_n-1 independent Bernoulli(0.02) dummy variables and n=100, at q_n/n=0.9 there are 89 dummies; the probability that at least one dummy column is identically zero is essentially one, so W has rank deficiency and Assumption 2 fails. Moreover, singleton dummy cells are common, producing (M_n)_ii=0 and making \\acute u_i = \\tilde u_i/(M_n)_ii undefined, which violates Assumption 4. The paper does not state how such replications were treated. The simulation evidence should be recomputed with a design that enforces full rank and min_i(M_n)_ii>0 (for example, by redrawing W until those conditions hold), and the effective number of replications should be reported.","section":"Section 5, Tables 1-3"}],"minor_comments":[{"comment":"The notation \\acuteΣ_n(β0) and \\hatΣ_n(β0) is used in the definition of a_n(β0) before the reader is told that \\acuteΣ_n(β0) is the null-imposed analogue of the meat in \\acuteΩ_n; consider defining both quantities directly after the display.","section":"Section 3, Eq. (3)"},{"comment":"The proof mixes o_{p*}(1) and o_p(1) for bootstrap quantities; for readability, state explicitly in each display whether the high-probability statement is under P or under P^* in probability.","section":"Appendix B, displays after Eq. (4)"},{"comment":"In all three tables the HC0 and HCK rejection frequencies are identical at q_n/n=0.9 (0.581, 0.574 and 0.583 respectively); this is likely a symptom of rank-deficient dummy columns, and the tables or the surrounding text should explain this coincidence.","section":"Tables 1-3, q_n/n=0.9 row"}],"recommendation":"major_revision","confidential_remarks":"The proposed method is simple and the claimed gap is real, so the paper is worth pursuing. However, the proof of Theorem 1 is not self-contained at the bootstrap CLT step, and the simulation design as described cannot be run without additional rules for rank-deficient W. I would like to see a revised version in which the Lindeberg condition is verified under stated assumptions, or the assumptions are strengthened appropriately, and the Monte Carlo DGP is corrected. If those points are resolved, I would support publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look, but the main theorem has a real hole. The adjustment factor is a clean idea, and as far as the cited literature goes, this is the first proposed valid wild bootstrap for linear regressions with q_n/n not vanishing under heteroskedasticity. That alone makes it interesting. The Monte Carlo results are also striking: the bootstrap controls size where HC0, HCK, and HCA over-reject badly, and the panel extension is sensible.\n\nThe problem is the proof of Theorem 1. The conditional multiplier CLT from van der Vaart and Wellner (Lemma 2.9.5) requires a Lindeberg condition: the largest squared coefficient in the linear form, relative to the sum of squares, must vanish. The proof instead asserts max_i sigma_i*^2 = O_p(1) and uses that to bound remainder terms. That bound is not implied by Assumptions 1–5; with only fourth moments on the errors, max_i |u_tilde_i| can be of order n^{1/4}, so max_i sigma_i*^2 can grow. More importantly, after the adjustment factor cancels, the effective coefficients are proportional to vhat_i * u_tilde_i, and the Lindeberg ratio max_i (vhat_i^2 u_tilde_i^2) / sum_k (vhat_k^2 u_tilde_k^2) need not vanish. The proof also uses max_i u_tilde_i^2 = O_p(1) in an earlier step, which is a stronger claim than the assumptions justify. So the central theorem is not established as stated.\n\nThere is also a simulation design issue. The Monte Carlo uses Bernoulli dummies with pi = 0.01 or 0.02 at n = 100. That means a nontrivial fraction of replications will have a dummy that equals 1 for exactly one observation, making (M_n)_ii = 0 and the whole procedure undefined, since division by (M_n)_ii appears in Steps 2 and 4. The paper never says how those cases are handled. If they are simply included, the reported results are not evaluating the proposed method. This is a fixable reporting problem, but it is a serious gap in the evidence as presented.\n\nOn balance: the idea is genuinely new and the direction is right, but the proof gap is load-bearing. The paper deserves a referee, not because it is correct, but because the question is important and the adjustment factor is a credibly promising construction. A serious referee would need to see a verified Lindeberg condition or a counterexample, and the simulations should either use a DGP that satisfies Assumption 4 or explicitly handle the degenerate cases.","headline":"Promising enough for a referee, but Theorem 1 is not proven as stated: the multiplier CLT step lacks a Lindeberg condition, and the simulations use a Bernoulli-dummy design that can make the procedure undefined.","tokens_in":12387,"tokens_out":4755,"would_cite":false,"duration_ms":49222,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F40","62J05","62E20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that a modification to the wild bootstrap—multiplying null-restricted residuals by an adjustment factor built from a cross-fit variance estimator—gives asymptotically valid t-tests in linear regressions with many…","keywords":["wild bootstrap","many covariates","heteroskedasticity","linear regression","bootstrap validity","high-dimensional inference","residual-maker matrix","cross-fit variance estimator"],"falsifier":"Set $n=100$, $q_n=90$, and include among the controls a dummy variable that equals 1 for exactly one observation; at that observation $(M_n)_{ii}=0$, so the bootstrap estimator of the variance divides by zero and the method cannot be computed. Running this design across many Monte Carlo draws and recording the minimum of $(M_n)_{ii}$ would show Assumption 4 is violated, and the claimed uniform approximation cannot hold there.","tokens_in":11342,"feed_emoji":"📊","tokens_out":10095,"duration_ms":94914,"temperature":0.7,"pith_summary":"This paper asks whether the wild bootstrap can be made reliable for inference on one coefficient in a linear regression when the number of control variables is a large fraction of the sample size and the errors are heteroskedastic. Standard wild bootstrap is known to fail in this regime, and normal-theory variance estimators over-reject badly. The proposed fix multiplies the null-restricted residuals by an adjustment factor before drawing bootstrap weights, and the main theorem shows the bootstrap t-statistic and the true t-statistic have the same limiting distribution uniformly over the real line. If right, practitioners get a straightforward bootstrap that holds its size even when controls are 90 percent of the observations, which the simulations confirm for n=100.","feed_headline":"Bootstrap fix keeps t-tests valid as covariates near sample size","feed_subtitle":"One adjustment factor keeps null-rejection rates near nominal even with 90 controls per 100 observations.","key_machinery":"The load-bearing object is the adjustment factor $a_n(\\beta_0)$, together with the residual-maker matrix $M_n = I_n - W(W'W)^{-1}W'$ whose diagonal entries the bootstrap divides by. The factor is chosen so that the bootstrap meat $n^{-1}\\sum_i \\hat v_i^2 y^*_i \\acute u^*_i$ approximates the null-imposed cross-fit quantity $\\tilde\\Sigma_n(\\beta_0)$; the $\\max\\{\\cdot,1/n\\}$ guard keeps the square root well defined when the estimated variance is not positive. The proof decomposes the difference between bootstrap and sample variance estimators into three remainder terms and shows each is conditionally $o_{p^*}(1)$ using the conditional multiplier central limit theorem, the boundedness of $(\\min_i (M_n)_{ii})^{-1}$, and a small-maximum fitted-value condition on the null-restricted predictions.","core_discovery":"The central claim is Theorem 1: under the paper's Assumptions 1–5, when the null hypothesis is true, $\\sup_c |F_n(c)-F^*_n(c)| = o_p(1)$, so the modified wild bootstrap t-statistic is asymptotically valid even when $q_n/n$ does not shrink to zero. The modification is to generate bootstrap errors as $u^*_i = a_n(\\beta_0)\\omega^*_i \\tilde u_i(\\beta_0)$, where $\\tilde u_i$ are null-restricted residuals and $a_n(\\beta_0)=\\sqrt{\\max\\{\\tilde\\Sigma_n(\\beta_0),1/n\\}/\\hat\\Sigma_n(\\beta_0)}$. This adjustment makes the bootstrap version of the cross-fit-style variance estimator line up with its sample counterpart, and the proof shows that conditional on the data the bootstrap t-statistic converges to a standard normal in probability while the actual t-statistic does the same. A direct consequence is that percentile-t bootstrap tests control size under many covariates and heteroskedasticity, a combination for which the paper argues no proven-valid bootstrap existed before.","pith_inferences":["A likely practical extension is to clusters: applying the same adjustment factor within cluster blocks would be natural, but the theorem's independence assumption does not cover cluster dependence, so that extension would need fresh proof.","The guard $\\max\\{\\tilde\\Sigma_n,1/n\\}$ suggests a direct diagnostic: practitioners can compare the adjusted and unadjusted variance estimates, and a large gap indicates the regime where normal critical values are unreliable and the bootstrap adjustment matters.","The proof's reliance on $(\\min_i (M_n)_{ii})^{-1}=O_p(1)$ implies the method should be used with caution when controls include singleton indicators or nearly saturated dummies; checking the diagonal of $M_n$ before running the bootstrap would flag unstable cases.","The paper's Monte Carlo designs use Bernoulli dummies with small success probabilities, so a natural stress test is to record the minimum diagonal of $M_n$ in those same designs to see whether Assumption 4 is actually met or whether finite-sample performance is carried by the $1/n$ guard."],"forward_implications":["For empirical work, the procedure gives t-tests and confidence intervals for a treatment effect that remain correctly sized when controls are numerous and errors are heteroskedastic, with no need to know whether $q_n/n$ is small or large.","In the simulations with $n=100$, the modified bootstrap holds null rejection frequencies between roughly 0.031 and 0.051 across $q_n/n = 0.1$ to $0.9$, while the best normal-based estimator reaches 0.172 and HC0/HCK reach 0.581.","When $q_n/n \\to 0$ the adjustment factor converges to 1, so the procedure reduces to the standard null-imposed wild bootstrap.","The percentile-t bootstrap is recommended over the percentile version, and the paper sketches a score-bootstrap extension for vector coefficients and linear hypotheses.","The panel fixed-effects simulations show the same size control when the many covariates are group dummies, with bootstrap rejection frequencies near nominal even at 50 groups."],"supporting_citations":[{"why":"Supplies the many-covariates asymptotic framework and the heteroskedasticity-robust variance estimator whose consistency requires $q_n/n$ to shrink; its assumptions are reused and its simulations provide the normal-theory benchmark.","marker":"Cattaneo et al. (2018b)"},{"why":"Proposes the cross-fit variance estimator whose construction the bootstrap adjustment factor follows; the proof of Theorem 1 reuses its decomposition steps.","marker":"Jochmans (2022)"},{"why":"Establishes that the standard wild bootstrap requires $q_n^{1+\\delta}/n$ to vanish, the limitation the modification is designed to overcome.","marker":"Mammen (1993)"},{"why":"Documents bootstrap inconsistency in high-dimensional linear models, motivating the need for a modified bootstrap procedure.","marker":"El Karoui and Purdom (2018)"},{"why":"Supplies the conditional multiplier central limit theorem (Lemma 2.9.5) used to show the bootstrap t-statistic converges to a standard normal.","marker":"van der Vaart and Wellner (1996)"}],"fun_headline_variants":["Bootstrap tweak valid even with as many controls as observations","Modified wild bootstrap corrects size with many covariates","Bootstrap fix tames high-dimensional controls","Bootstrap method proven valid for many-covariate regressions","New wild bootstrap stays accurate with many controls"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The procedure divides by the diagonal entries of the residual-maker matrix and requires every one of them to stay bounded away from zero; if any observation is completely or almost completely pinned down by the controls, the bootstrap variance estimate is undefined or unstable and the theorem's conditions fail.","fun_headline_variants_meta":{"raw":{"variants":["Bootstrap tweak valid even with as many controls as observations","Modified wild bootstrap corrects size with many covariates","Bootstrap fix tames high-dimensional controls","Bootstrap method proven valid for many-covariate regressions","New wild bootstrap stays accurate with many controls"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000535,"raw_usage":{"total_tokens":2515,"prompt_tokens":829,"completion_tokens":1686,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":445,"completion_tokens_details":{"reasoning_tokens":1611}},"tokens_in":445,"tokens_out":1686,"duration_ms":13154,"temperature":1.0,"reasoning_tokens":1611,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:37:18.391086+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Set $n=100$, $q_n=90$, and include among the controls a dummy variable that equals 1 for exactly one observation; at that observation $(M_n)_{ii}=0$, so the bootstrap estimator of the variance divides by zero and the method cannot be computed. Running this design across many Monte Carlo draws and recording the minimum of $(M_n)_{ii}$ would show Assumption 4 is violated, and the claimed uniform approximation cannot hold there.","supporting_citations":[{"cited_title":"(2022): Heteroscedasticity-robust inference in linear regression models with many covariates, Journal of the American Statistical Association, 117, 887--896","cited_arxiv_id":null,"evidence_quote":"Proposes the cross-fit variance estimator whose construction the bootstrap adjustment factor follows; the proof of Theorem 1 reuses its decomposition steps."},{"cited_title":"(1993): Bootstrap and wild bootstrap for high dimensional linear models, The Annals of Statistics, 21, 255--285","cited_arxiv_id":null,"evidence_quote":"Establishes that the standard wild bootstrap requires $q_n^{1+\\delta}/n$ to vanish, the limitation the modification is designed to overcome."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents bootstrap inconsistency in high-dimensional linear models, motivating the need for a modified bootstrap procedure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the conditional multiplier central limit theorem (Lemma 2.9.5) used to show the bootstrap t-statistic converges to a standard normal."}],"review_version":1}