{"id":"34c376d0-8d12-473e-a71e-73a8d48c2fbc","arxiv_id":"2607.24324","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A diffusion pairs bootstrap recovers correct OLS variance in proportional high-dimensional linear models under score approximation, while terminal Wasserstein consistency alone does not.","lead":"Classical bootstrap methods miscalibrate variance in high-dimensional linear regression. This paper replaces empirical pairs resampling with samples from a diffusion model and proves variance consistency when the learned score is accurate enough along the diffusion path.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"Theorem 3.10 rests on simultaneous pathwise score accuracy and global Lipschitz control; that combination is not established for the MLP denoisers used in the experiments.","rationale":"The reader identified the same load-bearing issue. I did not find a more serious internal break in the route from Assumptions 3.1–3.3 to Theorem 3.10: the χ² propagation, product change of measure, Gram lower-tail transfer, conditional-law approximation, and leave-one-out variance arguments are consistently organized around the stated assumptions. The important caveat is that those assumptions are substantially stronger than a terminal Wasserstein or average score-loss guarantee. Because the paper explicitly scopes Theorem 3.10 conditionally and supplies structured attainability examples rather than claiming coverage for arbitrary neural denoisers, this does not warrant rejection. It justifies retaining a conditional verdict until either the MLP scores are linked quantitatively to Assumption 3.3 or the empirical method is shown to agree with an oracle score satisfying the assumptions.","tokens_in":91350,"tokens_out":807,"duration_ms":390595,"concrete_test":"Use the Gaussian benchmark with β=0, σ²=1, where the exact OU score is s_t(z)=-z. For increasing n and several κ=p/n, train the same MLP and compute I_n=∫ E||ŝ_{n,t}(Z_t)+Z_t||₂⁴dt by Monte Carlo. Simultaneously estimate sup_t Lip(ŝ_{n,t}-s_t) from the operator norm of its Jacobian over OU samples, corroborated by a spectral-norm or interval-bound global certificate. If I_n does not approach zero or the Lipschitz certificate exceeds a feasible L⋆, the theorem does not cover the reported MLP results; if both hold and calibration tracks the oracle linear score, the practical bridge is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Assumption 3.3 is doing more than ordinary score-estimation error control. In the entropy calculation of Lemmas D.2–D.4, the uniform bound sup_t Lip(e_{n,t}) ≤ L⋆ is used to absorb the mixed term in the density-ratio identity and obtain sup_r χ²(ρ̂_{n,r},ρ_{n,r}) = o_p(1). That terminal χ² control is then raised to the n-fold product law in Lemma D.8 to transfer the true Gram-matrix lower tail to the generated design; without it, the inverse-moment bounds needed by Theorem 3.10 can fail even when terminal W₄ error vanishes, as Examples 5.1–5.2 illustrate. Section 4 verifies the needed pair of conditions for carefully structured plug-in score classes, but the experiments train an MLP DDPM. A fixed SiLU network may be globally Lipschitz for fixed weights, but its Lipschitz constant can be large and training-dependent, and vanishing denoising loss does not imply a dimension-free uniform-in-time Lipschitz bound below L⋆. Thus the practical procedure's calibration gains remain outside the theorem. This is an applicability gap, not an internal inconsistency: the paper states Theorem 3.10 conditionally and gives non-vacuous structured examples.","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript proposes a \"diffusion pairs bootstrap\" for linear models: rather than resampling observed pairs (X_i, Y_i) from the empirical distribution, it trains a diffusion model on the joint observations and draws bootstrap samples from the terminal law of the learned reverse OU process. The motivation is the well-documented failure of classical pairs and residual bootstraps in the proportional regime p/n → κ ∈ (0,1), where they exhibit opposite variance distortions. The paper proves: (i) fixed-p variance and distributional consistency (Theorem 2.12, Corollary 2.13) under strong log-concavity of the target plus an integrated L^q score-approximation assumption with local Lipschitz control (Assumption 2.6); (ii) proportional-regime variance consistency for OLS contrasts under strongly log-concave designs with general well-conditioned covariance and Gaussian noise (Theorem 3.10), via a χ² density-ratio analysis of the Fokker–Planck equations, anti-concentration, and Gram-matrix lower-tail bounds; (iii) attainability of the score assumptions for structured classes (Gaussian, AR(1), fixed-rank perturbations, a product exponential family) with minimax risk bounds of order log p/n (Theorems 4.1, 4.3); (iv) counterexamples showing terminal W_4 convergence alone does not imply variance consistency (§5); and (v) simulations across ten design/error combinations showing improved variance and Type I error calibration relative to classical bootstraps.","tokens_in":80868,"tokens_out":5579,"duration_ms":172417,"significance":"If the results hold, the paper makes a solid contribution at the intersection of bootstrap theory and generative modeling. Its strengths are concrete: (i) a variance-consistency theorem in the proportional regime for a clearly defined new procedure, stated conditionally on explicit score assumptions rather than on unverifiable closeness of the fitted law; (ii) elementary and convincing counterexamples (§5) showing terminal W4 convergence is insufficient — a falsifiable, load-bearing clarification of why the diffusion structure matters; (iii) minimax-type upper bounds of order log p/n for integrated score risk in structured classes, establishing non-vacuity of the assumptions; and (iv) experiments benchmarked against independent theoretical quantities (Wishart 1/{n(1−κ)} and elliptical m_λ(κ)/n) rather than against quantities defined from the fitted diffusion itself. The main limitation is that the practical MLP-based procedure used in all experiments lies outside the verified assumption class; the paper is honest about this but should foreground it more.","major_comments":[{"comment":"Theorem 3.10 is conditional on Assumption 3.3, in particular (4): sup_{0≤t≤Tn} Lip(e_{n,t}) ≤ L⋆ < 1/4 with probability → 1. This uniform-in-time Lipschitz control is load-bearing — it drives the absorption in Lemmas D.2–D.4 that yields sup_r χ²(ρ̂_{n,r},ρ_{n,r}) = o_p(1), which Lemma D.8 then lifts to the M-fold product law to transfer the Gram lower tail. Section 4 verifies the pair (integrated L⁴ error, Lipschitz threshold) only for structured plug-in score classes (Gaussian, AR(1), fixed-rank, product exponential family with δ_NG small). The experiments (§7, Table 2) train an MLP DDPM: a fixed SiLU network is globally Lipschitz for fixed weights, but its constant is training-dependent, and vanishing denoising loss does not imply a dimension-free bound below L⋆. Thus the implemented procedure is outside the theorem — an applicability gap, not an inconsistency, and Remark 3.4 acknowled","section":"§3, Assumption 3.3; §4; §7"},{"comment":"Item 5 of Assumption 3.1 states that the joint distribution 'belongs to a class for which a fitted score can satisfy Assumption 3.3.' Since Theorem 3.10 already imposes Assumption 3.3 directly, this item is redundant as stated and reads as assuming part of the conclusion inside the model assumption. Please either remove it or restate it as a remark clarifying that §4 provides sufficient classes; as written it blurs what is a distributional assumption versus a score-estimation assumption.","section":"§3, Assumption 3.1(5)"},{"comment":"The variance-ratio target is Var(c_n^T β̂) = σ²_n E[c_n^T S_n^{-1} c_n], of exact order n^{-1}; the proof of Theorem 3.10 shows both the conditional-response term and the conditional-mean term match at o_p(n^{-1}). The mean-term argument (Efron–Stein plus Sherman–Morrison, pp. 86–88) requires E_{bµ}|r_n(X)|⁸ = O_p(1) and E_{bµ}‖X‖⁸₂ = O_p(p⁴_n), obtained via conditional Jensen against χ². These steps are compressed relative to the care given elsewhere; since the ratio conclusion is sensitive to getting o_p(n^{-1}) rather than O_p(n^{-1}) here, please expand the justification of (121) and the display following (122), in particular the bound |h*_1| ≤ |X*_1^T a*_{-1}| and the uniformity of the leave-one-out negative moments over the n deletions.","section":"§D.1, proof of Theorem 3.10"}],"minor_comments":[{"comment":"Abstract: 'terminalW 4' — missing space/math delimiter. Also, the phrase 'in settings not covered by our theory' would be more informative if it named the specific gap (score-error Lipschitz control for generic learned denoisers).","section":"Abstract"},{"comment":"Remark 2.9 excludes discretization error, but the experiments use a 200-step DDPM. Please add a remark connecting the continuous-time theory to the discrete sampler, or a small sensitivity check over the number of sampling steps.","section":"§2, Remark 2.9; §A.3"},{"comment":"For the i.i.d. Laplace design the benchmark 1/{n(1−κ)} is used (§A.5: 'the same first-order benchmark'). Please justify that c^T(X^TX)^{-1}c concentrates at 1/(n(1−κ)) for i.i.d. non-Gaussian entries, with a citation, since this is not the Wishart calculation used for the Gaussian case.","section":"§7.2 and §A.5"},{"comment":"The jackknife uses a Gaussian interval while all bootstrap procedures use percentile intervals (§A.4). Please note this asymmetry, as it affects the Type I error comparison across methods.","section":"§A.4"},{"comment":"Notation: λ is overloaded — radial multiplier λ_i in elliptical designs, spike parameter in the fixed-rank class, threshold λ_n, and the constant λ⋆ of (3). Table 6 partially flags this; please disambiguate in the text where collisions occur (e.g., §6 and §7.2 vs. §4.1).","section":"§1.4, Tables 4–6"},{"comment":"Figure 1 appears in §1 before any method definitions; the caption should at least name the procedures plotted (pairs, residual, jackknife, smoothed pairs, diffusion pairs, diffusion residual) and define the variance-ratio denominator.","section":"§1, Figure 1"},{"comment":"Theorem 3.10 requires Gaussian regression noise (Assumption 3.1(1)) while the experiments include Laplace errors; Remark 3.6 sketches an extension to strongly log-concave noise. Please state explicitly in §7 that the Laplace-error experiments are outside the proved result, to avoid the reader over-reading 'generally improves Type I error calibration.'","section":"§3, Remark 3.6; §7"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript cites several very recent arXiv preprints, most notably Tran, Ye, Ding and Han (2026), \"Generative modeling for the bootstrap,\" which appears to be closely concurrent work. The overlap in scope is substantial enough that the editor may wish to confirm the authors' novelty positioning relative to that preprint. Also worth noting: the gap between the theory (structured plug-in scores) and the implemented method (MLP DDPM) is the paper's soft underbelly; the authors acknowledge it in Remark 3.4, but a journal readership will likely press on it, so the requested strengthening of the discussion is in the authors' own interest."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real contribution here is not “use a diffusion model for bootstrap.” It is the analytic point that terminal W4 is not enough for OLS variance, plus a workable pathwise-score route that controls Gram lower tails. The §5 mixture counterexamples are elementary and decisive: vanishing W4 with a rare near-degenerate design component blows up inverse Gram moments. That lesson is worth keeping even if you never train a DDPM.\n\nWhat they do well is scoped carefully. Fixed-p and proportional variance theorems (2.12, 3.10) are built from coupling, Fokker–Planck regularity, χ² density-ratio ODEs, and anti-concentration, not hand-waving. Residual diffusion is correctly dismissed: fitted residuals are already variance-shrunk in the proportional regime, so learning their law cannot fix the geometry. Experiments (ten design/error combinations, fixed architecture, n=500) show pairs and jackknife overshoot, residual undershoots, and diffusion pairs largely restores the Wishart/elliptical benchmarks—including Laplace and some elliptical cases outside the theory. Citations to El Karoui–Purdom and the generative-bootstrap line are honest.\n\nThe soft spot is exactly the stress-test note, and it is an applicability gap rather than a hole in the logic. Theorem 3.10 needs integrated L4 score error and a uniform Lip(e)≤L⋆ bound along the whole OU path so the entropy argument yields terminal χ² and product-law tail transfer. Section 4 proves that package for structured plug-ins (Gaussian, AR(1), fixed-rank, product exponential) at rate log p/n. The experiments use an MLP DDPM; vanishing denoising loss does not give you that dimension-free Lipschitz control. The paper states the theorem conditionally and does not pretend otherwise. Still, the practical calibration claim is empirical, not covered.\n\nMinor: no public code in the manuscript, discretization error ignored, Gaussian noise in the high-d theorem (they flag this). None of that sinks the main argument.\n\nThis is for people who care about high-dimensional bootstrap geometry and generative models as distribution estimators with pathwise control—not a drop-in software paper. I would send it to referees. Engage with the theory and the counterexamples; treat the MLP results as promising evidence, not theorem coverage.","headline":"Solid theory paper: pathwise score control for diffusion pairs bootstrap, clean W4 counterexamples, and real calibration gains; the MLP experiments sit outside the Lipschitz/score assumption.","tokens_in":80066,"tokens_out":566,"would_cite":true,"duration_ms":15195,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F40","62J05","62G09"],"pacs":[],"model":"grok-4.5","headline":"A diffusion pairs bootstrap can recover the right OLS variance in high dimensions when classical pairs and residual bootstraps systematically miscalibrate.","keywords":["bootstrap","diffusion models","high-dimensional inference","linear models","score approximation","pairs bootstrap","variance calibration","proportional asymptotics"],"falsifier":"Train the paper’s diffusion pairs procedure under a strongly log-concave proportional design where the integrated pathwise score error and Lipschitz bound can be checked or forced to fail; if variance ratios stay near one when the score condition holds and blow up or stay biased when it fails—while W4 remains small—the central claim is supported or refuted.","tokens_in":79983,"feed_emoji":"📉","tokens_out":1031,"duration_ms":21990,"temperature":0.7,"pith_summary":"In linear regression when the number of covariates is a fixed fraction of the sample size, the usual pairs bootstrap tends to overstate uncertainty and the residual bootstrap tends to understate it. This paper argues that the fix is not better resampling of the observed data, but replacing the empirical joint law of (X, Y) with samples from a diffusion model trained on those pairs. Under score approximation along the full noise path—not merely closeness of the final generated distribution—the bootstrap variance of ordinary least-squares contrasts matches the true variance, in both fixed dimension and the proportional high-dimensional regime. Terminal Wasserstein closeness alone is shown to be insufficient, because rare nearly singular designs can wreck inverse Gram moments while barely moving average transport cost. Simulations across Gaussian and non-Gaussian designs suggest the method improves variance and Type I error calibration even outside the theorems’ assumptions.","feed_headline":"Diffusion bootstrap fixes high-d regression variance","feed_subtitle":"Learning the joint law beats pairs and residual resampling when p grows with n","key_machinery":"Pathwise score approximation along the OU reverse SDE/Fokker–Planck flow: integrated L4 score error plus a uniform Lipschitz bound on the score error, used to get χ² control of generated versus true laws and polynomial lower tails on the generated Gram matrix (hence inverse-moment bounds for bootstrap OLS variance).","core_discovery":"Diffusion pairs bootstrap—drawing bootstrap pairs from a learned reverse diffusion on the joint observations—achieves variance consistency for OLS linear contrasts when the learned score tracks the true score along the Ornstein–Uhlenbeck path with vanishing integrated fourth-moment error and controlled Lipschitz score error. That pathwise score control yields density-ratio and Gram lower-tail bounds that terminal W4 convergence cannot supply, and the result covers strongly log-concave designs with well-conditioned covariance in the proportional regime p/n → κ ∈ (0,1).","pith_inferences":["If pathwise score control is the real bottleneck, diagnostics that monitor reverse-drift stability or generated Gram lower tails may be more informative than terminal sample-quality scores alone.","The counterexamples suggest a broader template: any generative bootstrap whose samples can place vanishing mass on nearly singular designs may fail variance calibration despite good average distributional metrics.","Extending the theory beyond Gaussian noise and strong log-concavity would likely hinge on replacing the uniform Poincaré/log-Sobolev and one-dimensional small-ball inputs, not on changing the pairs-versus-residual distinction.","Practitioners using diffusion bootstrap for other M-estimators would still need inverse-moment or anti-concentration control specific to those estimating equations."],"forward_implications":["Classical pairs bootstrap conservatism and residual bootstrap anti-conservatism in the proportional regime need not be inevitable if the joint law is learned rather than resampled empirically.","Bootstrap validity for OLS variance can be reduced to score estimation quality along the diffusion path, separating generative accuracy from Gram geometry control.","n-atomic estimators (including the empirical measure) cannot be W4-consistent for high-dimensional Gaussians, while structured scores can still be estimated at rate about log(p)/n.","Diffusion residual bootstrap is not a substitute: fitted residuals are already variance-shrunk by high-dimensional projection, so learning their law cannot restore the true error scale.","The same architecture can improve calibration for i.i.d. non-Gaussian designs and reduce (but not always erase) distortion under heterogeneous elliptical designs."],"fun_headline_variants":["Diffusion pairs bootstrap calibrates variance in high-d linear models","Score-tracked diffusion bootstrap yields consistent OLS variance","Learned joint diffusion law fixes high-d bootstrap variance","Pathwise score control gives variance-consistent diffusion bootstrap","Diffusion bootstrap improves Type I error over pairs and residual"],"cache_read_input_tokens":65664,"weakest_assumption_plain":"The fitted diffusion score must stay close to the true score along the entire noise path, with a uniform Lipschitz bound on the error; the variance theorems stand or fall with that condition, which is proven only for structured model classes, not for arbitrary neural denoisers.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion pairs bootstrap calibrates variance in high-d linear models","Score-tracked diffusion bootstrap yields consistent OLS variance","Learned joint diffusion law fixes high-d bootstrap variance","Pathwise score control gives variance-consistent diffusion bootstrap","Diffusion bootstrap improves Type I error over pairs and residual"]},"model":"grok-4.5","effort":"low","cost_usd":0.003819,"raw_usage":{"total_tokens":1151,"prompt_tokens":663,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":38188000,"prompt_tokens_details":{"text_tokens":663,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":427,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":663,"tokens_out":61,"duration_ms":6716,"temperature":1.0,"reasoning_tokens":427,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T18:04:28.581527+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train the paper’s diffusion pairs procedure under a strongly log-concave proportional design where the integrated pathwise score error and Lipschitz bound can be checked or forced to fail; if variance ratios stay near one when the score condition holds and blow up or stay biased when it fails—while W4 remains small—the central claim is supported or refuted.","supporting_citations":[],"review_version":1}