{"id":"2053b1df-af32-45f5-95d9-f0fc3ff084a3","arxiv_id":"2412.17181","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Matching-based ATE estimators have explicit non-asymptotic Gaussian and multiplier-bootstrap approximation bounds with rates in sample size, number of matches, and treatment balance.","lead":"This paper proves explicit error bounds for how quickly matching-based estimates of treatment effects become bell-curve shaped, and gives a multiplier bootstrap method with finite-sample guarantees. It matters because it converts asymptotic confidence intervals in causal inference into non-asymptotic ones that account for the number of matches and treatment balance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 5.1 and Theorem 6.1 both inherit their rates from Lemma A.4, whose proof imports non-asymptotic inequalities from Lin et al. without reproducing them; this is the clearest place where a hidden assumption could break the stated B3 term and the bootstrap high-probability bound.","rationale":"I reviewed the proof route: Theorem 7.1 (Malliavin-Stein) gives Gaussian approximation for E_n; Lemma 7.1 scales to sigma via a variance upper bound; Lemma A.4 is the only place where the non-asymptotic density-ratio rate is imported from prior work. The B3 term in Theorem 5.1 is exactly the variance approximation error, and Theorem 6.1 uses B3 both in the bound and in the probability of success. The reader's weakest_assumption identifies the same point. I also checked for internal inconsistencies in the score-function representation (7.5) and the stabilization radius; the score representation is valid, and the tail bound is plausible up to constants that can depend on the fixed support X. The bootstrap proof is a three-step comparison of conditional Gaussian variances and uses Theorem 5.1 as a bridge; no circularity is apparent. The main missing piece is the detailed verification of the imported inequalities in Lemma A.4. Because the rest of the argument is coherent, I would not change the reader's CONDITIONAL verdict; the stated rates should be treated as conditional on Lemma A.4 being fully non-asymptotically valid. An independent re-derivation of Lemma A.4 would settle the concern.","tokens_in":54419,"tokens_out":15724,"duration_ms":147218,"concrete_test":"Independently re-derive Lemma A.4 by spelling out the cited inequalities from Lin et al. (2023b, Proof of Theorems B.3/B.4, S3.30-S3.37) with all constants, on the event I = {|n0-np0| < np0/2}, and verify each step under Assumptions A/B. In particular, check whether the lower-bound term Cp(1-rratio)(n/M)e^{M-r0 n p0 - M log M + M log(r0 n p0)} and the variance bound can be combined without an extra eta^{-1} or M log n/(n eta) factor; if such a factor appears, B3 and the Theorem 6.1 probability statement must be revised accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central finite-sample claims (Theorem 5.1 and the bootstrap Theorem 6.1) rest on the variance approximation in Lemma 7.1, whose input is the density-ratio estimation bound in Lemma A.4. The B3 term is a direct component of the Theorem 5.1 bound; the bootstrap bound and its 'with probability at least 1 - 16(B3 and 1)' statement also scale with B3. Lemma A.4 is therefore load-bearing. The proof of Lemma A.4 is not self-contained: it says it adapts Lin et al. (2023a, proofs of Theorems B.3 and B.4) by replacing asymptotic o(n^{-\\gamma}) simplifications with 'the first inequalities' from those proofs, but it does not reproduce those inequalities or verify that they hold non-asymptotically under the stated assumptions. For example, the lower-bound step replaces an asymptotic term by Cp(1-rratio)(n/M)e^{M-r0 n p0 - M log M + M log(r0 n p0)} 'according to the first inequality' in the cited proof, and the variance decomposition uses cited S3.34/S3.36 with the expectation bounded by the modified (A.5). If any of the imported inequalities implicitly uses p0 fixed, delta small relative to g_min, or an ordering such as n0 tending to infinity before M diverges, the resulting rate could pick up an extra factor in eta or M log n/(n eta). That would change B3, hence the Corollary 5.1 rate and the high-probability statement in Theorem 6.1. I do not see an internal contradiction, but the missing details make the central rate a conditional rather than fully verified claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops non-asymptotic Gaussian approximation bounds for bias-corrected covariate- and rank-based matching estimators of the average treatment effect. The main covariate result, Theorem 5.1, bounds the Kolmogorov distance between sqrt(n)(tau_hat_M^bc - tau) and a centered normal with variance sigma^2 by an explicit sum B1+B2+B3 depending on the number of matches M, the balance parameter eta, the dimension m, and the moment parameter p. Theorem 5.2 extends the result to phi-transformed rank matching, and Theorem 6.1 gives analogous bounds for a multiplier bootstrap procedure, conditionally on the data with high probability. The proof strategy is a three-step decomposition: stabilization plus Malliavin-Stein for the main term, a quantitative variance approximation, and a bias bound; the technical core is the non-asymptotic density-ratio estimation bound in Lemma A.4.","tokens_in":1655,"tokens_out":1665,"duration_ms":85770,"significance":"If the stated theorems hold, the paper makes a substantial contribution: it provides the first finite-sample Gaussian and bootstrap approximation rates for matching ATE estimators with a diverging number of matches, makes the dependence on treatment balance explicit, and offers a data-driven bootstrap with quantifiable accuracy. The stabilization viewpoint and the explicit variance approximation are valuable technical innovations. However, the central rates in Theorems 5.1 and 6.1 rest on Lemma A.4, whose proof imports non-asymptotic inequalities from Lin et al. (2023a) without reproducing or fully verifying them. This is a load-bearing gap, not a presentation issue. The proof of the rank-based bootstrap assertion is also deferred rather than supplied. I therefore view the manuscript as promising but not yet fully verified in its current form.","major_comments":[{"comment":"Lemma A.4 is load-bearing: its density-ratio bound enters the B3 term of Theorem 5.1, the variance approximation in Lemma 7.1, and the probability statement 1 - min(16B3,1) in Theorem 6.1. The proof, however, is not self-contained. After equation (A.3), the text says it adapts Lin et al. (2023a, proofs of Theorems B.3 and B.4) by replacing asymptotic o(n^{-gamma}) simplifications with 'the first inequalities' from those proofs, referring to S3.30-S3.31 and S3.34-S3.37, but it does not reproduce those inequalities or verify that they hold under the present non-asymptotic settings, where M can diverge, eta may decay, and p0 is not assumed fixed. If any of the inherited inequalities implicitly uses an ordering such as n0 going to infinity before M diverges, or delta small relative to g_min, the resulting rate could acquire extra factors in eta or M log n/(n eta), changing B3 and hence the Corollary 5.1 rate and the bootstrap high-probability bound. The authors should either give a complete proof of the non-asymptotic density-ratio bound or state and prove the imported inequalities under the assumptions of Theorem 5.1.","section":"Appendix A, Lemma A.4"},{"comment":"The rank-based assertion of Theorem 6.1 is dismissed at the end of Appendix C as following 'mutatis mutandis'. This is not a routine adaptation. The rank-based proof must handle the additional term Delta E_phi,n, the estimation error of phi-hat appearing in Lemmas B.1 and B.2, and the different variance lower bound L(mu, mu-hat, phi, phi-hat, n). In particular, the bound on sqrt(n) E|Delta E_phi,n| in equation (E.11) is obtained only after choosing epsilon, epsilon_1 and delta of the order (M/(n eta))^{1/m1}, and no verification is given that the Cattaneo et al. (2023) inequalities used for that choice remain valid in the non-asymptotic regime of Theorem 5.2. The omitted proof should be supplied or replaced with a detailed verification of the three bootstrap steps in the rank-based setting.","section":"Appendix C, Theorem 6.1, rank-based case"},{"comment":"The bound for sqrt(n) E|Delta E_phi,n| in Lemma B.2 and (E.11) contains factors (delta epsilon_1)^{-m} multiplying the estimation error of phi-hat, and the free parameters epsilon, epsilon_1, delta are subsequently collapsed into a single choice delta ~ (M/(n eta))^{1/m1}. This step needs an explicit infimum over delta and a demonstration that the constraints delta >= (M/(n eta))^{1/m1}, epsilon, epsilon_1 of order delta are compatible with the small-delta assumptions in the cited work. Without such a verification, the term reported in Theorem 5.2's B6 may be optimistic, and the rank-based Gaussian approximation rate would not follow as stated.","section":"Appendix E, Lemma B.2 and equation (E.11)"}],"minor_comments":[{"comment":"The discussion after Corollary 5.2 says that for m = 1 the last summand in B1_6 equals M^{-1/6}, but the displayed formula in Corollary 5.2 gives 1/M^{1/4} for m = 1; the displayed rate uses M^{-1/4}. Please reconcile the text with the theorem statement.","section":"Section 5.2, text near (5.4)"},{"comment":"In the sup term of Lemma B.2, the expression appears as ||(phi-hat - phi)(x) - (phi-hat - phi)(x)||_8; the second evaluation should be at y to match the intended oscillation modulus.","section":"Appendix E, Lemma B.2"},{"comment":"The phrase 'with probabilities at least 1 - 16B3 ^ 1' and its analogue for B6 should be written as 1 - min(16B3,1) and 1 - min(16B6,1) to match the proof and avoid ambiguity.","section":"Theorem 6.1"},{"comment":"Assumption C.4 writes P(D = 1 | phi_omega(X) = phi_omega(x)) = P(D = 1 | X = x), but the notation is ambiguous about whether the conditioning argument is the transformed variable or the original variable; please state the equality explicitly in terms of sigma-algebras or events.","section":"Section 4.2.1, Assumption C.4"}],"recommendation":"major_revision","confidential_remarks":"This is a promising manuscript with a clear and interesting proof strategy, and I do not see an internal contradiction in the main derivation. The blocking issue is verifiability: Lemma A.4 imports non-asymptotic bounds without reproducing them, and the rank-based bootstrap proof is omitted. These are fixable in a revision and do not, in my view, require rejection. I would encourage the editor to send the revision back to a mathematical statistics referee who can check the imported inequalities in detail."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is a genuine advance: it gives the first non-asymptotic Gaussian approximation bounds for bias-corrected matching-based ATE estimators with the number of matches M diverging, with explicit rates in n, M, and the treatment balance eta, and it does the same for a multiplier bootstrap. Second, the headline rates ride on Lemma A.4, a non-asymptotic density-ratio bound that is load-bearing for the B3 term and the bootstrap high-probability statement, and its proof is not self-contained. Treat the rates as conditional until that lemma is fully verified.\n\nWhat is actually new: combining stabilization theory with the Malliavin-Stein method for matching estimators is a real departure from the usual asymptotic CLT route. The explicit dependence on M and eta is not in Abadie-Imbens or Lin et al. The paper also covers rank-based (CDF-transformed) matching, which is a useful extension. The three-step proof structure is clear and the variance bound in Lemma 7.1 is careful. The citation pattern is honest; they cite the external stabilization theorems and acknowledge the concurrent Lin-Han work on naive bootstrap consistency.\n\nWhere the soft spots are, in proportion. Lemma A.4 is the main one. The proof says it adapts Lin et al. (2023a, proofs of Theorems B.3/B.4) by replacing asymptotic o(n^{-gamma}) simplifications with 'the first inequalities' from those proofs, but it does not reproduce those inequalities or show they hold non-asymptotically under the stated assumptions. For example, the lower-bound step imports a term Cp(1-rratio)(n/M)e^{M-r0 n p0 - M log M + M log(r0 n p0)} 'according to the first inequality' without derivation. If any of those inherited inequalities assumes fixed p0 or a specific order of limits, the rates in eta or M would shift. I do not see an internal contradiction, and the paper is transparent that it is modifying earlier proofs, but the claim of a 'rigorous fully non-asymptotic bound' outruns what is written.\n\nTwo minor points: the rank-based bootstrap proof is a 'mutatis mutandis' statement (checkable, but omitted), and there are no simulations—fine for a theory paper, but the bounds have large polynomial exponents and readers should not expect tight constants.\n\nBottom line: this is a solid paper, aimed at mathematical statisticians in causal inference and geometric probability. It deserves a serious referee. The main gap is fixable by making Lemma A.4 self-contained. Send it to review; require that revision.","headline":"A genuine advance in non-asymptotic theory for matching ATE estimators with diverging M, but the headline rates rest on a density-ratio lemma whose proof imports unverified inequalities; treat the rates as conditional until that lemma is fully proved.","tokens_in":55280,"tokens_out":3515,"would_cite":true,"duration_ms":32130,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62G09","62G20","60F05","60G55"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes explicit finite-sample Gaussian approximation bounds for bias-corrected matching ATE estimators and proves that a multiplier bootstrap achieves the same order.","keywords":["nearest-neighbor matching","average treatment effect","Gaussian approximation","Berry–Esseen bound","multiplier bootstrap","stabilization","Malliavin–Stein method","rank-based matching"],"falsifier":"Simulate the density-ratio setting of Lemma A.4 in one dimension with known densities satisfying Assumption A, for $n=10^2,10^3,10^4$ and $M=\\lfloor n^\\alpha\\rfloor$ with $\\alpha\\in(0,1)$: compute the Monte Carlo MSE of the density-ratio estimator and divide it by the lemma's right-hand side $\\eta^{-2}(M/(n\\eta))^{1/m}+\\delta_{H1}+(\\delta_{H2}+1)/(\\eta^2 M)+\\delta_{H3}$. If this ratio grows without bound as $n$ increases, Lemma A.4 is false, and with it the stated $B_3$ rate.","tokens_in":54188,"feed_emoji":"📊","tokens_out":11551,"duration_ms":99133,"temperature":0.7,"pith_summary":"Nearest-neighbor matching is a standard causal-inference tool, but until now its confidence intervals were justified only asymptotically, with no handle on how the number of matches or the balance between treatment groups affects when the normal approximation becomes accurate. This paper proves that, under standard distributional and regression assumptions, the bias-corrected matching ATE estimator, centered and scaled, is close to a Gaussian in Kolmogorov distance with an explicit error depending on sample size $n$, match count $M$, and the treatment-probability lower bound $\\eta$. The main theorem gives $d_K(\\sqrt{n}(\\hat\\tau_M^{bc}-\\tau), N(0,\\sigma^2)) \\le C(B_1+B_2+B_3)$, with a simplified one-dimensional balanced-data rate $M^{40/(8+p)}n^{-1/2}+M^{-1/2}$. It also proves that a multiplier bootstrap approximates the same distribution at essentially the same rate, so data-driven non-asymptotic confidence intervals become possible. The rank-based matching variant that uses component-wise ranks is treated in parallel, with additional dependence on the error in estimating the transformation.","feed_headline":"Explicit Gaussian rates for matching ATE estimators","feed_subtitle":"Bounds track match count and treatment balance, making confidence intervals non-asymptotically valid.","key_machinery":"Score-function localization combined with a Malliavin-Stein bound for stabilizing functionals. The estimator's main term $nE_n$ is re-expressed as a sum of score functions $\\xi_n(\\tilde{x}_j,\\tilde{X}_n)$, and a score at $j$ is affected by other points only through its $M$ nearest neighbors in the opposite treatment arm, so the radius of stabilization is the $M$-NN distance. Lemma A.1 gives the exponential tail $P(R_n \\ge r) \\le e^2\\exp(-V_m g_{\\min}\\eta/((2M)\\vee 8)\\, n r^m)$. Theorem 7.1, the paper's refinement of a known normal-approximation bound for stabilizing functionals of binomial point processes, tracks constants through $c(M,\\eta,p) \\asymp (M/(\\zeta\\eta))^5 \\vee 1$ and converts the tail and a $4+p$ moment condition into the $S_1$-$S_5$ terms. Lemma A.4 supplies the fully non-asymptotic density-ratio bound on which the variance term $B_3$ rests, and Lemma 7.1 turns it into quantitative upper and lower bounds on $n\\operatorname{Var}E_n$. The bias term $B_2$ is handled by a moment bound whose $\\eta^{-k/m}$ factors come from the binomial counts of treated and control units.","core_discovery":"The paper's central claim is that the bias-corrected nearest-neighbor matching estimator of the average treatment effect has a Gaussian approximation whose error can be bounded explicitly in terms of $n$, the number of matches $M$, and the minimal treatment probability $\\eta$. Writing $\\hat\\tau_M^{bc}=E_n+(B_M-\\hat B_M)$, the main term $E_n$ is shown to be a stabilizing functional of the binomial point process of covariates, treatments, and errors, so a Malliavin-Stein bound for stabilizing functionals applies. Theorem 5.1 states $d_K(\\sqrt{n}(\\hat\\tau_M^{bc}-\\tau), N(0,\\sigma^2)) \\le C(B_1+B_2+B_3)$, where $B_1$ controls the Gaussian approximation of $E_n$, $B_2$ controls the regression-bias correction, and $B_3$ controls convergence of the sample variance to the limiting variance $\\sigma^2$; the last term rests on a new fully non-asymptotic density-ratio bound. In the balanced one-dimensional case this yields $M^{40/(8+p)}n^{-1/2}+M^{-1/2}$. Theorem 6.1 then shows that the multiplier bootstrap statistic $\\sqrt{n}(\\hat\\tau_M^{boot}-\\hat\\tau_M^{bc})$ approximates $\\sqrt{n}(\\hat\\tau_M^{bc}-\\tau)$ conditional on the data, with probability at least $1-16(B_3\\wedge 1)$, at the same order, making data-driven intervals with non-asymptotic coverage possible.","pith_inferences":["The stabilization route used here is generic: any estimator whose leading term is a sum of scores with exponentially decaying interaction radius can be fed through Theorem 7.1, so the same template should yield non-asymptotic CLTs for other nearest-neighbor statistics, random forests, and geometric functionals, as the paper itself suggests for K-fold versions.","Because Lemma A.4 is freestanding, it can be reused outside ATE estimation to make density-ratio-based semiparametric inference finite-sample, offering a testable check of whether the earlier asymptotic density-ratio rates were conservative.","Balancing the two leading terms of the simplified one-dimensional bound suggests an optimal match-count scaling of order $n^{(8+p)/(88+p)}$, roughly $n^{1/10}$ for $p=1$; this is a prediction about simulation-based tuning, not a result proven here.","The bootstrap construction, residuals multiplied by Gaussian variables plus a Gaussian centering term, provides a recipe that could be carried over to other causal estimators, since it only needs a Gaussian approximation theorem for the underlying estimator."],"forward_implications":["Theorem 5.1 turns directly into a coverage lower bound of $1-\\alpha - C(B_1+B_2+B_3)$ for confidence intervals built from the limiting variance.","The multiplier bootstrap gives data-driven intervals whose approximation error is the same order as the Gaussian approximation, holding conditionally on the sample with probability at least $1-16(B_3\\wedge 1)$.","The number of matches may diverge with $n$ at explicit rates under $M \\le C_0 n\\eta$ and $n\\eta^2 \\ge C_1$, with simplified conditions $M^{-1}\\log n=o(1)$ and $n^{-1}M\\log n=o(1)$; this contrasts with the fixed-$M$ failure of the naive bootstrap.","The bounds quantify treatment balance: when $\\eta=o(M)$ the $B_1$ term blows up, so imbalanced designs require smaller $M$, and the same parameter controls the variance term $B_3$.","For the empirical-CDF rank-based estimator, the one-dimensional rate is $M^{40/(8+p)}n^{-1/2}+M^{-1/4}$, worse in the match-count term than the covariate-based $M^{-1/2}$, showing that the choice of metric affects the speed of Gaussian convergence."],"supporting_citations":[{"why":"Defines the nearest-neighbor matching ATE estimator and shows its bias is non-negligible when the covariate dimension exceeds one.","marker":"Abadie and Imbens (2006)"},{"why":"Introduces the bias-corrected estimator that is the object of this paper and establishes its asymptotic normality for fixed M.","marker":"Abadie and Imbens (2011)"},{"why":"Provides the main-term plus bias decomposition, the diverging-M asymptotic normality, and the density-ratio proof whose first inequalities Lemma A.4 makes fully non-asymptotic.","marker":"Lin et al. (2023a)"},{"why":"Supplies the normal-approximation theorem for stabilizing functionals of binomial processes that Theorem 7.1 refines with explicit constants.","marker":"Lachièze-Rey et al. (2019)"},{"why":"Supplies the second-order Poincaré inequality used in the proof of Theorem 7.1.","marker":"Lachièze-Rey and Peccati (2017)"},{"why":"Introduces the general phi-transformed rank-based estimator studied in Theorem 5.2, including the empirical-CDF rank special case and its estimation-error bounds.","marker":"Cattaneo et al. (2023)"},{"why":"Shows the naive bootstrap fails for fixed M, framing the contribution of the multiplier bootstrap with diverging M.","marker":"Abadie and Imbens (2008)"},{"why":"Proposes rank-based matching using component-wise ranks, the estimator family analyzed in Section 3.2 and Theorem 5.2.","marker":"Rosenbaum (2010)"}],"fun_headline_variants":["Non-asymptotic bounds for matching ATE estimators","Gaussian and bootstrap rates for matching ATE","Explicit rates for matching ATE via Malliavin-Stein","Data-driven confidence intervals for matching ATE","Precise Gaussian approximation for matching ATE"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything downstream depends on Lemma A.4's claim that certain inequalities taken from an earlier asymptotic proof of density-ratio convergence remain valid word-for-word when the asymptotic simplifications are stripped away; if one of those inherited inequalities silently relied on the asymptotic regime it came from, the variance term $B_3$ and hence the bootstrap comparison fail.","fun_headline_variants_meta":{"raw":{"variants":["Non-asymptotic bounds for matching ATE estimators","Gaussian and bootstrap rates for matching ATE","Explicit rates for matching ATE via Malliavin-Stein","Data-driven confidence intervals for matching ATE","Precise Gaussian approximation for matching ATE"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000159,"raw_usage":{"total_tokens":1241,"prompt_tokens":971,"completion_tokens":270,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":195}},"tokens_in":587,"tokens_out":270,"duration_ms":2561,"temperature":1.0,"reasoning_tokens":195,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:43:58.049534+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the density-ratio setting of Lemma A.4 in one dimension with known densities satisfying Assumption A, for $n=10^2,10^3,10^4$ and $M=\\lfloor n^\\alpha\\rfloor$ with $\\alpha\\in(0,1)$: compute the Monte Carlo MSE of the density-ratio estimator and divide it by the lemma's right-hand side $\\eta^{-2}(M/(n\\eta))^{1/m}+\\delta_{H1}+(\\delta_{H2}+1)/(\\eta^2 M)+\\delta_{H3}$. If this ratio grows without bound as $n$ increases, Lemma A.4 is false, and with it the stated $B_3$ rate.","supporting_citations":[],"review_version":1}