{"id":"75db0b19-de43-4951-8898-8a2b583b8e7a","arxiv_id":"1908.05607","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Under a global undersmoothing condition, the m-th order spline HAL-MLE is asymptotically efficient as a plug-in estimator for pathwise differentiable target parameters, with applications to average treatment effects and the integral of the squared density.","lead":"This paper studies the spline highly adaptive lasso (HAL), a flexible function estimator, and gives conditions under which plug-in estimates of smooth target parameters reach the semiparametric efficiency bound. The conditions require undersmoothing the fit so that the estimator approximately solves the efficient influence curve equation, and the authors demonstrate the approach on average treatment effects and density integrals.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The ATE theorem's n^-1/4 rate for the projection G0n is a load-bearing assumption, not a consequence of undersmoothing; the basis selected for Q0 need not approximate G0, so the efficiency claim is only conditional.","rationale":"The reader identified the unproved approximation-rate assumptions for nuisance parameters and projections as the weakest point, and my reading agrees: the load-bearing step for the ATE example is the rate ||G0n - G0||_{P0} = OP(n^-1/4), and for the density example it is the projection-residual condition (14). The paper's own language, 'appears to be a reasonable condition' and 'should generally also be rich enough', concedes that these are not proved. This is not an internal inconsistency: if Theorem 3 is read as a conditional statement with that rate as an assumption, the proof structure is coherent, and the main decomposition in Theorem 2 is standard. The concern is that the abstract and introduction present the undersmoothing condition as the sufficient and practically verifiable condition for efficiency, while the two worked examples require additional, non-verifiable approximation properties of the basis selected by the HAL-MLE. The suggested simulation is a direct check because it isolates the mechanism by which the assumption can fail: the selected basis is optimized for Q0, not for G0. I do not see a reason to change the reader's CONDITIONAL verdict; the issue is real but addressable by either proving a general approximation-rate theorem for such projections under explicit smoothness and richness conditions or by restating the efficiency result as conditional on the projection rate. No ad hominem is intended; the critique concerns the argument, not the authors.","tokens_in":36515,"tokens_out":5746,"duration_ms":67748,"concrete_test":"Run a simulation in the ATE setting with a two-dimensional covariate W=(W1,W2), where Q0 depends only on W1 (e.g., logit Q0 = W1) and G0 depends only on W2 (e.g., logit G0 = 2W2 - 1). For sample sizes n = 500, 1000, 2000, 4000, fit the undersmoothed HAL-MLE Qn satisfying condition (7), and compute G0n as the L2(P0) projection of G0 onto the linear span of the basis functions with nonzero coefficients in Qn. Record sqrt(n)||G0n - G0||_{P0} and compare it with the target n^-1/4 rate; also record n times the Monte Carlo variance of the plug-in estimator against the efficiency bound. If sqrt(n)||G0n - G0||_{P0} does not remain stochastically bounded or diverges, the assumption in Theorem 3 fails in a setting that satisfies the undersmoothing condition, and the efficiency conclusion for this example is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 3 states that if ||G0n - G0||_{P0} = OP(n^-1/4), then the undersmoothed HAL-MLE plug-in is asymptotically efficient for the ATE, but this rate is not proved. In Section 4.4 the paper says only that the condition 'appears to be reasonable' and that the selected basis functions 'should generally also be rich enough' to approximate G0. This is a genuine gap, not a stylistic one. The second-order remainder in the ATE example is bounded by (1/delta)||Qn - Q0||_{P0} ||G0n - G0||_{P0}; with ||Qn - Q0||_{P0} = OP(n^-1/4 - alpha), the G0n rate of OP(n^-1/4) is exactly what makes the remainder oP(n^-1/2). If G0n converges more slowly, the plug-in can fail to be efficient even when the global undersmoothing condition (7) is satisfied, because the score equation alone does not control the remainder. There is also a structural reason the rate is not implied: G0n is the projection of the true propensity score onto the linear span of the basis functions that Qn happened to select, and those basis functions are chosen to fit Q0. If Q0 depends on one set of covariates and G0 depends on another, the selected span can be rich for Q0 while being poor for G0, so no argument from the rate of Qn establishes the required rate for G0n. The density example has the same pattern: Theorem 4 assumes the projection-residual condition (14), P0{Π⊥_n(D*(Qn)-D*(Q0))} = oP(n^-1/2), with only a parenthetical sufficient condition that the projection operator has operator norm OP(n^-1/4), again without proof. Theorems 3 and 4 are thus internally consistent conditional statements, but they do not demonstrate the advertised unconditional efficiency conclusion for the two examples. The paper itself flags these as assumptions to verify, and no verification is provided.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper defines an m-th order Spline HAL-MLE over a class of cadlag functions with bounded m-th order sectional variation norm, proves a spline representation theorem, and studies when a single undersmoothed HAL-MLE plug-in is asymptotically efficient for a pathwise differentiable target parameter. The main device is Theorem 1: if the HAL-MLE solves score equations for a sufficiently rich set of submodels, measured by the data-checkable undersmoothing condition (7), then the empirical mean of an approximation D*_n(Qn,G0) of the efficient influence curve is oP(n^{-1/2}). Theorem 2 then gives the standard TMLE-style efficiency argument under second-order remainder, Donsker, and L2-consistency conditions. The paper applies the result to the average treatment effect (Section 4, Theorem 3) and to the integral of the square of a density (Section 5, Theorem 4), and reports simulations for both examples.","tokens_in":36996,"tokens_out":6133,"duration_ms":62513,"significance":"If the main theorem is correct, the undersmoothing condition is a useful and non-parameter-specific criterion: a single HAL-MLE can be efficient for many smooth functionals simultaneously without a separate TMLE update. The paper ships self-contained proofs of the spline representation theorem (Appendix A), of Theorem 1 (Appendix C), and of the TMLE-style efficiency argument (Appendix D), and the simulation evidence supports the finite-sample behavior. The central gap is the unproved rate at which the span of basis functions selected for Qn approximates the nuisance parameter G0 in the ATE example, and the analogous projection condition in the density example; this makes the example theorems conditional rather than unconditional.","major_comments":[{"comment":"The assumption ||G0n - G0||_{P0} = OP(n^{-1/4}) is load-bearing and is not proved. In the text this is justified only by the statement that the condition 'appears to be a reasonable condition' and that the selected basis functions 'should generally also be rich enough' to approximate G0 at that rate. This matters because R20(...,...) <= (1/delta)||Qn - Q0||_{P0} ||G0n - G0||_{P0}; combined with the established ||Qn - Q0||_{P0} = OP(n^{-1/4-alpha}), the n^{-1/4} rate on G0n is exactly the threshold for R20 = oP(n^{-1/2}). The selected span is chosen to fit Q0, not G0; if Q0 and G0 depend on different coordinates, the span can be rich for Q0 and poor for G0, so no argument from the rate of Qn establishes the required rate for G0n. The theorem should be restated as conditional on this projection rate, or the rate should be proved under explicit additional conditions on G0 and the basis selection.","section":"§4.4, Theorem 3"},{"comment":"The density example contains the same structural gap. Theorem 4 assumes P0{Π_perp_n(D*(Qn) - D*(Q0))} = oP(n^{-1/2}) in equation (14), with only a parenthetical sufficient condition that the projection operator norm ||Π_perp_n|| = OP(n^{-1/4}), and the construction also simply states 'We will assume that ||fn(Qn) - f(Qn)||_{P0} = oP(n^{-1/4})'. None of these projection or residual rates is proved for the HAL basis selected for Qn. Since this condition is exactly what makes the approximation D*_n(Qn) - D*(Qn) negligible in the proof of Theorem 2, the efficiency conclusion for this example is conditional on an unverified structural assumption. As with Theorem 3, the paper should either prove the required rate or state the theorem with the projection condition as an explicit assumption rather than as a parenthetical remark.","section":"§5, Theorem 4, Eq. (14)"},{"comment":"The abstract and introduction state the efficiency result as an unconditional claim for the undersmoothed HAL-MLE in the two examples. Given the unproved nuisance approximation rates in Theorems 3 and 4, the present version supports only conditional efficiency statements for those examples. The claims should be softened until the missing rates are supplied, or the theorems should be presented as conditional results with the projection conditions listed among the assumptions.","section":"Abstract and Section 1"}],"minor_comments":[{"comment":"In the statement of Theorem 1, 'PnD*_n(Qn,G0) = oP(n^{-1.2})' should be 'PnD*_n(Qn,G0) = oP(n^{-1/2})'; the same typo appears in the inline text following condition (7).","section":"Theorem 1"},{"comment":"Both sections refer to 'condition (25)' before any equation (25) has appeared; equation (25) is only defined later in Appendix E. The numbering should be changed so that the referenced condition appears in the main text or the cross-reference is clarified.","section":"§4.4 and §5"},{"comment":"In the proof of Lemma 2, after supposing P(mn < m0) > delta, the proof defines A = {m : m > m0} and then asserts P(mn in A) > delta; the intended set is A = {m : m < m0}. This is a typo but makes the proof difficult to follow.","section":"Appendix B, Lemma 2"},{"comment":"The phrase 'this reduces to the assumption that O(min_{s,j} Pn(W(s) >= ws,j)) = OP(...)' is awkward and formally incorrect; the min of the empirical probabilities is already a random variable, so it should not be wrapped in O(). The intended statement is that the minimum itself is OP(...).","section":"§4.4 and §5"},{"comment":"The simulation section correctly acknowledges that the practical undersmoothing selector requires a constant that the theory does not specify, and that the choice of constant affects performance. This is a useful caveat, but the paper should state explicitly that the simulation-based targeted criterion is a heuristic for the theoretical condition (7), not a proven instance of it.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The core proof of Theorem 1 is coherent and the undersmoothing idea is valuable, so I am not recommending rejection. However, the example theorems are stated more strongly than what is proved: the n^{-1/4} projection rate for G0n in Theorem 3 and the projection-residual condition (14) in Theorem 4 are assumed rather than established. I would like the revision to either prove these rates under explicit conditions or clearly present the examples as conditional efficiency theorems, and to adjust the abstract and introduction accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere is my read on the Spline HAL undersmoothing paper. The genuinely new and valuable piece is Theorem 2: an explicit, data-checkable undersmoothing condition (min over selected basis functions of the empirical score is o_P(n^{-1/2})) that turns a single HAL-MLE plug-in into an efficient estimator for any smooth pathwise differentiable target. The m-th order spline representation theorem is also a real extension of the zero-order HAL result, and the proof strategy is the standard exact second-order expansion plus Donsker equicontinuity. I believe the central mechanism works.\n\nThe soft spots are where the reader and stress-test put them. Theorems 3 and 4, for ATE and density square, both assume rather than prove the nuisance approximation rate. In Section 4.4 the paper needs ||G0n - G0|| = O_P(n^{-1/4}) for the projection of the true propensity score onto the span of basis functions selected by Qn. The text says this 'appears to be reasonable' and that the selected basis 'should generally also be rich enough.' That is not a proof, and it is not a minor gap: the second-order remainder for ATE is exactly the product of ||Qn-Q0|| and ||G0n-G0||, so the n^{-1/4} rate is load-bearing. The density example has the same shape with condition (14), where the projection residual is assumed o_P(n^{-1/2}) with only a parenthetical sufficient condition on the operator norm.\n\nI would not call the paper misleading; it labels these as assumptions to verify, and the simulations are encouraging even without code. But the advertised unconditional efficiency conclusion for the examples is conditional on rates that are not established and may fail if the basis selected for Q0 is poor for G0. The main theorem stands on its own, so this is an addressable revision rather than a fatal flaw.\n\nThis paper deserves a serious referee. A careful referee should push the authors to either prove the missing rates under concrete conditions or state them as explicit assumptions in the theorem statements. I would also ask for code and data for the simulations. The citation pattern is fine: prior HAL work is cited, and the new claims are distinguishable.\n\nVerdict: engage it, conditional. If the authors close the G0n-rate gap, this becomes the standard reference for HAL-based plug-in efficiency.","headline":"The undersmoothed HAL efficiency theorem is a real and useful step forward, but the paper's two headline applications rest on unproved nuisance-approximation rates that are exactly what the main theorem was supposed to deliver.","tokens_in":37462,"tokens_out":738,"would_cite":true,"duration_ms":10232,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62G08","62G20","62F12"],"pacs":[],"model":"deepseek-v4-flash","headline":"A single over-fitted HAL-MLE is asymptotically efficient for every smooth target at once.","keywords":["asymptotic efficiency","highly adaptive lasso","undersmoothing","pathwise differentiable parameter","canonical gradient","sectional variation norm","spline basis functions","plug-in estimator"],"falsifier":"Simulate an ATE model in which $G_0$ has a narrow local feature not aligned with the observed support points (for example, a steep dip in the propensity score between two observed covariate values), while $Q_0$ is smooth; compute $\\|G_{0n}-G_0\\|_{P_0}$ for $n=10^3,10^4,10^5$ and the scaled quantities $\\sqrt{n}\\,\\|P_nD^*(Q_n,G_0)\\|$ and $\\sqrt{n}(\\Psi(Q_n)-\\Psi(Q_0))$. If the empirical projection rate is slower than $n^{-1/4}$, the scaled bias and score will not vanish, giving a concrete instance where Theorem 3's assumption fails even though condition (7) holds.","tokens_in":36307,"feed_emoji":"📈","tokens_out":8292,"duration_ms":73627,"temperature":0.7,"pith_summary":"This paper tries to establish that a single, deliberately over-fitted version of the Highly Adaptive Lasso estimator—the m-th order Spline HAL-MLE, with its L1 bound increased past the cross-validation choice—is asymptotically efficient for any smooth ('pathwise differentiable') feature of the function being estimated. The reason is that a larger L1 bound forces the lasso fit to include very sparsely supported basis functions, and the score equations such an MLE solves then approximate the efficient influence curve equation, the key requirement for efficiency. If right, applied researchers could estimate a whole menu of target parameters (average treatment effects, density functionals) with one plug-in fit instead of running a separate targeted estimator for each. The argument is carried by a representation theorem: any cadlag function with finite sectional variation norm is an infinitesimal linear combination of tensor-product spline basis functions, so the HAL-MLE is a lasso over that basis and its score equations are exactly parametric score equations.","feed_headline":"One undersmoothed HAL fit yields efficient estimates of many targets","feed_subtitle":"With the lasso penalty large enough, a single fit solves the efficient score equation for any smooth target.","key_machinery":"The load-bearing object is the m-th order spline basis representation and the score equations it generates. A cadlag function with finite sectional variation norm is represented as $Q(x)=Q(0)+\\sum_{\\bar{s}(m)}\\int \\varphi_{\\bar{s}(m),x}(z_{s_m})\\,dQ^m_{\\bar{s}(m)}(z_{s_m})$, an infinitesimal linear combination of tensor products of up to m-th order splines, with the $L_1$-norm of the coefficients equal to the m-th order sectional variation norm. The HAL-MLE minimizes empirical risk over such representations with $\\|\\beta\\|_1 \\le C_n$. Because it is an MLE, it solves the score equation $P_n S_h(Q_n)=0$ for every direction $h$ that preserves the $L_1$ constraint; the undersmoothing condition (7) makes one of these directions match the canonical gradient score, yielding $P_n D^*_n(Q_n,G_0)=o_P(n^{-1/2})$.","core_discovery":"The central claim, Theorem 2, is that if the m-th order Spline HAL-MLE $Q_n$ is computed with an $L_1$-bound $C_n$ large enough that the sparsest selected basis function satisfies condition (7), then the plug-in estimator $\\Psi(Q_n)$ is asymptotically efficient for the target parameter $\\Psi(Q_0)$: its scaled bias vanishes faster than $n^{-1/2}$ and its variance reaches the efficiency bound, provided the second-order remainder is $o_P(n^{-1/2})$ and the canonical gradients stay in a Donsker class. Because condition (7) is global and not parameter specific, the same over-fitted $Q_n$ is efficient for every smooth pathwise differentiable functional at once. The representation theorem for the class $D^m[0,\\tau]$—functions with finite m-th order sectional variation norm—is what makes this concrete: each such function is an infinitesimal linear combination of tensor-product spline bases, so the HAL-MLE is literally a lasso on spline basis coefficients.","pith_inferences":["A testable consequence of the paper's logic is that you could post-select any number of target parameters after seeing one undersmoothed HAL fit and report efficient estimates for all of them; this would invert the usual TMLE workflow where the target must be declared before estimation.","The unproved $n^{-1/4}$ projection-rate assumption suggests a practical diagnostic: track $\\|G_{0n}-G_0\\|_{P_0}$ (and the residual in condition (14)) across sample sizes; if the empirical rate is slower, enlarging $C_n$ alone may not restore efficiency and the basis itself may need to be enriched near rough parts of $G_0$.","The paper's discussion implies a comparative conjecture that could be settled by simulation: undersmoothed HAL-MLE should beat HAL-TMLE when the nuisance $G_0$ is as hard to estimate as $Q_0$ or when positivity is weak, and lose when $G_0$ is strongly constrained.","A refinement of condition (7) suggested by the proof—checking the projection residual of $D^*(Q_n,G_0)$ onto the span of selected basis functions rather than only the minimum score—might give a more stable finite-sample undersmoothing selector than the paper's global rate condition."],"forward_implications":["With one HAL-MLE computed at a large enough $L_1$ bound, every pathwise differentiable smooth functional of the fitted function is estimated efficiently, so no separate targeted step per target is needed.","Condition (7) can be checked from the data alone, which yields a concrete rule: take the smallest $C_n$ above the cross-validated choice for which $\\min_{s,j} \\|P_n \\frac{d}{dQ_n}L(Q_n)(\\varphi_{s,j})\\| = o_P(n^{-1/2})$.","Undersmoothing does not sacrifice the HAL-MLE's own rate: the $L_1$ bound may even grow slowly with $n$ while preserving the faster-than-$n^{-1/4}$ convergence and the Donsker property.","The smoothness-adaptive version (selecting $m$ by cross-validation) inherits the efficiency result, since under separated rates the cross-validation selector picks the true smoothness order $m_0$ with probability tending to 1.","In the two worked examples—average treatment effect and the integral of the square of the density—the undersmoothed plug-in reaches the efficiency bound in simulations, while the cross-validated HAL-MLE does not."],"supporting_citations":[{"why":"Defines the original m=0 HAL-MLE and its rate, the estimator generalized here; also the source of the TMLE-style efficiency proof template.","marker":"van der Laan (2015)"},{"why":"Supplies the HAL-MLE implementation and the faster-than-n^{-1/4} rate used in the examples.","marker":"Benkeser and van der Laan (2016)"},{"why":"Underlies the sectional variation norm and the measure representation on which the spline basis representation builds.","marker":"Gill et al., 1995"},{"why":"Establishes the canonical-gradient characterization of efficiency that the plug-in estimator must satisfy.","marker":"Bickel et al., 1997"},{"why":"Introduces targeted minimum loss estimation, whose standard efficiency expansion Theorem 2 adapts to the undersmoothed HAL-MLE.","marker":"van der Laan and Rubin, 2006"},{"why":"Provides the supremum-norm convergence of HAL used in Lemma 1 to weaken the support condition.","marker":"van der Laan and Bibaut"}],"fun_headline_variants":["Undersmoothed HAL: one fit achieves efficiency for all smooth targets","Spline HAL with global undersmoothing: efficient for every smooth target","One lasso on splines: efficient for all pathwise differentiable targets","Adaptive smoothness via spline HAL: one fit, all efficient targets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The examples rely on the assumption, stated but not proved in Section 4.4 and Theorem 4, that the projection of the true nuisance $G_0$ onto the basis functions selected by the undersmoothed HAL fit converges to $G_0$ at rate $O_P(n^{-1/4})$ (or an analogous projection-residual condition); if that projection is slower, the second-order remainder is not $o_P(n^{-1/2})$ and the efficiency conclusion for those examples fails.","fun_headline_variants_meta":{"raw":{"variants":["Undersmoothed HAL: one fit achieves efficiency for all smooth targets","Spline HAL with global undersmoothing: efficient for every smooth target","One lasso on splines: efficient for all pathwise differentiable targets","Adaptive smoothness via spline HAL: one fit, all efficient targets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000634,"raw_usage":{"total_tokens":3019,"prompt_tokens":1133,"completion_tokens":1886,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":749,"completion_tokens_details":{"reasoning_tokens":1805}},"tokens_in":749,"tokens_out":1886,"duration_ms":12014,"temperature":1.0,"reasoning_tokens":1805,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:20:02.255897+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate an ATE model in which $G_0$ has a narrow local feature not aligned with the observed support points (for example, a steep dip in the propensity score between two observed covariate values), while $Q_0$ is smooth; compute $\\|G_{0n}-G_0\\|_{P_0}$ for $n=10^3,10^4,10^5$ and the scaled quantities $\\sqrt{n}\\,\\|P_nD^*(Q_n,G_0)\\|$ and $\\sqrt{n}(\\Psi(Q_n)-\\Psi(Q_0))$. If the empirical projection rate is slower than $n^{-1/4}$, the scaled bias and score will not vanish, giving a concrete instance where Theorem 3's assumption fails even though condition (7) holds.","supporting_citations":[],"review_version":1}