Pith. sign in

REVIEW 1 major objections 5 minor 27 references

A note on the properties of the confidence set for the local average treatment effect obtained by inverting the score test

T0 review · 1 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The score confidence set for the local average treatment effect takes only six possible shapes, and under strong instruments it coincides with the standard Wald interval up to a 1/n error.

desk verdict Score confidence set paper has a solid core in Sections 3–4, but Theorem 2's weak-instrument formula has a sign error; it deserves peer review after that section is fixed. read the letter →

arxiv 2506.10449 v1 pith:6BLNAYTR submitted 2025-06-12 math.ST stat.MEstat.TH

classification math.STstat.MEstat.TH MSC 62F2562G0562G20
keywords localaveragetreatmenteffectscoreconfidencesetweakinstrumentsdoublemachinelearninginfluencefunctionuniformlyvalidsetsAnderson-Rubintestweak-instrumentasymptotics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper studies a confidence set for the local average treatment effect that is built by inverting a score test based on the estimated nonparametric influence function. Such a set is attractive because it remains uniformly valid even when the instrument is arbitrarily weak, whereas ordinary Wald intervals can badly under-cover. The authors prove that the set can only take six forms—finite interval, infinite ray(s), the whole line, a point, or empty—and that it has infinite diameter exactly when the data fail to reject the hypothesis that the instrument has no average effect on treatment. Their main theorem shows that at a fixed law with sufficiently fast nuisance estimation, the score set agrees with the doubly robust Wald interval up to endpoints of order 1/n, so in strong-instrument settings it sacrifices no precision; when the efficient influence function equals the nonparametric one, it is asymptotically no wider than any regular Wald interval. The paper also shows that under weak-instrument asymptotics the doubly robust estimator is biased and heavy-tailed, while the score set keeps correct coverage.

What carries the argument

The central object is the score confidence set Cn = {θ : |Sn(θ)| ≤ z1−α/2}, where Sn(θ) is the empirical mean of the estimated influence function ψb,η − θψa,η over its empirical standard deviation. The key identity is that |Sn(θ)|² ≤ z² is equivalent to aθ² + bθ + c ≤ 0 with coefficients a, b, c built from empirical means of ψa,η, ψb,η, and their squares and products, so the set's geometry is controlled by the sign of ∆ = b² − 4ac and the sign of a. Theorem 1 is carried by proving P(a > 0, ∆ > 0) → 1 and then expanding the two roots of the quadratic to show they match the Wald endpoints up to OP(1/n); the product-rate condition (5) is what makes those empirical coefficients converge to their population limits fast enough.

What would settle it

Run the Section 6 weak-instrument simulation with π = 0.15/√n but replace the nuisance fit for the treatment-assignment probability with an estimator whose L2 error decays like $n^{{-1/4}}$, so condition (5) fails: the paper predicts the DRML Wald interval under-covers while the score set holds coverage, so observing the score set also under-cover would contradict the central claim. Alternatively, in a law where EPψa,ηP = 0, Theorem 1 does not apply and the paper predicts Cn is unbounded with high probability, so finding a finite Cn there would falsify the necessity of that condition.

Watch

Extended reading notes

Core claim

The claim is that, at any fixed law P where conditions (2)-(6) hold and EPψa,ηP ≠ 0, the score confidence set Cn equals [bφ − z1−α/2 bσ/√n + OP(1/n), bφ + z1−α/2 bσ/√n + OP(1/n)] with probability tending to one, where bφ is the doubly robust machine learning (DRML) estimator and $bσ^{2}$ estimates the variance of the nonparametric influence function φ1_P(O). Since $bσ^{2}$ estimates the semiparametric efficiency bound whenever the efficient influence function is the nonparametric one, the score set is asymptotically at least as short as any Wald interval from a regular asymptotically linear estimator. In weak-instrument asymptotics the same DRML estimator converges to a ratio of correlated normals with infinite mean, so its Wald interval is not a reliable tool there. The six-form classification follows from viewing Cn as the sublevel set of the quadratic inequality aθ² + bθ + c ≤ 0, whose coefficients are explicit empirical sums of the estimated influence functions.

Load-bearing premise

The load-bearing premise is that the product of the L2 errors of the estimated treatment-assignment probability and the slower of the estimated outcome and treatment regressions converges faster than $n^{{-1/2}}$; if the machine learning nuisance estimates converge too slowly, the score set may no longer coincide with the doubly robust Wald interval.

Editorial extensions

If this is right

  • At strong instruments, practitioners can use the score confidence set instead of the DRML Wald interval without asymptotically sacrificing precision, because the two sets agree to OP(1/n).
  • An unbounded score set acts as a direct diagnostic: it occurs exactly when the test of zero average instrument effect on treatment does not reject, signaling that the data are uninformative about the local average treatment effect.
  • In models where the instrument assignment probability is known, such as randomized encouragement designs, the score set is asymptotically no wider than the Wald interval from any regular estimator, so it inherits efficiency.
  • The DRML point estimator should not be reported under weak instruments: its limiting distribution is a ratio of normal variables with infinite mean, so its Wald interval can be misleadingly short.
  • The provided algorithm and software implementation make the score confidence set immediately usable in double machine learning pipelines.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I infer that the six-form classification is not specific to the local average treatment effect: any score test whose statistic is a ratio of an empirical mean to an empirical standard deviation of a single score will produce a confidence set with the same six shapes, giving a general template for weak-instrument-robust inference.
  • I infer that the product-rate condition (5) could be relaxed if one only wants uniform validity rather than the OP(1/n) endpoint match, but then the asymptotic equivalence with the Wald interval would hold at a slower rate; the paper does not claim this.
  • I infer that the diagnostic property (infinite diameter iff non-rejection of the zero-effect test) is an honest alternative to first-stage F-statistics for detecting weak instruments, because it directly targets identifiability of the estimand rather than an arbitrary threshold.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. The paper studies the confidence set obtained by inverting a score test based on the nonparametric influence function for the local average treatment effect under possibly weak instruments. It makes four main contributions: Proposition 1 classifies the possible geometric forms of the score confidence set; Theorem 1 shows that at any fixed law satisfying standard double-machine-learning regularity conditions, the score confidence set coincides with the Wald interval based on the doubly robust estimator up to O_P(1/n) endpoint differences; Corollary 1 draws an optimality conclusion about the diameter of the score set relative to regular asymptotically linear estimators; and Theorem 2 derives a weak-instrument limiting distribution for the doubly robust estimator that is non-normal and has infinite mean. The paper also contains simulations, a real-data application to Glovo experiments, and an implementation in the DoubleML package.

Significance. If the fixed-law results are correct, the paper provides a practically important bridge: a weak-instrument-robust confidence set that, at strong instruments, asymptotically matches the precision of the DRML Wald interval, and that is asymptotically no wider than any regular Wald interval when the efficient influence function is the nonparametric one. The simulation code and DoubleML implementation are concrete and reproducible, and the application to real experiments is useful. The weak-instrument result in Theorem 2 is one of the paper's stated contributions, but its exact statement is internally inconsistent with Lemma 1; the qualitative message of non-normality and infinite mean appears to survive, but the stated limiting law needs correction. For these reasons the paper is a solid methods note provided the weak-instrument theorem is fixed.

major comments (1)
  1. [§5, Theorem 2 and Lemma 2] Theorem 2 is internally inconsistent with Lemma 1. Lemma 1 defines (Na, Nb) as the limit of sqrt(n)(Pn ψa,ηPn - EPn ψa,ηPn, Pn ψb,ηPn - EPn ψb,ηPn), so under Condition 1 one has sqrt(n) Pn ψa,ηPn -> Na + ca and sqrt(n) Pn ψb,ηPn -> Nb + cb. Consequently, under Condition 2 the limit of the DRML estimator minus the estimand must be (Nb+cb)/(Na+ca) - cb/ca, which equals (ca Nb - cb Na)/(ca(Na+ca)). The expression in Theorem 2, (ca Nb + cb Na)/(ca^2 - ca Na), differs in the sign of the cb Na term and in the sign of Na in the denominator. The same sign error appears in Lemma 2 and in its proof: the factor 1 - sqrt(n)(EPn ψa - Pn ψa)/ca should be 1 + sqrt(n)(Pn ψa - EPn ψa)/ca, since the latter is the object converging to Na. The qualitative conclusion that the limiting law is non-normal with infinite mean survives, because the corrected ratio still has a nondegenerate normal denominator and the Marsaglia heavy-tail argument applies, but the exact limiting law stated in Theorem 2 is not correct as written.
minor comments (5)
  1. [§4, proof of Theorem 1, display (8)] In the proof of Theorem 1, the displayed equality after (8) states EP{ψb,ηP + φ(P)ψa,ηP}^2 where the correct expression is EP{ψb,ηP - φ(P)ψa,ηP}^2; because only positivity of the displayed quantity is needed, this appears to be a typo, but it should be corrected.
  2. [§4, proof of Theorem 1] The notation z^2_{1-α2} appears in the same proof and should be z^2_{1-α/2}.
  3. [§5 and Introduction] There are several typos: 'strenght' should be 'strength' in Section 5 after Condition 1; 'arbirtrarily' should be 'arbitrarily' in the Introduction; and 'far below the nominal below' should be 'far below the nominal level'.
  4. [§6] The sample sizes are reported as n in {1500, 4500, 7500, 10500, 12000}, but Figure 1's horizontal axis begins at 2000; the axis and the reported grid should be aligned.
  5. [Proposition 1] The text promises that the confidence set can take one of six forms, but Proposition 1 lists seven cases; it would be clearer to say that Cases 1-6 are the six geometric types and Case 7 is a degenerate subcase.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: fixed-law equivalence and weak-instrument limit are derived from explicit conditions, not from the target claims.

full rationale

The paper's central claims are derived rather than imported. Proposition 1 is an algebraic classification of the quadratic inequality that defines C_n from S_n(θ), so it is a direct consequence of the definition and not a renamed input. Theorem 1 is proved from the stated regularity conditions (2)-(6) and the standard asymptotic linearity of the DRML estimator; the proof computes the roots of the quadratic and shows they are the Wald endpoints up to O_P(1/n), which is an actual asymptotic derivation, not a definitional identity. Corollary 1 uses the external semiparametric efficiency bound (Van der Vaart, Theorem 25.20) rather than a self-citation. Theorem 2 is argued from Conditions 1-2 through Lemmas 1-2; even if the displayed limit (7) contains an algebraic sign error, that is a correctness concern, not circularity, because the target distribution is not assumed as an input. Self-citations to Smucler et al. (2019, 2025) are background references for the rate double-robustness condition and for the classical non-uniformity impossibility; the substantive assumptions are stated explicitly in (2)-(6) and are not conclusions of this paper. No fitted parameter is relabeled as a prediction, and no uniqueness claim is imported from the authors' prior work. Thus no circular step is exhibited.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper introduces no free parameters or invented entities. Its theorems are parameter-free given the stated regularity conditions; simulation constants such as pi are illustrative design choices, not fitted quantities. The main imported ingredients are standard DML rate conditions, a high-level weak-instrument condition, causal assumptions for the LATE interpretation, and Ma (2023)'s uniform-validity theorem.

assumptions (5)
  • domain assumption Nuisance estimators satisfy regularity conditions (2)-(6), including the product rate condition (5).
    Needed for Theorem 1 and Corollary 1. If machine learning nuisance estimates converge too slowly, the claimed equality between the score confidence set and the Wald interval may fail.
  • domain assumption Condition 2: sqrt(n) Pn(psi_a,beta - psi_a,eta) = oPn(1) and sqrt(n) Pn(psi_b,beta - psi_b,eta) = oPn(1) under weak-instrument sequences.
    High-level condition used to prove Theorem 2. The paper cites Takatsu et al. (2023) for primitive sufficient conditions but does not verify them in the simulation design.
  • domain assumption Corollary 1 restricts to models where the efficient influence function equals the nonparametric influence function, such as a known propensity score.
    Needed for the optimal-diameter claim. The paper gives a known-propensity example but does not analyze a fully nonparametric model.
  • domain assumption LATE identification conditions: exclusion Y(z,a)=Y(z), independence of Z given X, monotonicity, and overlap.
    Background assumptions used to interpret the functional phi(P) as the local average treatment effect. They are not used in the statistical derivations themselves.
  • standard math Ma (2023): the score confidence set is uniformly valid over models allowing arbitrarily weak instruments.
    Imported external theorem used to frame the contribution. The paper does not reprove uniform validity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A note on the properties of the confidence set for the local average treatment effect obtained by inverting the score test." pith.science (2026). https://pith.science/paper/6BLNAYTR

@misc{pith2026250610449,
  author       = {Pith},
  title        = {Pith review of: A note on the properties of the confidence set for the local average treatment effect obtained by inverting the score test},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6BLNAYTR}},
  note         = {Machine review of arXiv:2506.10449}
}
read the original abstract

We study the properties of the score confidence set for the local average treatment effect in non and semiparametric instrumental variable models. This confidence set is constructed by inverting a score test based on an estimate of the nonparametric influence function for the estimand, and is known to be uniformly valid in models that allow for arbitrarily weak instruments; because of this, the confidence set can have infinite diameter at some laws. We characterize the six possible forms the score confidence set can take: a finite interval, an infinite interval (or a union of them), the whole real line, an empty set, or a single point. Moreover, we show that, at any fixed law, the score confidence set asymptotically coincides, up to a term of order 1/n, with the Wald confidence interval based on the doubly robust estimator which solves the estimating equation associated with the nonparametric influence function. This result implies that, in models where the efficient influence function coincides with the nonparametric influence function, the score confidence set is, in a sense, optimal in terms of its diameter. We also show that under weak instrument asymptotics, where the strength of the instrument is modelled as local to zero, the doubly robust estimator is asymptotically biased and does not follow a normal distribution. A simulation study confirms that, as expected, the doubly robust estimator performs poorly when instruments are weak, whereas the score confidence set retains good finite-sample properties in both strong and weak instrument settings. Finally, we provide an algorithm to compute the score confidence set, which is now available in the DoubleML package for double machine learning.

Figures

Figures reproduced from arXiv: 2506.10449 by the authors.

Figure 1
Figure 1. Empirical coverage of the score confidence set and the DRML Wald confidence interval in the weak instrument [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Empirical coverage of the score confidence set and the DRML Wald confidence interval in the strong instrument [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Median length of the score confidence set and the DRML Wald confidence interval in the strong instrument setting. [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Ratio of the diameters of the score confidence set and the DRML Wald confidence interval for Promo 1 and Promo [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 18 canonical work pages

  1. [1]

    Anderson, T. W. and Rubin, H. (1949). Estimation of the parameters of a single equation in a complete system of stochastic equations. The Annals of mathematical statistics , 20(1):46--63

  2. [2]

    W., Moreira, M

    Andrews, D. W., Moreira, M. J., and Stock, J. H. (2006). Optimal two-sided invariant similar tests for instrumental variables regression. Econometrica , 74(3):715--752

  3. [3]

    and Mikusheva, A

    Andrews, I. and Mikusheva, A. (2016). Conditional inference with a functional nuisance parameter. Econometrica , 84(4):1571--1612

  4. [4]

    H., and Sun, L

    Andrews, I., Stock, J. H., and Sun, L. (2019). Weak instruments in instrumental variables regression: Theory and practice. Annual Review of Economics , 11(1):727--753

  5. [5]

    S., and Spindler, M

    Bach, P., Chernozhukov, V., Kurz, M. S., and Spindler, M. (2022). Doubleml-an object-oriented implementation of double machine learning in python. Journal of Machine Learning Research , 23(53):1--6

  6. [6]

    Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., and Hansen, C. (2018). Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal , 21(1):C1--C68

  7. [7]

    Dufour, J.-M. (1997). Some impossibility theorems in econometrics with applications to structural and dynamic models. Econometrica: Journal of the Econometric Society , pages 1365--1387

  8. [8]

    Elbers, B. (2023). Encouragement designs and instrumental variables for a/b testing. https://engineering.atspotify.com/2023/08/encouragement-designs-and-instrumental-variables-for-a-b-testing/

Show all 27 references
  1. [9]

    Forter, C. (2017). Two-stage least squares for a/b tests. Retrieved from https://blog.twitch.tv/en/2017/06/30/two-stage-least-squares-for-a-b-tests-669d07f904f7/. Accessed: April 8, 2025

  2. [10]

    Gleser, L. J. and Hwang, J. T. (1987). The nonexistence of 100 (1- )\ The Annals of Statistics , pages 1351--1362

  3. [11]

    Imbens, G. W. and Angrist, J. D. (1994). Identification and estimation of local average treatment effects. Econometrica , 62(2):467--475

  4. [12]

    Kennedy, E. H. (2024). Semiparametric doubly robust targeted double machine learning: a review. Handbook of Statistical Methods for Precision Medicine , pages 207--236

  5. [13]

    S., McCrary, J., Moreira, M

    Lee, D. S., McCrary, J., Moreira, M. J., and Porter, J. (2022). Valid t-ratio inference for iv. American Economic Review , 112(10):3260--3290

  6. [14]

    Ma, Y. (2023). Identification-robust inference for the late with high-dimensional covariates. arXiv preprint arXiv:2302.09756

  7. [15]

    Marsaglia, G. (2006). Ratios of normal variables. Journal of Statistical Software , 16:1--10

  8. [16]

    Mikusheva, A. (2010). Robust confidence sets in the presence of weak instruments. Journal of Econometrics , 157(2):236--247

  9. [17]

    L., Rotnitzky, A., and Robins, J

    Ogburn, E. L., Rotnitzky, A., and Robins, J. M. (2015). Doubly robust estimation of the local average treatment effect curve. Journal of the Royal Statistical Society Series B: Statistical Methodology , 77(2):373--396

  10. [18]

    Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al. (2011). Scikit-learn: Machine learning in python. the Journal of machine Learning research , 12:2825--2830

  11. [19]

    Rotnitzky, A., Smucler, E., and Robins, J. M. (2021). Characterization of parameters with a mixed bias property. Biometrika , 108(1):231--238

  12. [20]

    M., and Rotnitzky, A

    Smucler, E., Robins, J. M., and Rotnitzky, A. (2025). On the asymptotic validity of confidence sets for linear functionals of solutions to integral equations. arXiv preprint arXiv:2502.16673

  13. [21]

    Smucler, E., Rotnitzky, A., and Robins, J. M. (2019). A unifying approach for doubly-robust l1 regularized estimation of causal contrasts. arXiv preprint arXiv:1904.03737

  14. [22]

    and Stock, J

    Staiger, D. and Stock, J. (1997). Instrumental variables regression with weak instruments. Econometrica , 65(3):557--586

  15. [23]

    Stock, J. H. and Wright, J. H. (2000). Gmm with weak identification. Econometrica , 68(5):1055--1096

  16. [24]

    W., Kennedy, E., Kelz, R., and Keele, L

    Takatsu, K., Levis, A. W., Kennedy, E., Kelz, R., and Keele, L. (2023). Doubly robust machine learning for an instrumental variable study of surgical care for cholecystitis. arXiv preprint arXiv:2307.06269

  17. [25]

    Van Der Laan, M. J. and Rubin, D. (2006). Targeted maximum likelihood learning. The international journal of biostatistics , 2(1)

  18. [26]

    Van der Vaart, A. W. (2000). Asymptotic statistics , volume 3. Cambridge university press

  19. [27]

    and Tchetgen Tchetgen, E

    Wang, L. and Tchetgen Tchetgen, E. (2018). Bounded, efficient and multiply robust estimation of average treatment effects using instrumental variables. Journal of the Royal Statistical Society Series B: Statistical Methodology , 80(3):531--550

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.