{"id":"d172cbb4-8e67-4aa6-87ae-7a73342defee","arxiv_id":"2506.10449","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The score confidence set for local average treatment effects is proved to take only six shapes, to match the Wald interval at strong instruments, and to stay valid when instruments are weak, while the DRML estimator fails.","lead":"The paper is a statistical study of a confidence interval for treatment effects with weak instruments, showing it comes in six possible shapes and behaves like the usual Wald interval when instruments are strong. It also proves that the standard double-machine-learning estimator becomes badly biased when instruments are weak, while this score-based interval keeps its promised coverage.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2's weak-instrument limit is internally inconsistent with Lemma 1; the correct limit is (ca Nb - cb Na)/(ca(ca+Na)), not Eq. (7).","rationale":"The paper's central message is that the score confidence set is a reliable weak-instrument-robust tool. That message rests on Ma's uniformity result, Proposition 1, Theorem 1, and Corollary 1. The most fragile load-bearing condition for Theorem 1 is the product-rate assumption (5), but this is a standard high-level condition and it holds in the paper's simulations and applications because mP is known. The place where the paper is least secure is Theorem 2 and its supporting Lemma 2: the stated limiting distribution is not what follows from Lemma 1. This is an internal inconsistency, not a disagreement with consensus. I also checked the fixed-law result independently: despite typographical and scaling oddities in the proof of Theorem 1, the endpoint expansion can be completed and the theorem appears correct. The concrete check above settles the Theorem 2 issue. The reader flagged the same theorem-level issue in the rationale, though their formal weakest_assumption was the product-rate condition, hence partial agreement. The error does not overturn the paper's practical recommendation, so the conditional verdict stands unchanged.","tokens_in":14754,"tokens_out":19028,"duration_ms":201550,"concrete_test":"Re-derive Theorem 2 from Lemma 1 without using Lemma 2's algebraic step: since sqrt(n) Pn psi_a -> Na + ca and sqrt(n) Pn psi_b -> Nb + cb, the continuous mapping theorem gives bphi - phi(Pn) -> (Nb+cb)/(Na+ca) - cb/ca = (ca Nb - cb Na)/(ca(ca+Na)). Compare this limit with Eq. (7); if the two differ for some Sigma_ab with nonzero cross-covariance, Theorem 2 is incorrect as stated. A secondary check is to simulate once with ca = cb = 1 and Sigma_ab = [[1, 0.5], [0.5, 1]] and compare the empirical distribution of the ratio with both formulas.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Lemma 1 states that (Na, Nb), the limit of sqrt(n)(Pn psi_a - EPn psi_a, Pn psi_b - EPn psi_b), is N(0, Sigma_ab). Under Condition 1, sqrt(n) Pn psi_a -> Na + ca and sqrt(n) Pn psi_b -> Nb + cb. Therefore, under Condition 2, bphi - phi(Pn) = Pn psi_b/Pn psi_a - cb/ca -> (Nb+cb)/(Na+ca) - cb/ca = (ca Nb - cb Na)/(ca(Na+ca)). Eq. (7) instead states (ca Nb + cb Na)/(ca(ca - Na)). The two expressions coincide only if Na is redefined as the negative of the Lemma 1 fluctuation and the covariance of the pair is changed; with (Na,Nb) ~ N(0, Sigma_ab) as stated, the theorem's formula is wrong. The proof of Lemma 2 contains an algebraic slip: after multiplying by (1 - sqrt(n)(E psi_a - Pn psi_a)/ca), the displayed equation does not imply the subsequent formula for the ratio minus cb/ca. The qualitative conclusion (non-normality, infinite mean) survives, so the paper's practical warning stands, but the exact theorem is incorrect as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the confidence set obtained by inverting a score test based on the nonparametric influence function for the local average treatment effect under possibly weak instruments. It makes four main contributions: Proposition 1 classifies the possible geometric forms of the score confidence set; Theorem 1 shows that at any fixed law satisfying standard double-machine-learning regularity conditions, the score confidence set coincides with the Wald interval based on the doubly robust estimator up to O_P(1/n) endpoint differences; Corollary 1 draws an optimality conclusion about the diameter of the score set relative to regular asymptotically linear estimators; and Theorem 2 derives a weak-instrument limiting distribution for the doubly robust estimator that is non-normal and has infinite mean. The paper also contains simulations, a real-data application to Glovo experiments, and an implementation in the DoubleML package.","tokens_in":14913,"tokens_out":13087,"duration_ms":159138,"significance":"If the fixed-law results are correct, the paper provides a practically important bridge: a weak-instrument-robust confidence set that, at strong instruments, asymptotically matches the precision of the DRML Wald interval, and that is asymptotically no wider than any regular Wald interval when the efficient influence function is the nonparametric one. The simulation code and DoubleML implementation are concrete and reproducible, and the application to real experiments is useful. The weak-instrument result in Theorem 2 is one of the paper's stated contributions, but its exact statement is internally inconsistent with Lemma 1; the qualitative message of non-normality and infinite mean appears to survive, but the stated limiting law needs correction. For these reasons the paper is a solid methods note provided the weak-instrument theorem is fixed.","major_comments":[{"comment":"Theorem 2 is internally inconsistent with Lemma 1. Lemma 1 defines (Na, Nb) as the limit of sqrt(n)(Pn ψa,ηPn - EPn ψa,ηPn, Pn ψb,ηPn - EPn ψb,ηPn), so under Condition 1 one has sqrt(n) Pn ψa,ηPn -> Na + ca and sqrt(n) Pn ψb,ηPn -> Nb + cb. Consequently, under Condition 2 the limit of the DRML estimator minus the estimand must be (Nb+cb)/(Na+ca) - cb/ca, which equals (ca Nb - cb Na)/(ca(Na+ca)). The expression in Theorem 2, (ca Nb + cb Na)/(ca^2 - ca Na), differs in the sign of the cb Na term and in the sign of Na in the denominator. The same sign error appears in Lemma 2 and in its proof: the factor 1 - sqrt(n)(EPn ψa - Pn ψa)/ca should be 1 + sqrt(n)(Pn ψa - EPn ψa)/ca, since the latter is the object converging to Na. The qualitative conclusion that the limiting law is non-normal with infinite mean survives, because the corrected ratio still has a nondegenerate normal denominator and the Marsaglia heavy-tail argument applies, but the exact limiting law stated in Theorem 2 is not correct as written.","section":"§5, Theorem 2 and Lemma 2"}],"minor_comments":[{"comment":"In the proof of Theorem 1, the displayed equality after (8) states EP{ψb,ηP + φ(P)ψa,ηP}^2 where the correct expression is EP{ψb,ηP - φ(P)ψa,ηP}^2; because only positivity of the displayed quantity is needed, this appears to be a typo, but it should be corrected.","section":"§4, proof of Theorem 1, display (8)"},{"comment":"The notation z^2_{1-α2} appears in the same proof and should be z^2_{1-α/2}.","section":"§4, proof of Theorem 1"},{"comment":"There are several typos: 'strenght' should be 'strength' in Section 5 after Condition 1; 'arbirtrarily' should be 'arbitrarily' in the Introduction; and 'far below the nominal below' should be 'far below the nominal level'.","section":"§5 and Introduction"},{"comment":"The sample sizes are reported as n in {1500, 4500, 7500, 10500, 12000}, but Figure 1's horizontal axis begins at 2000; the axis and the reported grid should be aligned.","section":"§6"},{"comment":"The text promises that the confidence set can take one of six forms, but Proposition 1 lists seven cases; it would be clearer to say that Cases 1-6 are the six geometric types and Case 7 is a degenerate subcase.","section":"Proposition 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a methodological note with an applied component and a software contribution; it fits the journal's scope if the weak-instrument theorem is corrected. The sign inconsistencies in Theorem 2 and Lemma 2 are substantive but local, and the practical conclusion of those results appears to survive. I do not see any issue with the citation pattern or novelty disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. The main result—at any fixed law, the score confidence set equals the DRML Wald interval up to O_P(1/n), with the optimal-diameter corollary—is real and well argued. The weak-instrument distribution for the DRML estimator in Theorem 2 is wrong as stated: a sign inconsistency with Lemma 1 makes the formula incorrect. The qualitative conclusion still survives, but the formula needs fixing.\n\nWhat I like. The shape characterization (Proposition 1) is elementary but useful, especially the observation that infinite diameter occurs exactly when the test of zero average instrument effect on treatment fails to reject. Theorem 1 and Corollary 1 give a clean picture: at strong instruments the score set is essentially the Wald interval, and in models where the efficient influence function is the nonparametric one, no regular estimator's Wald interval can beat it asymptotically. The simulations support this, and the DoubleML implementation plus the Glovo data make the practical intent concrete.\n\nSoft spots. Condition (5) is a standard but strong product-rate requirement, and the simulations never check it: the weak-instrument DGP sets the propensity score to the truth, so nuisance rates are trivial. The claim about optimal diameter is only for the model where the nonparametric influence function is efficient; the paper says this, but it is worth remembering when citing Corollary 1.\n\nThe load-bearing flaw is in Section 5. In Lemma 1, (Na,Nb) is the limit of sqrt(n)(Pn psi_a - EPn psi_a, Pn psi_b - EPn psi_b). Under Condition 1, sqrt(n) Pn psi_a -> ca+Na and sqrt(n) Pn psi_b -> cb+Nb, so the limiting ratio minus cb/ca is (ca Nb - cb Na)/(ca(ca+Na)). The paper's Lemma 2 and Theorem 2 give (ca Nb + cb Na)/(c2a - caNa). The algebra in the proof of Lemma 2 slips at the 'which implies that' step: the rearrangement should leave a factor 1/(1-A), not the expression they write. Since A does not vanish in probability, the stated distribution is not the limit. The non-normality and infinite mean conclusions are untouched because the correct denominator can be arbitrarily close to zero, but the exact formula and any simulation of it need to be re-done.\n\nBottom line: this is a genuine contribution with a solid core in Sections 3–4 and a fixable flaw in Section 5. I would send it to a serious referee, with a request to correct the weak-instrument theorem and to verify whether the conclusions survive without Condition (5) being checked in the simulations. After that fix, I'd happily cite it.","headline":"Score confidence set paper has a solid core in Sections 3–4, but Theorem 2's weak-instrument formula has a sign error; it deserves peer review after that section is fixed.","tokens_in":15540,"tokens_out":6793,"would_cite":true,"duration_ms":63451,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F25","62G05","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The score confidence set for the local average treatment effect takes only six possible shapes, and under strong instruments it coincides with the standard Wald interval up to a 1/n error.","keywords":["local average treatment effect","score confidence set","weak instruments","double machine learning","influence function","uniformly valid confidence sets","Anderson-Rubin test","weak-instrument asymptotics"],"falsifier":"Run the Section 6 weak-instrument simulation with π = 0.15/√n but replace the nuisance fit for the treatment-assignment probability with an estimator whose L2 error decays like $n^{{-1/4}}$, so condition (5) fails: the paper predicts the DRML Wald interval under-covers while the score set holds coverage, so observing the score set also under-cover would contradict the central claim. Alternatively, in a law where EPψa,ηP = 0, Theorem 1 does not apply and the paper predicts Cn is unbounded with high probability, so finding a finite Cn there would falsify the necessity of that condition.","tokens_in":14473,"feed_emoji":"🎯","tokens_out":6769,"duration_ms":73071,"temperature":0.7,"pith_summary":"This paper studies a confidence set for the local average treatment effect that is built by inverting a score test based on the estimated nonparametric influence function. Such a set is attractive because it remains uniformly valid even when the instrument is arbitrarily weak, whereas ordinary Wald intervals can badly under-cover. The authors prove that the set can only take six forms—finite interval, infinite ray(s), the whole line, a point, or empty—and that it has infinite diameter exactly when the data fail to reject the hypothesis that the instrument has no average effect on treatment. Their main theorem shows that at a fixed law with sufficiently fast nuisance estimation, the score set agrees with the doubly robust Wald interval up to endpoints of order 1/n, so in strong-instrument settings it sacrifices no precision; when the efficient influence function equals the nonparametric one, it is asymptotically no wider than any regular Wald interval. The paper also shows that under weak-instrument asymptotics the doubly robust estimator is biased and heavy-tailed, while the score set keeps correct coverage.","feed_headline":"Score set matches Wald interval at strong instruments","feed_subtitle":"The score confidence set stays honest under weak instruments and gives a built-in diagnostic when data are uninformative.","key_machinery":"The central object is the score confidence set Cn = {θ : |Sn(θ)| ≤ z1−α/2}, where Sn(θ) is the empirical mean of the estimated influence function ψb,η − θψa,η over its empirical standard deviation. The key identity is that |Sn(θ)|² ≤ z² is equivalent to aθ² + bθ + c ≤ 0 with coefficients a, b, c built from empirical means of ψa,η, ψb,η, and their squares and products, so the set's geometry is controlled by the sign of ∆ = b² − 4ac and the sign of a. Theorem 1 is carried by proving P(a > 0, ∆ > 0) → 1 and then expanding the two roots of the quadratic to show they match the Wald endpoints up to OP(1/n); the product-rate condition (5) is what makes those empirical coefficients converge to their population limits fast enough.","core_discovery":"The claim is that, at any fixed law P where conditions (2)-(6) hold and EPψa,ηP ≠ 0, the score confidence set Cn equals [bφ − z1−α/2 bσ/√n + OP(1/n), bφ + z1−α/2 bσ/√n + OP(1/n)] with probability tending to one, where bφ is the doubly robust machine learning (DRML) estimator and $bσ^{2}$ estimates the variance of the nonparametric influence function φ1_P(O). Since $bσ^{2}$ estimates the semiparametric efficiency bound whenever the efficient influence function is the nonparametric one, the score set is asymptotically at least as short as any Wald interval from a regular asymptotically linear estimator. In weak-instrument asymptotics the same DRML estimator converges to a ratio of correlated normals with infinite mean, so its Wald interval is not a reliable tool there. The six-form classification follows from viewing Cn as the sublevel set of the quadratic inequality aθ² + bθ + c ≤ 0, whose coefficients are explicit empirical sums of the estimated influence functions.","pith_inferences":["I infer that the six-form classification is not specific to the local average treatment effect: any score test whose statistic is a ratio of an empirical mean to an empirical standard deviation of a single score will produce a confidence set with the same six shapes, giving a general template for weak-instrument-robust inference.","I infer that the product-rate condition (5) could be relaxed if one only wants uniform validity rather than the OP(1/n) endpoint match, but then the asymptotic equivalence with the Wald interval would hold at a slower rate; the paper does not claim this.","I infer that the diagnostic property (infinite diameter iff non-rejection of the zero-effect test) is an honest alternative to first-stage F-statistics for detecting weak instruments, because it directly targets identifiability of the estimand rather than an arbitrary threshold."],"forward_implications":["At strong instruments, practitioners can use the score confidence set instead of the DRML Wald interval without asymptotically sacrificing precision, because the two sets agree to OP(1/n).","An unbounded score set acts as a direct diagnostic: it occurs exactly when the test of zero average instrument effect on treatment does not reject, signaling that the data are uninformative about the local average treatment effect.","In models where the instrument assignment probability is known, such as randomized encouragement designs, the score set is asymptotically no wider than the Wald interval from any regular estimator, so it inherits efficiency.","The DRML point estimator should not be reported under weak instruments: its limiting distribution is a ratio of normal variables with infinite mean, so its Wald interval can be misleadingly short.","The provided algorithm and software implementation make the score confidence set immediately usable in double machine learning pipelines."],"supporting_citations":[{"why":"Introduced the score confidence set and proved its uniform validity over models with arbitrarily weak instruments.","marker":"Ma (2023)"},{"why":"Defined the double/debiased machine learning framework and the estimator bφ whose asymptotic linearity underlies Theorem 1.","marker":"Chernozhukov et al. (2018)"},{"why":"Established the regular asymptotic linearity of bφ under conditions (2)-(6), which Theorem 1 directly invokes.","marker":"Takatsu et al. (2023)"},{"why":"Derived the weak-instrument asymptotic distribution of two-stage least squares, the benchmark that Theorem 2 extends to nonparametric models.","marker":"Staiger and Stock (1997)"},{"why":"Proved the analogous six-form classification for Anderson-Rubin confidence sets, which Proposition 1 mirrors.","marker":"Mikusheva (2010)"},{"why":"Supplies the definitions and convolution theorem for regular asymptotically linear estimators used in Corollary 1.","marker":"Van der Vaart (2000)"},{"why":"Provides the distribution theory for ratios of normal variables that implies the infinite mean in Theorem 2.","marker":"Marsaglia (2006)"},{"why":"Shows finite-length confidence sets cannot be uniformly valid over weak-instrument models, justifying the unbounded forms.","marker":"Gleser and Hwang (1987)"}],"fun_headline_variants":["Score confidence set takes six possible forms.","Under strong instruments, score set matches Wald.","Weak instruments do not break score confidence set.","New algorithm computes score confidence set in DoubleML.","Score set is optimal at strong, robust at weak instruments."],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the product of the L2 errors of the estimated treatment-assignment probability and the slower of the estimated outcome and treatment regressions converges faster than $n^{{-1/2}}$; if the machine learning nuisance estimates converge too slowly, the score set may no longer coincide with the doubly robust Wald interval.","fun_headline_variants_meta":{"raw":{"variants":["Score confidence set takes six possible forms.","Under strong instruments, score set matches Wald.","Weak instruments do not break score confidence set.","New algorithm computes score confidence set in DoubleML.","Score set is optimal at strong, robust at weak instruments."]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000485,"raw_usage":{"total_tokens":2454,"prompt_tokens":1070,"completion_tokens":1384,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":686,"completion_tokens_details":{"reasoning_tokens":1315}},"tokens_in":686,"tokens_out":1384,"duration_ms":16305,"temperature":1.0,"reasoning_tokens":1315,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:27:36.720893+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Section 6 weak-instrument simulation with π = 0.15/√n but replace the nuisance fit for the treatment-assignment probability with an estimator whose L2 error decays like $n^{{-1/4}}$, so condition (5) fails: the paper predicts the DRML Wald interval under-covers while the score set holds coverage, so observing the score set also under-cover would contradict the central claim. Alternatively, in a law where EPψa,ηP = 0, Theorem 1 does not apply and the paper predicts Cn is unbounded with high probability, so finding a finite Cn there would falsify the necessity of that condition.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defined the double/debiased machine learning framework and the estimator bφ whose asymptotic linearity underlies Theorem 1."},{"cited_title":"Doubly robust machine learning for an instrumental variable study of surgical care for cholecystitis","cited_arxiv_id":"2307.06269","evidence_quote":"Established the regular asymptotic linearity of bφ under conditions (2)-(6), which Theorem 1 directly invokes."},{"cited_title":"and Stock, J","cited_arxiv_id":null,"evidence_quote":"Derived the weak-instrument asymptotic distribution of two-stage least squares, the benchmark that Theorem 2 extends to nonparametric models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Proved the analogous six-form classification for Anderson-Rubin confidence sets, which Proposition 1 mirrors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the distribution theory for ratios of normal variables that implies the infinite mean in Theorem 2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows finite-length confidence sets cannot be uniformly valid over weak-instrument models, justifying the unbounded forms."}],"review_version":1}