REVIEW 3 major objections 4 minor 43 references
A generalized plurality rule identifies multiple treatment effects with invalid instruments, and a sampling confidence interval keeps nominal coverage at parametric rate.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 07:18 UTC pith:RQ23EEP6
load-bearing objection Solid extension of single-treatment IV robustness to multiple exposures, with a genuinely new sampling CI, but the parametric-rate length claim leans on an untestable high-level condition (Assumption 3) that can fail; coverage is fine. the 3 major comments →
Identification and Robust Inference for Multiple Treatment Effects with Possibly Invalid Instruments
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central claim is that for p_d ≥ 1 treatments and p_z instruments, the vector of causal effects is identified by the generalized plurality rule, which says the true effect β* receives votes from more instruments than any other candidate effect, where each instrument 'votes' for every candidate on its hyperplane L_k = {β : Υ_{k·}(β−β*) = π_k}. A sufficient, easy-to-check condition is the generalized majority rule: |V*| > (|S*| + p_d − 1)/2, where V* is the set of truly valid instruments and S* the set of relevant ones. For inference, the paper constructs a sampling confidence interval (SCI): for each just-identifying subset it estimates the validity of every instr
What carries the argument
The key geometric object is the hyperplane voting interpretation of instruments: in a multiple-treatment linear IV model, each relevant instrument defines a (p_d−1)-dimensional hyperplane in the p_d-dimensional effect space, and an instrument votes for any candidate effect lying on its hyperplane. The generalized plurality rule (Condition 2) requires β* to lie on more hyperplanes than any other candidate; Theorem 1 shows that the count condition |V*| > (|S*| + h_0 − 1)/2 suffices, where h_0 encodes the rank structure of the relevance matrix and reduces to p_d in the benign case. For inference, the mechanism is 'perturb-and-aggregate': Gaussian perturbations of the validity estimates generate
Load-bearing premise
The procedure's parametric-rate length guarantee rests on Assumption 3's inequality (39), which says that for every just-identifying subset with strongly invalid instruments, at most (|S*|+p_d−1)/2 instruments can have centering bias inside the local-to-zero window C_1 log(C_2/α0)/√n.
What would settle it
Rerun the paper's own simulation design (S1) with n=5000 and τ=0.05, where one instrument is locally invalid and hard thresholding undercovers badly: the sampling interval must cover the true effect at about the nominal 95% level while its length shrinks with n. If empirical coverage is materially below 95% or the interval length does not shrink at the 1/√n rate, Theorems 2 and 3 are falsified. Alternatively, construct a design satisfying (39) but with p_d=2, p_z=9, and 5 valid instruments, and check that the SCI length remains O(1/√n).
If this is right
- Applied researchers can report honest confidence intervals for multiple endogenous treatments without knowing which instruments are valid, as long as the valid instruments exceed the majority bound.
- The method works with summary-level Mendelian randomization data, where only SNP-exposure and SNP-outcome coefficients and standard errors are available.
- The SCI avoids a full grid search over the effect space by searching over just-identifying subsets, keeping the computation feasible for moderate dimensions.
- The parametric-rate length guarantee means the procedure does not achieve coverage by inflating the interval width.
- For a single treatment the generalized majority rule reduces to the classical majority rule, so the results unify single- and multi-treatment IV inference.
Where Pith is reading between the lines
- The paper leaves open whether replacing the Gaussian perturbation with a bootstrap or wild bootstrap would improve finite-sample coverage; that trade-off is not explored.
- The same perturb-and-aggregate scheme could be adapted to other post-selection inference problems, such as confidence intervals after moment selection in GMM, wherever a plurality condition can be formulated.
- If the local plurality condition (39) fails, the length guarantee breaks down; an adaptive widening of the SCI could restore honesty at the cost of conservatism.
- Proposition 1 suggests that under random coefficients, the generalized plurality rule holds almost surely once |V*| > p_d, implying that the majority rule may be unnecessarily restrictive for identification.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the linear instrumental variable model with multiple endogenous treatments and possibly invalid instruments. It introduces generalized plurality and majority rules for identifying the vector of treatment effects, emphasizing the geometric distinction from the single-treatment case. For inference, it proposes a Sampling Confidence Interval (SCI) that aggregates TSLS intervals over perturbed estimates of instrument validity, with the aim of remaining valid even when hard-thresholding cannot separate locally invalid from valid instruments. The main theoretical results are Theorem 2 (asymptotic coverage of the SCI under a generalized majority condition) and Theorem 3 (parametric-rate length under Assumption 3). The paper also reports simulations and a Mendelian randomization application.
Significance. The identification results are a useful extension of single-treatment plurality logic to multiple treatments, with a clear geometric interpretation and a transparent sufficient condition. The sampling scheme is a nontrivial generalization of Guo (2023) that avoids multidimensional grid search, and the supplementary proofs are detailed. However, the advertised parametric-rate length guarantee rests entirely on Assumption 3, Eq. (39), a high-level local-plurality condition that is neither derived from primitive conditions nor verifiable from data. In addition, the adaptive tuning of C0 in Remark 4 does not match the fixed constant used in the proof of Proposition 3. These issues prevent the paper from fully establishing its central efficiency claim, although the coverage theorem is more robust.
major comments (3)
- [Section 4, Assumption 3, Eq. (39); Theorem 3] The parametric-rate length guarantee (41) is obtained only under a high-level condition not implied by the identification conditions. A concrete failure: |S*|=10, p_d=2, |V*|=6, so the generalized majority rule (16) holds. Let five invalid IVs pass exactly through a false effect beta_f, and two further invalid IVs have |pi_k^(ell)| <= C1 log(C2/alpha0)/sqrt(n) for the subset ell at beta_f. Then Condition 2 holds, but the left side of (39) is at least 7 > 5.5, so Assumption 3 fails. The corresponding TSLS interval is then included in the SCI and is centered near beta_f, so Len(alpha)=O(1), not O(log/sqrt(n)). Thus the length claim is not a byproduct of identification; it is an extra, untestable local-plurality assumption. Please derive (39) from primitives or reframe the length guarantee as conditional.
- [Section 3.3, Remark 4 and proof of Proposition 3] The proof sets a specific constant C0 in rho_n(M) that depends on alpha0 and the dimension, but Algorithm 1 uses an adaptive rule: start at 0.05 and multiply by 1.25 until more than 5% of the generalized-majority selections are nonempty. No argument shows that this adaptive choice preserves the event in Proposition 3 or the coverage/length bounds. The simulations and application implement the adaptive rule, so the reported performance is not the procedure for which theorems are proved. Either prove guarantees for the adaptive C0 or fix C0 as in the proof and report sensitivity.
- [Section 4, Assumption 2 and Remark 5] Assumption 2(b) requires every just-identifying subset H in H* to have eigenvalue lambda_min(Upsilon_H^T Upsilon_H) bounded away from zero. This is much stronger than Condition 1, which only requires the full valid set to have full column rank. In Mendelian randomization, valid IVs are often weak individually, so there may exist just-identifying valid subsets with very small eigenvalues under Condition 1. Under such subsets, the TSLS estimator is not sqrt(n)-consistent, and Proposition 2, Proposition 3, and Theorem 2 are unavailable. Remark 5 acknowledges weak identification is out of scope, but the abstract and title do not state this restriction. It should be stated prominently.
minor comments (4)
- [Section 3.3, Eq. (28)] The symbol M is used both for the number of resamples and for the filtered set in (28). Use a distinct notation such as N to avoid confusion.
- [Section 3.2, Eq. (27)] The definition of locally invalid IVs uses the TSHT candidate set Lhat from (26), but the sampling procedure later uses different candidate sets bT_m. Clarify whether the concept is relative to TSHT or to the SCI procedure.
- [Section 5-6] The method enumerates all subsets of size p_d from S_hat. For small p_z this is fine, but the paper does not discuss computational feasibility for larger p_z. A brief complexity discussion would be helpful.
- [Section 4, Assumption 3] The constants C1 and C2 in (38)-(39) are introduced as absolute constants but never specified or connected to primitive parameters. The theorem would be more informative if these constants were made explicit or shown to depend only on identifiable quantities.
Circularity Check
No significant circularity: the main theorems are derived from stated assumptions and standard asymptotic arguments; the high-level Assumption 3 is an extra regularity condition, not a restatement of the conclusion.
full rationale
The paper's central claims are not circular. Theorems 2 and 3 are proved from Proposition 3 plus standard TSLS asymptotics and union-bound arguments. Proposition 3 itself is a genuine probabilistic result: it shows that by sampling normal perturbations of the estimated validity vector, with high probability one perturbation lands within a shrinking neighborhood of the true π*, so the corresponding selected IV set contains all valid IVs. This does not assume the coverage or length conclusion. The proof of Theorem 2 then shows that, conditional on this event, at least one TSLS interval in the union SCI(α) is valid, and hence the union covers β*1; this is standard union-inference logic, not a tautology. Theorem 3's parametric-rate length bound does rely on Assumption 3, especially inequality (39), but the paper transparently labels this as an additional assumption rather than deriving it from the conclusion. Indeed, the paper states that (40) is 'slightly stronger than the generalized plurality rule in Condition 2, accounting for the variable selection error in finite samples.' Whether (39) is primitive, verifiable, or too strong is a legitimate robustness/assumption-strength concern, not evidence of circularity. The self-citations to Guo et al. (2018) and Guo (2023) are used as prior baselines, comparisons, and tuning heuristics; the identification theorems and the inference proofs do not reduce to these citations. In particular, the sampling idea is generalized from Guo (2023), but the coverage and length guarantees are proved here directly from the paper's own conditions, not imported from Guo (2023). No fitted parameter is renamed as a prediction, and no uniqueness theorem from the authors' prior work is invoked to force the choice of procedure. The real-data application and simulations are external checks rather than part of the derivation. Therefore there is no specific reduction of a claimed result to its own inputs, and the appropriate circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- C0 (ρ_n(M) scaling) =
Theory: 0.5·(2/c(α0))^{1/|S*|}; implementation: start 0.05, multiply by 1.25 until >5% nonempty (Remark 4)
- α0 (screening level) =
α/20 (Remark 4)
- M (number of resamples) =
1000 (Remark 4)
- C* (TSHT threshold constant) =
0.5 in the main simulations
axioms (7)
- domain assumption Linear, additive, constant-effect potential outcome model (1): Y(d,z,x)−Y(d',z',x') = (d−d')'β* + (x−x')'ξ* + (z−z')'κ*, with E(Y(0,0,0)|Z,X) = Z'η*+X'ζ*.
- domain assumption Assumption 1: i.i.d. sub-Gaussian (W_i, v_i) with bounded eigenvalue spectra.
- domain assumption Assumption 2: strong relevance: max_j |Υ*_k,j| > c_Υ for every k∈S*, and min_{H∈H*} λ_min(Υ_H*'Υ_H*) > c_Υ.
- domain assumption Condition 1: rank(Υ*_{V*·}) = p_d.
- domain assumption Identification condition: Condition 2 (generalized plurality) or sufficient majority condition (15) |V*| > (|S*|+h0−1)/2.
- ad hoc to paper Assumption 3, Eq. (39): for each strongly invalid just-identifying subset ℓ, the number of IVs with |π^(ℓ)_k| ≤ C_1 log(C_2/α0)/√n is at most (|S*|+p_d−1)/2.
- domain assumption Summary-data independence: SNPs independent, and bΓ and bΥ (and their entries) asymptotically independent (Section D).
read the original abstract
The instrumental variable (IV) method is widely used to infer causal effects in observational studies with unmeasured confounding, but invalid instruments can compromise both population identification and finite-sample inference. This paper studies linear IV models with multiple endogenous treatments and possibly invalid instruments. Identification is more delicate than in the single-treatment setting because a single instrument no longer identifies a scalar candidate effect; instead, each relevant instrument defines a hyperplane in the multidimensional effect space. For identification of multiple treatment effects, we introduce generalized plurality and majority rules which require a sufficiently large number of IVs to be valid. For inference, data-dependent instrument selection may fail to separate certain invalid IVs from valid ones, leading to undercoverage of confidence intervals when these invalid instruments are mistakenly selected as valid. We propose a sampling confidence interval for each treatment effect, which is robust to IV selection errors. We establish asymptotic coverage and parametric-rate length of our sampling confidence interval under regularity conditions and illustrate this method in a Mendelian randomization application.
Figures
Reference graph
Works this paper leans on
-
[1]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
Confidence intervals for causal effects with invalid instruments by using two-stage hard thresholding with voting , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2018 , publisher=
2018
-
[2]
2013 , publisher=
Discovery and refinement of loci associated with lipid levels , journal=. 2013 , publisher=
2013
-
[3]
Cell Genomics , volume=
Deciphering proteins in Alzheimer’s disease: A new Mendelian randomization method integrated with AlphaFold3 for 3D structure prediction , author=. Cell Genomics , volume=. 2024 , publisher=
2024
-
[4]
American journal of epidemiology , volume=
Efficient design for Mendelian randomization studies: subsample and 2-sample instrumental variable estimators , author=. American journal of epidemiology , volume=. 2013 , publisher=
2013
-
[5]
International journal of epidemiology , volume=
Mendelian randomization with invalid instruments: effect estimation and bias detection through Egger regression , author=. International journal of epidemiology , volume=. 2015 , publisher=
2015
-
[6]
Biometrika , volume=
Using negative controls to identify causal effects with invalid instrumental variables , author=. Biometrika , volume=. 2025 , publisher=
2025
-
[7]
Biometrika , volume=
Semiparametric efficient G-estimation with invalid instrumental variables , author=. Biometrika , volume=. 2023 , publisher=
2023
-
[8]
Journal of Machine Learning Research , volume=
Selective machine learning of the average treatment effect with an invalid instrumental variable , author=. Journal of Machine Learning Research , volume=
-
[9]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
GENIUS-MAWII: for robust Mendelian randomization with many weak invalid instruments , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2024 , publisher=
2024
-
[10]
Journal of the American Statistical Association , pages=
Robust inference for federated meta-learning , author=. Journal of the American Statistical Association , pages=. 2025 , publisher=
2025
-
[11]
Annual Review of Statistics and Its Application , volume=
Identification and inference with invalid instruments , author=. Annual Review of Statistics and Its Application , volume=. 2024 , publisher=
2024
-
[12]
Robust Mendelian Randomization Analysis by Automatically Selecting Valid Genetic Instruments with Applications to Identify Plasma Protein Biomarkers for Alzheimer’s Disease , author=
-
[13]
Journal of Applied Econometrics , year=
Agglomerative hierarchical clustering for selecting valid instrumental variables , author=. Journal of Applied Econometrics , year=
-
[14]
Journal of the Royal Statistical Society Series B: Statistical Methodology , pages=
On the instrumental variable estimation with many weak and invalid instruments , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , pages=. 2024 , publisher=
2024
-
[15]
Journal of Business & Economic Statistics , volume=
Using heteroscedasticity to identify and estimate mismeasured and endogenous regressor models , author=. Journal of Business & Economic Statistics , volume=. 2012 , publisher=
2012
-
[16]
Proceedings of the National Academy of Sciences , volume=
Mendelian randomization for causal inference accounting for pleiotropy and sample structure using genome-wide summary statistics , author=. Proceedings of the National Academy of Sciences , volume=. 2022 , publisher=
2022
-
[17]
Journal of Econometrics , volume=
Select the valid and relevant moments: An information-based LASSO for GMM with many moments , author=. Journal of Econometrics , volume=. 2015 , publisher=
2015
-
[18]
arXiv preprint arXiv:2310.08063 , year=
Uniform Inference for Nonlinear Endogenous Treatment Effects with High-Dimensional Covariates , author=. arXiv preprint arXiv:2310.08063 , year=
-
[19]
Journal of Business & Economic Statistics , volume=
A heteroscedasticity-robust overidentifying restriction test with high-dimensional covariates , author=. Journal of Business & Economic Statistics , volume=. 2025 , publisher=
2025
-
[20]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
The confidence interval method for selecting valid instrumental variables , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2021 , publisher=
2021
-
[21]
Review of Economics and Statistics , volume=
Endogenous treatment effect estimation with a large and mixed set of instruments and control variables , author=. Review of Economics and Statistics , volume=. 2024 , publisher=
2024
-
[22]
Journal of the American statistical Association , volume=
Instrumental variables estimation with some invalid instruments and its application to Mendelian randomization , author=. Journal of the American statistical Association , volume=. 2016 , publisher=
2016
-
[23]
2005 , publisher=
Matrix Algebra , author=. 2005 , publisher=
2005
-
[24]
2012 , publisher=
Matrix Analysis , author=. 2012 , publisher=
2012
-
[25]
Econometrica , volume=
Consistent moment selection procedures for generalized method of moments estimation , author=. Econometrica , volume=. 1999 , publisher=
1999
-
[26]
Journal of the American Statistical Association , volume=
Sensitivity analysis for instrumental variables regression with overidentifying restrictions , author=. Journal of the American Statistical Association , volume=. 2007 , publisher=
2007
-
[27]
arXiv preprint arXiv:2301.00718 , year=
Robust inference for federated meta-learning , author=. arXiv preprint arXiv:2301.00718 , year=
-
[28]
Journal of the American Statistical Association , pages=
Statistical inference for maximin effects: Identifying stable associations across multiple studies , author=. Journal of the American Statistical Association , pages=. 2023 , publisher=
2023
-
[29]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
Post-selection problems and a solution using searching and sampling , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2023 , publisher=
2023
-
[30]
Journal of Econometrics , volume=
The validity of instruments revisited , author=. Journal of Econometrics , volume=. 2012 , publisher=
2012
-
[31]
Quantitative Economics , volume=
Simple and honest confidence intervals in nonparametric regression , author=. Quantitative Economics , volume=. 2020 , publisher=
2020
-
[32]
Statistical Science , volume=
The GENIUS approach to robust Mendelian randomization inference , author=. Statistical Science , volume=. 2021 , publisher=
2021
-
[33]
arXiv preprint arXiv:2012.14823 , year=
Bias-aware inference in regularized regression models , author=. arXiv preprint arXiv:2012.14823 , year=
Pith/arXiv arXiv 2012
-
[34]
Quantitative Economics , volume=
Sensitivity analysis using approximate moment condition models , author=. Quantitative Economics , volume=. 2021 , publisher=
2021
-
[35]
Journal of Econometrics , pages=
Testing underidentification in linear models, with applications to dynamic panel and asset pricing models , author=. Journal of Econometrics , pages=. 2021 , publisher=
2021
-
[36]
Econometric theory , volume=
Testing identifiability and specification in instrumental variable models , author=. Econometric theory , volume=. 1993 , publisher=
1993
-
[37]
Journal of the American Statistical Association , volume=
On the use of the lasso for instrumental variables estimation with some invalid instruments , author=. Journal of the American Statistical Association , volume=. 2019 , publisher=
2019
-
[38]
arXiv preprint arXiv:2208.05278 , year=
Selecting Valid Instrumental Variables in Linear Models with Multiple Exposure Variables: Adaptive Lasso and the Median-of-Medians Estimator , author=. arXiv preprint arXiv:2208.05278 , year=
-
[39]
Detecting Invalid Instruments using L_1 -
Han, Chirok , journal=. Detecting Invalid Instruments using L_1 -
-
[40]
American Journal of Epidemiology , volume=
Multivariable Mendelian Randomization: The Use of Pleiotropic Genetic Variants to Estimate Causal Effects , author=. American Journal of Epidemiology , volume=. 2015 , publisher=
2015
-
[41]
Structural variation in amyloid-
Qiang, Wei and Yau, Wai-Ming and Lu, Jun-Xia and Collinge, John and Tycko, Robert , journal=. Structural variation in amyloid-. 2017 , publisher=
2017
-
[42]
Nature Genetics , volume=
A unified framework for joint-tissue transcriptome-wide association and Mendelian randomization analysis , author=. Nature Genetics , volume=. 2020 , publisher=
2020
-
[43]
American Economic Review , pages=
Contamination Bias in Linear Regressions , author=. American Economic Review , pages=. 2024 , publisher=
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.