Pith. sign in

REVIEW 2 major objections 3 minor 3 references

Possibilistic Instrumental Variable Regression with Potentially Invalid Instruments

T0 review · 2 major / 3 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read This paper proves that a possibilistic instrumental-variable method keeps finite-sample coverage of causal-effect intervals even when the only available instrument is invalid, as long as the analyst's violation set contains the true exogene

desk verdict New possibilistic IV sensitivity method with a clean construction, but the finite-sample coverage guarantee in Prop 2 is not actually proven. read the letter →

arxiv 2511.16029 v3 pith:63VIVQ6N submitted 2025-11-20 stat.ME econ.EMmath.STstat.TH

classification stat.MEecon.EMmath.STstat.TH
keywords instrumentalvariablespossibilitytheoryinvalidinstrumentsexogeneityviolationsfinite-samplecoveragepartialidentificationsensitivityanalysisposterior
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish a method for instrumental-variable regression that does not require the analyst to bet on whether the instrument is valid. The core idea is to perform posterior inference on the treatment effect conditionally on a user-specified set of possible exogeneity violations, using possibility theory rather than standard probability. The key theoretical result is a finite-sample coverage guarantee: if the violation set contains the true violation, the calibrated ('validified') uncertainty intervals for the treatment effect are valid confidence sets at any level, even with a single potentially invalid instrument. A sympathetic reader would care because validity of instruments is often uncertain, and existing approaches either require many instruments or lose coverage when instruments fail. The paper also gives practical approximations and demonstrates on two real datasets that qualitative conclusions can survive plausible violations.

What carries the argument

The central object is the conditional posterior possibility f(β | α∈A, W): a curve over treatment-effect values β obtained by maximising the structural posterior possibility over all α in the violation set A and over the covariance Σ. For each β, the best α is the projection of t(β) = γ̂1 − βγ̂2 onto A under the Z'Z metric; values of β whose implied α can lie inside A get possibility 1, forming the partial-identification plateau. The validification transform then converts this raw posterior into π_w(β|A) = P_β(f(β|α∈A,W) ≤ f(β|α∈A,W=w)), which is the probability that a resampled dataset produces a posterior no larger than the observed one. Because this is a probability integral transform, it

What would settle it

Take a single-instrument design with true α = 0.05, set A = [−0.1, 0.1], and compute the Monte Carlo validified 95% interval over thousands of datasets at n = 50. If empirical coverage falls materially below 0.95, the finite-sample guarantee does not survive the plug-in/Monte Carlo implementation; if it stays at or above 0.95, the practical method matches the theory.

Watch

Extended reading notes

Core claim

The paper's central claim is Proposition 2: after applying the validification transform to the conditional posterior possibility, for any δ in [0,1] one has sup_β P_β(π_W(β|A) ≤ δ) ≤ δ whenever A contains the true value of the exogeneity-violation vector α. In other words, the upper level sets of the validified posterior possibility are valid 100(1−δ)% confidence sets for the treatment effect β in finite samples. The posterior itself is built by mapping the reduced-form estimates to structural parameters through a projection of t(β) = γ̂1 − βγ̂2 onto the violation set A under the metric induced by Z'Z; values of β whose implied α can stay inside A receive the highest possibility, forming a p

Load-bearing premise

Everything rests on the user-specified violation set A actually containing the true exogeneity violation α; this is untestable from the data, and the paper's own simulations show coverage collapsing when the set is wrong.

Editorial extensions

If this is right

  • If the violation set A contains the true violation vector, the reported 100(1−δ)% uncertainty intervals cover the true treatment effect with at least nominal frequency in finite samples, regardless of whether any instrument is actually valid.
  • The method yields a genuine posterior possibility function (possibly diffuse) for every violation set, so sensitivity analyses do not require a binary choice between valid and invalid instruments.
  • Widening A trades validity against informativeness: intervals remain valid but become more conservative, and a sufficiently large A makes inference completely uninformative.
  • A Monte Carlo approximation and a cheaper χ² approximation are provided; the paper argues the Monte Carlo version is preferable when β is not point-identified because the Gaussian approximation mishandles tails.
  • In the real-data examples, allowing plausible violation bounds leaves qualitative conclusions intact, while sufficiently large violation sets can erase significance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct but unstated consequence of the validification argument is that the same construction could yield simultaneous confidence bands or family-wise error control over multiple hypotheses about β, provided A is fixed before seeing the data.
  • The method's usefulness hinges on choosing A; a natural extension the paper leaves open is data-dependent selection of A with split-sample or Bonferroni-type corrections to preserve the finite-sample guarantee.
  • In multi-instrument settings, the projection-based mode corresponds to a 'least-violation' point estimate of β; this could be developed into an identification-robust test statistic, though the paper stops at interval estimation.
  • The paper's own simulations show the χ² approximation can undercover when the true α sits on a corner of the hypercube; a practical reading is to treat χ² as a screening device and use the Monte Carlo version for final reports.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. This paper proposes a possibilistic approach to instrumental variable regression when instruments may be invalid. The model is a linear structural equation with a direct-effect parameter α; the analyst specifies a set A of possible α values. The authors derive a closed-form expression for the posterior possibility of the treatment effect β conditional on α∈A (Prop. 1), then apply the 'validification' transform of Martin and Liu to define a validified posterior possibility π_w(β|A). Proposition 2 claims a finite-sample coverage guarantee for the upper-level sets of π_w whenever A contains the true α; Corollary 1 states the resulting confidence sets. Simulations with one and five instruments compare coverage with TSLS, PGMM, BudgetIV, gIVBMA, and CIIV, and two real-data examples illustrate the method.

Significance. If the coverage theorem is correct, the paper is a valuable contribution: it provides a principled way to handle a single potentially invalid instrument, yields intervals that are valid in finite samples under a user-specified violation set, and has computational benefits because the possibilistic posterior has closed form. The authors are transparent about the need for the violation set to contain the true α and about the trade-off between validity and informativeness. The paper includes reproducible code and a thorough comparison with existing methods. However, the central finite-sample guarantee (Prop. 2) is not established by the proof as written; the decisive monotonicity step is asserted without justification. Consequently, the contribution is conditional on a rigorous proof or a weaker statement.

major comments (2)
  1. [Appendix A.3] The proof of Prop. 2 hinges on the assertion that π_w(β|A) ≥ π_w(β|{α0}) for A⊇{α0}. This is not proven. The function f(β|α∈A,W) is a ratio of suprema over α, so both numerator and denominator change with A; pointwise monotonicity of the ratio is not immediate. Even if the ratio were pointwise increasing, the event in (4), {f_A(W) ≤ f_A(w)}, can be invariant under monotone transformations (e.g., if f_A = h(f_{α0}) with h increasing), so the inequality π_A ≥ π_{α0} does not follow. Since Corollary 1 and the finite-sample coverage claim in the abstract rest on this step, the proof is incomplete.
  2. [Section 3.3] The MC approximation of π_w requires sampling from P_β, but P_β is not fully specified: the sampling distribution of W depends on α, γ2, and Σ in addition to β. In practice one must plug in estimates, and the paper does not show that the finite-sample guarantee of Prop. 2 survives this plug-in step. The χ² approximation is asymptotic and cannot inherit the finite-sample guarantee either. The simulations appear to use the true nuisance parameters for the MC samples, so the empirical coverage does not validate the practical implementation. Please clarify how P_β is constructed and whether the reported coverage applies to the feasible version.
minor comments (3)
  1. [Section 2, Eq. (4)] The notation P_β is used without specifying the full parameter vector; the distribution of W also depends on α, γ2, and Σ. Please state the dependence explicitly or define a profile/plug-in distribution.
  2. [Section 4.1] The description of the MC approximation does not report the number of Monte Carlo samples M used. Please specify.
  3. [Appendix A.2] Figure A.1 is helpful but is not referenced in the main text; consider adding a cross-reference.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the central coverage claim is an external validification/PIT property, not a fitted input or self-referential derivation.

full rationale

The load-bearing coverage argument (Prop. 2 and Cor. 1, Appendix A.3) is not a fitted-input prediction: the validified posterior π_w(β|A) is defined in (4) as a probability integral transform under the true β-measure, and the proof for A={α0} invokes the standard first-order stochastic dominance of the PIT over Uniform (Casella and Berger), an external, non-self-cited result. No constants are tuned to force the simulated coverage, and the violation set A is user-specified rather than estimated from the data. The possibility-theory background is cited to standard literature (Zadeh; Dubois and Prade) and to co-author expositions (Houssineau), but those citations are background, not the basis of Prop. 2; the authors' own gIVBMA is used only as a comparator. The real weakness is not circularity: Appendix A.3 asserts without proof the monotonicity step 'the validified posterior becomes no more informative, i.e., π_w(β|A) ≥ π_w(β|{α0})' for A⊇{α0}. This is an omitted proof and a genuine correctness risk, but an unproved lemma is not a reduction of the conclusion to its inputs. Similarly, the practical MC/χ² approximations that rely on plug-in estimates are unsupported by the finite-sample theorem, but they are not fitted inputs renamed as predictions. The paper's own simulations showing coverage collapse when A is misspecified are the expected consequence of the theorem's explicit conditioning on A containing the true α, not a hidden circular step. No circular step meeting the quoted-evidence standard was found.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The method rests on a standard IV model and on the Martin-Liu validification framework. The only truly fragile input is the user's violation set A, which must contain the true α for the coverage theorem; the paper is explicit about this trade-off. No new entities or fitted constants are introduced.

free parameters (1)
  • Violation set A = user-specified (e.g., [−0.1, 0.1])
    The main input; coverage only holds if it contains the true α. Not fitted to data, but a hand-chosen tolerance.
assumptions (4)
  • domain assumption Structural model: Y_i = βX_i + Z_i α + ε_i, X_i = Z_i γ_2 + η_i with (ε,η) jointly Gaussian.
    Underlies all derivations; the reduced-form and the identification relation (2) follow from it.
  • standard math Validification transform (Martin & Liu 2013) makes π_w strongly valid for a singleton violation set.
    Prop. 2's singleton-case result is imported from the inferential models literature; the paper cites it rather than reproving it.
  • domain assumption The user-specified violation set A contains the true α_0.
    Core conditional assumption; without it the coverage guarantee fails, as the simulations with A={0} under α=0.5 show.
  • domain assumption Regularity: Z'Z invertible, γ_2≠0, and the model is correctly specified.
    Needed for the closed-form OLS/projection formulas and to define the reduced form.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Possibilistic Instrumental Variable Regression with Potentially Invalid Instruments." pith.science (2026). https://pith.science/paper/63VIVQ6N

@misc{pith2026251116029,
  author       = {Pith},
  title        = {Pith review of: Possibilistic Instrumental Variable Regression with Potentially Invalid Instruments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/63VIVQ6N}},
  note         = {Machine review of arXiv:2511.16029}
}
abstract

Instrumental variable regression is a common approach for causal inference in the presence of unobserved confounding. However, identifying valid instruments is often difficult in practice. In this paper, we propose a novel method based on possibility theory that performs posterior inference on the treatment effect, conditional on a user-specified set of potential violations of the instrument exogeneity assumption. Our method can provide valid results even when only a single, potentially invalid, instrument is available. Crucially, and in contrast with existing methods, we prove a finite-sample coverage guarantee for the exactly calibrated (validified) uncertainty intervals when the violation set contains the true value, and we provide practical MC/$\chi^2$ approximations. Simulation experiments and real-data applications indicate strong performance of the proposed approach.

Figures

Figures reproduced from arXiv: 2511.16029 by the authors.

Figure 1
Figure 1. The effect of institutions on economic growth: Validified posterior possibility functions under a perfectly valid instrument (α = 0) and potential violations A = [−0.1, 0.1] and A = [−0.4, 0.4]. The solid line is based on the χ 2 approximation, while the dashed line displays the Monte Carlo approximation. The dashed grey line indicates the 0.05 level, such that the 95% uncertainty interval for β includes all values … view at source ↗
Figure 2
Figure 2. The returns to schooling: Validified posterior possibility functions under a perfectly valid instrument (α = 0) and allowing for potential violations A = [0, 0.02] and A = [0, 0.04]. The solid line is based on the χ 2 approximation, while the dashed line displays the Monte Carlo approximation. The dashed grey line indicates the 0.05 level, such that the 95% uncertainty interval for β includes all values where the po… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

3 extracted references · 2 linked inside Pith

  1. [1]

    Acemoglu, D., Johnson, S., and Robinson, J. A. (2001). The Colonial Origins of Comparative Development: An Empirical Investigation.American Economic Review, 91(5):1369–1401. Armstrong, T. B. and Kolesár, M. (2021). Sensitivity analysis using approximate moment condition models. Quantitative Economics, 12(1):77–108. Card, D. (1995). Using Geographic Variat...

  2. [2]

    Chib, S., Shin, M., and Simoni, A. (2018). Bayesian Estimation and Comparison of Moment Condition Models.Journal of the American Statistical Association, 113(524):1656–1668. Cinelli, C. and Hazlett, C. (2025). An omitted variable bias framework for sensitivity analysis of instrumental variables.Biometrika, 112(2):asaf004. Conley, T. G., Hansen, C. B., and...

  3. [985]

    Guo, Z., Kang, H., Tony Cai, T., and Small, D. S. (2018). Confidence Intervals for Causal Effects with Invalid Instruments by Using Two-Stage Hard Thresholding with Voting.Journal of the Royal Statistical Society Series B: Statistical Methodology, 80(4):793–815. Hieu, N. M., Houssineau, J., Chada, N. K., and Delande, E. (2025). Decoupling epistemic and al...

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.