Pith. sign in

REVIEW 3 major objections 5 minor 28 references

When Screening Misleads: A Robust Mendelian Randomization Test for Reliable Causal Discovery

T0 review · 3 major / 5 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read Screening exposures for association first can make ordinary Mendelian randomization tests reject far too often when there is no causal effect; a decorrelated test fixes that.

desk verdict Solid selective-inference fix for a real MR workflow problem; math is careful, but the feasible test needs independent external exposure–outcome association data that GWAS practice often only approximates. read the letter →

arxiv 2607.10755 v1 pith:BXSEXBZA submitted 2026-07-12 stat.ME

classification stat.ME MSC 62F0362P1062E20
keywords Mendelianrandomizationselectiveinferencepost-screeningconditionaltypeIerrorinstrumentalvariablesdecorrelatedstatisticIVWestimatorWaldratio
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Mendelian randomization is often used only after an exposure and outcome already look associated. The paper shows that this everyday workflow can badly inflate false causal findings even when the instruments are valid and the true causal effect is zero. The association screen and the usual ratio or inverse-variance weighted statistic share random fluctuations under confounding, so conditioning on a significant association tilts the null distribution of the causal test. The authors derive the joint large-sample law of the observational association and the MR estimator, then build a decorrelated statistic that subtracts the component of the MR estimator that is linearly predictable from the association. With an independent external estimate of the same observational association, the corrected test keeps the right type I error after screening and can gain power when that external estimate is precise. The practical message is that screen-then-MR p-values should not be read as ordinary unconditional p-values.

What carries the argument

The feasible decorrelated statistic U2: it subtracts from the MR estimator a plug-in multiple of the main-sample observational association (centered by an independent external association estimate), using the influence-function covariance of the two estimators so the corrected numerator is first-order uncorrelated with the screen.

What would settle it

In a simulation or split-sample study with known zero causal effect and valid instruments, after two-sided association screening, compare conditional rejection rates of the classical Wald/IVW test versus U2: if the classical rate stays near the nominal level while U2 is not needed, or if U2 itself inflates when external association data are independent and correctly specified, the central claim fails.

Watch

Extended reading notes

Core claim

Under a correctly specified linear MR model with valid instruments and no causal effect, the classical Wald or fixed-effect IVW test can have severe conditional type I error inflation after two-sided observational association screening, while a feasible decorrelated statistic remains asymptotically standard normal given the screening event.

Load-bearing premise

The method needs an independent estimate of the same population observational exposure-outcome association from a comparable external study; standard SNP-to-exposure and SNP-to-outcome summaries alone are not enough.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies post-screening inference for Mendelian randomization when investigators first screen exposure–outcome pairs by observational association and then apply Wald/IVW MR only to selected pairs. Under a linear structural model with valid instruments, the authors show that the classical MR statistic can be marginally calibrated yet have severe conditional type I error inflation after association screening, because the association estimator and the MR estimator are first-order correlated under confounding. They derive joint asymptotic normality of (β̂_XY, θ̂_n) for the single-SNP Wald ratio and fixed-K IVW (Lemma 1, Theorems 1–2), characterize the selected null distribution via a bivariate Gaussian approximation, and propose a decorrelated statistic that subtracts the component of the MR estimator linearly predictable from β̂_XY. Feasibility requires an independent external estimate of the same population observational association β_XY; the implementable statistic U2 is shown to be asymptotically N(0,1) under the null both unconditionally and conditionally on association-screening events (Theorems 3–5). Simulations illustrate conditional inflation of classical MR and calibration/power of the proposed procedure, with diagnostics on robust vs classical ω11 and alternative ω22 targets.

Significance. If the result holds under the stated conditions, the paper identifies a practically important and under-discussed source of invalid MR inference that arises from the screen-then-MR workflow rather than from weak instruments or pleiotropy. The contribution is technically solid: joint CLT/delta-method expansions, explicit influence functions, post-selection validity theorems with detailed appendices, and a clear geometric account of selection-induced distortion. The robust ω11 analysis is a genuine methodological strength, because it shows that the decorrelation coefficient itself fails if a classical homoskedastic slope variance is used. The work is complementary to existing MR robustness tools and would be useful for exploratory/phenome-wide MR pipelines, provided the external-association requirement can be met or relaxed.

major comments (3)
  1. Section 4.2 (eqs. 36–40), Remark 1, Theorem 4–5, and Appendix D.2/D.4: the feasible post-screening validity claim for U2 rests on an independent external estimator of the same population β_XY, with independence used to force Cov{A, B−rA+r√ηC}=0. The paper correctly notes that standard SNP–exposure/outcome GWAS summaries alone are insufficient and flags sample overlap as future work (Section 6), but provides no robustness checks under partial overlap or population mismatch (different λ or σ_U²). Because this assumption is load-bearing for the implementable procedure and only approximately true in many GWAS workflows, the manuscript needs either (i) a limited-overlap extension/bound or (ii) simulation evidence quantifying how much overlap/mismatch can be tolerated before conditional type I error breaks.
  2. Section 2.3 and Figures 2–4: all reported Monte Carlo evidence is under a correctly specified linear model with valid instruments, Rademacher/Binomial genotypes, and Gaussian structural errors (with asymptotic theory allowing non-Gaussianity). The central practical claim is about exploratory association-to-causality workflows. At least one additional simulation block is needed under mild misspecification that is common in applications—e.g., modest horizontal pleiotropy, weak instruments, binary outcomes, or LD among instruments—to show whether U2 remains approximately calibrated or fails gracefully relative to classical MR. Without this, the advantage is demonstrated only in the ideal regime the theory assumes.
  3. Section 6 and the abstract’s “reliable causal discovery” framing: the method is a post-screening calibration of Wald/IVW under valid instruments, not a general causal-discovery procedure. The discussion already lists weak instruments, pleiotropy, LD, binary traits, and overlap as open. The abstract/introduction should more carefully bound the claim to post-screening type I error control under the stated structural model, and the discussion should give concrete guidance on when an external β_XY estimate is realistically available (e.g., prior observational studies of the same pair) versus when the method is not yet applicable.
minor comments (5)
  1. Figure 1 is described as an idealized diagnostic whose contours are density ellipses, not uniform regions; the caption is careful, but the main text could more explicitly warn readers not to equate Euclidean area with probability mass.
  2. Notation for screening level α_scr vs causal level a, and for ω^S_22 vs ω^M_22 vs package-style ω^pkg_22, is dense; a short notation table would help.
  3. Section 5.2’s package-style variance comparison is useful, but the manuscript should state more clearly that ω22 choice affects power/standardization only and is not the mechanism of post-screening validity.
  4. References to selective inference (Fithian et al., Taylor & Tibshirani, Lee et al.) are appropriate; a brief comparison to other post-selection IV/MR literature, if any, would help position the contribution.
  5. Typos/style: occasional spacing issues in math (e.g., “H γ 0”, “bpsel”) and long multi-line displays that could be tightened for readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: post-screening validity is derived from joint CLT/delta-method asymptotics and residualization, then checked in independent Monte Carlo designs under γ=0.

full rationale

The paper’s load-bearing chain is self-contained and not forced by definition or self-citation. Section 3 obtains the joint first-order limit of (β̂_XY, θ̂_n) from sample-covariance CLT plus delta-method expansions under a fixed-K structural IV model (Lemma 1, Theorems 1–2); no normality of structural errors is assumed, and the covariance entries ω11, ω12, ω22 are identified from influence functions rather than fitted to a target rejection rate. Section 4 defines the decorrelation coefficient r=ω12/ω11 so that the residual numerator is first-order uncorrelated with the screening statistic; under the bivariate Gaussian limit this implies asymptotic independence of the corrected statistic from association-screening events A_n (Theorems 3–5, Appendix D.4). That residualization is intentional methodology, not a circular “prediction”: the claim is a theorem under stated assumptions, and the implementable U2 plugs in a consistent influence-function covariance estimator (Proposition 1) plus an independent external association estimate. Type I error and power claims are then checked against nominal levels in Monte Carlo simulations under γ=0 (Figures 2–4), not calibrated to produce those levels. Selective-inference citations (Freedman; Fithian–Sun–Taylor; Taylor–Tibshirani; Lee et al.) are background framing by other authors, not self-citations that define the conclusion. The external-association independence assumption is load-bearing for feasibility and is a genuine robustness concern for GWAS practice, but it is an explicit modeling assumption, not a circular reduction of the result to its inputs. No self-definitional loop, fitted-input-as-prediction, uniqueness-from-authors, or renaming-of-known-result pattern is present.

Assumptions & free parameters 3 free parameters · 5 assumptions · 1 invented entities

The central claim rests on a standard linear IV structural model with valid instruments, fixed number of SNPs, uncorrelated variants, finite fourth moments, and a first-order joint Gaussian limit for (β̂_XY, θ̂_n). The feasible test further requires an independent external estimate of the observational association. No new physical entities are postulated; free choices are design/tuning quantities (screening level, external sample size, IVW weights) rather than fitted scientific constants.

free parameters (3)
  • screening level α_scr / threshold c
    Chosen by the analyst; defines the selection event A_n on which conditional type I error is evaluated (eq. 1).
  • external sample size m (or external SE of β̂^(s)_XY)
    Determines V1 and local power of U1/U2 (eqs. 38, 46); treated as known input, not estimated from the main sample.
  • IVW weights π_j (or π_j,n)
    Enter the multi-SNP estimator and influence function; assumed to converge to fixed positive limits (eq. 19).
assumptions (5)
  • domain assumption Valid instruments: G independent of (U,ε1,ε2); exclusion restriction holds so Y−γX = λU+ε2 with no direct G→Y path.
    Assumption 1(A2) and structural model (14); paper states theory does not address pleiotropy.
  • domain assumption Fixed-K asymptotics with pairwise uncorrelated SNPs and relevance (positive first-stage signal).
    Assumption 1(A1,A3,A4); LD would require matrix weights not developed here.
  • standard math Finite fourth moments and i.i.d. sampling so sample covariances admit a multivariate CLT and delta-method joint limit for (β̂_XY, θ̂_n).
    Assumption 1(A2) and Section 3 proof strategy; normality of structural errors is not required.
  • domain assumption External association estimator is independent of the main sample and targets the same population β_XY, with n/m→η∈(0,∞).
    Section 4.2 eqs. (36)–(38); required for feasible U1/U2 and post-selection independence.
  • standard math Screening events have positive limiting probability and boundary probability zero in the limiting Gaussian experiment.
    Theorem 3 and Appendix D.4; used to pass from joint convergence to conditional null N(0,1).
invented entities (1)
  • Decorrelated post-screening MR statistic U2 (and oracle/external benchmarks U0, U1)
    purpose: Remove first-order covariance between the MR estimator and the association-screening statistic so conditional type I error remains nominal.
    Defined in Section 4 as a statistical procedure, not a new biological object; independent evidence is the asymptotic theory and simulations under the stated model.

how reviews work

0 comments
Cite this review

Pith. "Pith review of When Screening Misleads: A Robust Mendelian Randomization Test for Reliable Causal Discovery." pith.science (2026). https://pith.science/paper/BXSEXBZA

@misc{pith2026260710755,
  author       = {Pith},
  title        = {Pith review of: When Screening Misleads: A Robust Mendelian Randomization Test for Reliable Causal Discovery},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BXSEXBZA}},
  note         = {Machine review of arXiv:2607.10755}
}
read the original abstract

Mendelian Randomization (MR) has been widely used as a standard approach for identifying causal effects in biomedical research. When the true exposure and outcome variables are unknown among many potential candidates, a common but problematic practice is to first screen for associations and then evaluate causal effects only among exposure-outcome pairs that show significant correlations. We demonstrate that the classical MR ratio estimator suffers from severe type I error inflation under this selection procedure. To address this issue, we propose a novel robust MR test that remains valid regardless of the prior association screening step when there is no true causal effect. We show that the proposed test consistently maintains the correct type I error rate, independent of the association test results. Furthermore, our method can incorporate summary statistics from previous association studies to improve the power of causal effect detection. Extensive simulation studies illustrate the advantages of the proposed method compared with the classical MR approach.

Figures

Figures reproduced from arXiv: 2607.10755 by the authors.

Figure 1
Figure 1. Geometric view of post-screening MR miscalibration under the bivariate Gaussian [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Conditional type I error versus empirical selection rate under two-sided observational [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Unconditional type I error versus empirical selection rate under the same settings as [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Conditional power as a function of γ after two-sided observational association screen￾ing. The curves compare the package-style statistic, the oracle benchmark U0, the external￾association benchmark U1, and the feasible decorrelated statistic U2 for different external …
Figure 5
Figure 5. Figure 5: Diagnostics for the ω11 variance target. The robust target is the influence-function variance in (47); the classical target is (48). Large robust–classical discrepancies occur in parameter regimes that include low-frequency instruments and concentrated first-stage sign…
Figure 6
Figure 6. Figure 6: Effect of the ω22 variance target on conditional power. The full first-order target and the package-style target agree at the null in these structural-model diagnostics, but differ away from the null. The sign of the difference reverses for negative versus positive cau…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

28 extracted references

  1. [1]

    Mendelian randomization: can genetic epidemi- ology contribute to understanding environmental determinants of disease?International Journal of Epidemiology, 32(1):1–22, 2003

    George Davey Smith and Shah Ebrahim. Mendelian randomization: can genetic epidemi- ology contribute to understanding environmental determinants of disease?International Journal of Epidemiology, 32(1):1–22, 2003

  2. [2]

    Mendelian randomization: genetic anchors for causal inference in epidemiological studies.Human Molecular Genetics, 23(R1):R89–R98, 2014

    George Davey Smith and Gibran Hemani. Mendelian randomization: genetic anchors for causal inference in epidemiological studies.Human Molecular Genetics, 23(R1):R89–R98, 2014

  3. [3]

    Association between C reactive protein and coronary heart disease: mendelian randomisation analysis based on individual participant data.BMJ, 342:d548, 2011

    C Reactive Protein Coronary Heart Disease Genetics Collaboration. Association between C reactive protein and coronary heart disease: mendelian randomisation analysis based on individual participant data.BMJ, 342:d548, 2011

  4. [4]

    Voight, Gina M

    Benjamin F. Voight, Gina M. Peloso, Marju Orho-Melander, Ruth Frikke-Schmidt, Maja Barbalic, Majken K. Jensen, George Hindy, Hilma H´ olm, Eric L. Ding, Toby Johnson, et al. Plasma HDL cholesterol and risk of myocardial infarction: a mendelian randomisation study.The Lancet, 380(9841):572–580, 2012

  5. [5]

    Holmes, Folkert W

    Michael V. Holmes, Folkert W. Asselbergs, Tom M. Palmer, Fotios Drenos, Matthew B. Lanktree, Christopher P. Nelson, Caroline E. Dale, Sandosh Padmanabhan, Chris Finan, Daniel I. Swerdlow, et al. Mendelian randomization of blood lipids for coronary heart disease.European Heart Journal, 36(9):539–550, 2015. 40

  6. [6]

    Thompson

    Stephen Burgess, Adam Butterworth, and Simon G. Thompson. Mendelian randomization analysis with multiple genetic variants using summarized data.Genetic Epidemiology, 37(7):658–665, 2013

  7. [7]

    Wade, Valeriia Haberland, Denis Baird, Charles Laurin, Stephen Burgess, Jack Bowden, Ryan Langdon, Vanessa Y

    Gibran Hemani, Jie Zheng, Benjamin Elsworth, Kaitlin H. Wade, Valeriia Haberland, Denis Baird, Charles Laurin, Stephen Burgess, Jack Bowden, Ryan Langdon, Vanessa Y. Tan, James Yarmolinsky, Hashem A. Shihab, Nicholas J. Timpson, David M. Evans, Caroline Relton, Richard M. Martin, George Davey Smith, Tom R. Gaunt, and Philip C. Haycock. The MR-Base platfor...

  8. [8]

    Timofeeva, Yuan He, Athina Spiliopoulou, Wen-Qing Wei, Angela Gifford, Hao Wu, Tony Varley, Peter K

    Xiaomeng Meng, Xia Li, Maria N. Timofeeva, Yuan He, Athina Spiliopoulou, Wen-Qing Wei, Angela Gifford, Hao Wu, Tony Varley, Peter K. Joshi, et al. Phenome-wide mendelian- randomization study of genetically determined vitamin D on multiple health outcomes using the UK Biobank study.International Journal of Epidemiology, 48(5):1425–1434, 2019

Show all 28 references
  1. [9]

    Robinson, John J

    Zhihong Zhu, Zhili Zheng, Futao Zhang, Yang Wu, Maciej Trzaskowski, Robert Maier, Matthew R. Robinson, John J. McGrath, Peter M. Visscher, Naomi R. Wray, and Jian Yang. Causal associations between risk factors and common diseases inferred from GWAS summary data.Nature Communic...

  2. [10]

    Thompson

    Stephen Burgess and Simon G. Thompson. Avoiding bias from weak instruments in mendelian randomization studies.International Journal of Epidemiology, 40(3):755–764, 2011

  3. [11]

    Mendelian randomization with invalid instruments: effect estimation and bias detection through Egger regression.Inter- national Journal of Epidemiology, 44(2):512–525, 2015

    Jack Bowden, George Davey Smith, and Stephen Burgess. Mendelian randomization with invalid instruments: effect estimation and bias detection through Egger regression.Inter- national Journal of Epidemiology, 44(2):512–525, 2015

  4. [12]

    Freedman

    David A. Freedman. A note on screening regression equations.The American Statistician, 37(2):152–155, 1983

  5. [13]

    Optimal inference after model selec- tion, 2014

    William Fithian, Dennis Sun, and Jonathan Taylor. Optimal inference after model selec- tion, 2014

  6. [14]

    Tibshirani

    Jonathan Taylor and Robert J. Tibshirani. Statistical learning and selective infer- ence.Proceedings of the National Academy of Sciences of the United States of America, 112(25):7629–7634, 2015

  7. [15]

    Lee, Dennis L

    Jason D. Lee, Dennis L. Sun, Yuekai Sun, and Jonathan E. Taylor. Exact post-selection inference, with application to the lasso.The Annals of Statistics, 44(3):907–927, 2016

  8. [16]

    Serfling.Approximation Theorems of Mathematical Statistics

    Robert J. Serfling.Approximation Theorems of Mathematical Statistics. John Wiley & Sons, New York, 1980. 41

  9. [17]

    van der Vaart.Asymptotic Statistics

    Aad W. van der Vaart.Asymptotic Statistics. Cambridge University Press, Cambridge, 1998

  10. [18]

    Newey and Daniel McFadden

    Whitney K. Newey and Daniel McFadden. Large sample estimation and hypothesis testing. In Robert F. Engle and Daniel L. McFadden, editors,Handbook of Econometrics, volume 4, pages 2111–2245. Elsevier, Amsterdam, 1994

  11. [19]

    Frank R. Hampel. The influence curve and its role in robust estimation.Journal of the American Statistical Association, 69(346):383–393, 1974

  12. [20]

    Peter J. Huber. The behavior of maximum likelihood estimates under nonstandard con- ditions. InProceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, volume 1, pages 221–233. University of California Press, 1967

  13. [21]

    A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity.Econometrica, 48(4):817–838, 1980

    Halbert White. A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity.Econometrica, 48(4):817–838, 1980

  14. [22]

    Marees, Hilde de Kluiver, Sven Stringer, Florence Vorspan, Emmanuel Curis, Cynthia Marie-Claire, and Eske M

    Andries T. Marees, Hilde de Kluiver, Sven Stringer, Florence Vorspan, Emmanuel Curis, Cynthia Marie-Claire, and Eske M. Derks. A tutorial on conducting genome-wide associa- tion studies: Quality control and statistical analysis.International Journal of Methods in Psychiatric R...

  15. [23]

    Martin, Hilary C

    Emil Uffelmann, Qin Qin Huang, Nchangwi Syntia Munung, Jantina de Vries, Yukinori Okada, Alicia R. Martin, Hilary C. Martin, Tuuli Lappalainen, and Danielle Posthuma. Genome-wide association studies.Nature Reviews Methods Primers, 1:59, 2021

  16. [24]

    E. C. Fieller. Some problems in interval estimation.Journal of the Royal Statistical Society. Series B (Methodological), 16(2):175–185, 1954

  17. [25]

    Cochran.Sampling Techniques

    William G. Cochran.Sampling Techniques. John Wiley & Sons, New York, 3 edition, 1977

  18. [26]

    Gary W. Oehlert. A note on the delta method.The American Statistician, 46(1):27–29, 1992

  19. [27]

    TwoSampleMR: Two sample mendelian randomiza- tion functions and interface to MR-Base

    MRC Integrative Epidemiology Unit. TwoSampleMR: Two sample mendelian randomiza- tion functions and interface to MR-Base. R package documentation and source code, 2026. Accessed 2026-06-29

  20. [28]

    TwoSampleMR package documentation, 2026

    MRC Integrative Epidemiology Unit.mr wald ratio: Perform 2 Sample IV Using Wald Ratio. TwoSampleMR package documentation, 2026. Accessed 2026-07-02. 42

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.