REVIEW 3 major objections 5 minor 28 references
When Screening Misleads: A Robust Mendelian Randomization Test for Reliable Causal Discovery
T0 review · 3 major / 5 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read Screening exposures for association first can make ordinary Mendelian randomization tests reject far too often when there is no causal effect; a decorrelated test fixes that.
desk verdict Solid selective-inference fix for a real MR workflow problem; math is careful, but the feasible test needs independent external exposure–outcome association data that GWAS practice often only approximates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The feasible decorrelated statistic U2: it subtracts from the MR estimator a plug-in multiple of the main-sample observational association (centered by an independent external association estimate), using the influence-function covariance of the two estimators so the corrected numerator is first-order uncorrelated with the screen.
What would settle it
In a simulation or split-sample study with known zero causal effect and valid instruments, after two-sided association screening, compare conditional rejection rates of the classical Wald/IVW test versus U2: if the classical rate stays near the nominal level while U2 is not needed, or if U2 itself inflates when external association data are independent and correctly specified, the central claim fails.
Extended reading notes
Core claim
Under a correctly specified linear MR model with valid instruments and no causal effect, the classical Wald or fixed-effect IVW test can have severe conditional type I error inflation after two-sided observational association screening, while a feasible decorrelated statistic remains asymptotically standard normal given the screening event.
Load-bearing premise
The method needs an independent estimate of the same population observational exposure-outcome association from a comparable external study; standard SNP-to-exposure and SNP-to-outcome summaries alone are not enough.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies post-screening inference for Mendelian randomization when investigators first screen exposure–outcome pairs by observational association and then apply Wald/IVW MR only to selected pairs. Under a linear structural model with valid instruments, the authors show that the classical MR statistic can be marginally calibrated yet have severe conditional type I error inflation after association screening, because the association estimator and the MR estimator are first-order correlated under confounding. They derive joint asymptotic normality of (β̂_XY, θ̂_n) for the single-SNP Wald ratio and fixed-K IVW (Lemma 1, Theorems 1–2), characterize the selected null distribution via a bivariate Gaussian approximation, and propose a decorrelated statistic that subtracts the component of the MR estimator linearly predictable from β̂_XY. Feasibility requires an independent external estimate of the same population observational association β_XY; the implementable statistic U2 is shown to be asymptotically N(0,1) under the null both unconditionally and conditionally on association-screening events (Theorems 3–5). Simulations illustrate conditional inflation of classical MR and calibration/power of the proposed procedure, with diagnostics on robust vs classical ω11 and alternative ω22 targets.
Significance. If the result holds under the stated conditions, the paper identifies a practically important and under-discussed source of invalid MR inference that arises from the screen-then-MR workflow rather than from weak instruments or pleiotropy. The contribution is technically solid: joint CLT/delta-method expansions, explicit influence functions, post-selection validity theorems with detailed appendices, and a clear geometric account of selection-induced distortion. The robust ω11 analysis is a genuine methodological strength, because it shows that the decorrelation coefficient itself fails if a classical homoskedastic slope variance is used. The work is complementary to existing MR robustness tools and would be useful for exploratory/phenome-wide MR pipelines, provided the external-association requirement can be met or relaxed.
major comments (3)
- Section 4.2 (eqs. 36–40), Remark 1, Theorem 4–5, and Appendix D.2/D.4: the feasible post-screening validity claim for U2 rests on an independent external estimator of the same population β_XY, with independence used to force Cov{A, B−rA+r√ηC}=0. The paper correctly notes that standard SNP–exposure/outcome GWAS summaries alone are insufficient and flags sample overlap as future work (Section 6), but provides no robustness checks under partial overlap or population mismatch (different λ or σ_U²). Because this assumption is load-bearing for the implementable procedure and only approximately true in many GWAS workflows, the manuscript needs either (i) a limited-overlap extension/bound or (ii) simulation evidence quantifying how much overlap/mismatch can be tolerated before conditional type I error breaks.
- Section 2.3 and Figures 2–4: all reported Monte Carlo evidence is under a correctly specified linear model with valid instruments, Rademacher/Binomial genotypes, and Gaussian structural errors (with asymptotic theory allowing non-Gaussianity). The central practical claim is about exploratory association-to-causality workflows. At least one additional simulation block is needed under mild misspecification that is common in applications—e.g., modest horizontal pleiotropy, weak instruments, binary outcomes, or LD among instruments—to show whether U2 remains approximately calibrated or fails gracefully relative to classical MR. Without this, the advantage is demonstrated only in the ideal regime the theory assumes.
- Section 6 and the abstract’s “reliable causal discovery” framing: the method is a post-screening calibration of Wald/IVW under valid instruments, not a general causal-discovery procedure. The discussion already lists weak instruments, pleiotropy, LD, binary traits, and overlap as open. The abstract/introduction should more carefully bound the claim to post-screening type I error control under the stated structural model, and the discussion should give concrete guidance on when an external β_XY estimate is realistically available (e.g., prior observational studies of the same pair) versus when the method is not yet applicable.
minor comments (5)
- Figure 1 is described as an idealized diagnostic whose contours are density ellipses, not uniform regions; the caption is careful, but the main text could more explicitly warn readers not to equate Euclidean area with probability mass.
- Notation for screening level α_scr vs causal level a, and for ω^S_22 vs ω^M_22 vs package-style ω^pkg_22, is dense; a short notation table would help.
- Section 5.2’s package-style variance comparison is useful, but the manuscript should state more clearly that ω22 choice affects power/standardization only and is not the mechanism of post-screening validity.
- References to selective inference (Fithian et al., Taylor & Tibshirani, Lee et al.) are appropriate; a brief comparison to other post-selection IV/MR literature, if any, would help position the contribution.
- Typos/style: occasional spacing issues in math (e.g., “H γ 0”, “bpsel”) and long multi-line displays that could be tightened for readability.
Circularity Check
No significant circularity: post-screening validity is derived from joint CLT/delta-method asymptotics and residualization, then checked in independent Monte Carlo designs under γ=0.
full rationale
The paper’s load-bearing chain is self-contained and not forced by definition or self-citation. Section 3 obtains the joint first-order limit of (β̂_XY, θ̂_n) from sample-covariance CLT plus delta-method expansions under a fixed-K structural IV model (Lemma 1, Theorems 1–2); no normality of structural errors is assumed, and the covariance entries ω11, ω12, ω22 are identified from influence functions rather than fitted to a target rejection rate. Section 4 defines the decorrelation coefficient r=ω12/ω11 so that the residual numerator is first-order uncorrelated with the screening statistic; under the bivariate Gaussian limit this implies asymptotic independence of the corrected statistic from association-screening events A_n (Theorems 3–5, Appendix D.4). That residualization is intentional methodology, not a circular “prediction”: the claim is a theorem under stated assumptions, and the implementable U2 plugs in a consistent influence-function covariance estimator (Proposition 1) plus an independent external association estimate. Type I error and power claims are then checked against nominal levels in Monte Carlo simulations under γ=0 (Figures 2–4), not calibrated to produce those levels. Selective-inference citations (Freedman; Fithian–Sun–Taylor; Taylor–Tibshirani; Lee et al.) are background framing by other authors, not self-citations that define the conclusion. The external-association independence assumption is load-bearing for feasibility and is a genuine robustness concern for GWAS practice, but it is an explicit modeling assumption, not a circular reduction of the result to its inputs. No self-definitional loop, fitted-input-as-prediction, uniqueness-from-authors, or renaming-of-known-result pattern is present.
Assumptions & free parameters
free parameters (3)
- screening level α_scr / threshold c
- external sample size m (or external SE of β̂^(s)_XY)
- IVW weights π_j (or π_j,n)
assumptions (5)
- domain assumption Valid instruments: G independent of (U,ε1,ε2); exclusion restriction holds so Y−γX = λU+ε2 with no direct G→Y path.
- domain assumption Fixed-K asymptotics with pairwise uncorrelated SNPs and relevance (positive first-stage signal).
- standard math Finite fourth moments and i.i.d. sampling so sample covariances admit a multivariate CLT and delta-method joint limit for (β̂_XY, θ̂_n).
- domain assumption External association estimator is independent of the main sample and targets the same population β_XY, with n/m→η∈(0,∞).
- standard math Screening events have positive limiting probability and boundary probability zero in the limiting Gaussian experiment.
invented entities (1)
-
Decorrelated post-screening MR statistic U2 (and oracle/external benchmarks U0, U1)
Cite this review
Pith. "Pith review of When Screening Misleads: A Robust Mendelian Randomization Test for Reliable Causal Discovery." pith.science (2026). https://pith.science/paper/BXSEXBZA
@misc{pith2026260710755,
author = {Pith},
title = {Pith review of: When Screening Misleads: A Robust Mendelian Randomization Test for Reliable Causal Discovery},
year = {2026},
howpublished = {\url{https://pith.science/paper/BXSEXBZA}},
note = {Machine review of arXiv:2607.10755}
}
read the original abstract
Mendelian Randomization (MR) has been widely used as a standard approach for identifying causal effects in biomedical research. When the true exposure and outcome variables are unknown among many potential candidates, a common but problematic practice is to first screen for associations and then evaluate causal effects only among exposure-outcome pairs that show significant correlations. We demonstrate that the classical MR ratio estimator suffers from severe type I error inflation under this selection procedure. To address this issue, we propose a novel robust MR test that remains valid regardless of the prior association screening step when there is no true causal effect. We show that the proposed test consistently maintains the correct type I error rate, independent of the association test results. Furthermore, our method can incorporate summary statistics from previous association studies to improve the power of causal effect detection. Extensive simulation studies illustrate the advantages of the proposed method compared with the classical MR approach.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Mendelian randomization: can genetic epidemi- ology contribute to understanding environmental determinants of disease?International Journal of Epidemiology, 32(1):1–22, 2003
George Davey Smith and Shah Ebrahim. Mendelian randomization: can genetic epidemi- ology contribute to understanding environmental determinants of disease?International Journal of Epidemiology, 32(1):1–22, 2003
2003
-
[2]
Mendelian randomization: genetic anchors for causal inference in epidemiological studies.Human Molecular Genetics, 23(R1):R89–R98, 2014
George Davey Smith and Gibran Hemani. Mendelian randomization: genetic anchors for causal inference in epidemiological studies.Human Molecular Genetics, 23(R1):R89–R98, 2014
2014
-
[3]
Association between C reactive protein and coronary heart disease: mendelian randomisation analysis based on individual participant data.BMJ, 342:d548, 2011
C Reactive Protein Coronary Heart Disease Genetics Collaboration. Association between C reactive protein and coronary heart disease: mendelian randomisation analysis based on individual participant data.BMJ, 342:d548, 2011
2011
-
[4]
Voight, Gina M
Benjamin F. Voight, Gina M. Peloso, Marju Orho-Melander, Ruth Frikke-Schmidt, Maja Barbalic, Majken K. Jensen, George Hindy, Hilma H´ olm, Eric L. Ding, Toby Johnson, et al. Plasma HDL cholesterol and risk of myocardial infarction: a mendelian randomisation study.The Lancet, 380(9841):572–580, 2012
2012
-
[5]
Holmes, Folkert W
Michael V. Holmes, Folkert W. Asselbergs, Tom M. Palmer, Fotios Drenos, Matthew B. Lanktree, Christopher P. Nelson, Caroline E. Dale, Sandosh Padmanabhan, Chris Finan, Daniel I. Swerdlow, et al. Mendelian randomization of blood lipids for coronary heart disease.European Heart Journal, 36(9):539–550, 2015. 40
2015
-
[6]
Thompson
Stephen Burgess, Adam Butterworth, and Simon G. Thompson. Mendelian randomization analysis with multiple genetic variants using summarized data.Genetic Epidemiology, 37(7):658–665, 2013
2013
-
[7]
Wade, Valeriia Haberland, Denis Baird, Charles Laurin, Stephen Burgess, Jack Bowden, Ryan Langdon, Vanessa Y
Gibran Hemani, Jie Zheng, Benjamin Elsworth, Kaitlin H. Wade, Valeriia Haberland, Denis Baird, Charles Laurin, Stephen Burgess, Jack Bowden, Ryan Langdon, Vanessa Y. Tan, James Yarmolinsky, Hashem A. Shihab, Nicholas J. Timpson, David M. Evans, Caroline Relton, Richard M. Martin, George Davey Smith, Tom R. Gaunt, and Philip C. Haycock. The MR-Base platfor...
2018
-
[8]
Timofeeva, Yuan He, Athina Spiliopoulou, Wen-Qing Wei, Angela Gifford, Hao Wu, Tony Varley, Peter K
Xiaomeng Meng, Xia Li, Maria N. Timofeeva, Yuan He, Athina Spiliopoulou, Wen-Qing Wei, Angela Gifford, Hao Wu, Tony Varley, Peter K. Joshi, et al. Phenome-wide mendelian- randomization study of genetically determined vitamin D on multiple health outcomes using the UK Biobank study.International Journal of Epidemiology, 48(5):1425–1434, 2019
2019
Show all 28 references
-
[9]
Robinson, John J
Zhihong Zhu, Zhili Zheng, Futao Zhang, Yang Wu, Maciej Trzaskowski, Robert Maier, Matthew R. Robinson, John J. McGrath, Peter M. Visscher, Naomi R. Wray, and Jian Yang. Causal associations between risk factors and common diseases inferred from GWAS summary data.Nature Communic...
2018
-
[10]
Thompson
Stephen Burgess and Simon G. Thompson. Avoiding bias from weak instruments in mendelian randomization studies.International Journal of Epidemiology, 40(3):755–764, 2011
2011
-
[11]
Mendelian randomization with invalid instruments: effect estimation and bias detection through Egger regression.Inter- national Journal of Epidemiology, 44(2):512–525, 2015
Jack Bowden, George Davey Smith, and Stephen Burgess. Mendelian randomization with invalid instruments: effect estimation and bias detection through Egger regression.Inter- national Journal of Epidemiology, 44(2):512–525, 2015
2015
-
[12]
Freedman
David A. Freedman. A note on screening regression equations.The American Statistician, 37(2):152–155, 1983
1983
-
[13]
Optimal inference after model selec- tion, 2014
William Fithian, Dennis Sun, and Jonathan Taylor. Optimal inference after model selec- tion, 2014
2014
-
[14]
Tibshirani
Jonathan Taylor and Robert J. Tibshirani. Statistical learning and selective infer- ence.Proceedings of the National Academy of Sciences of the United States of America, 112(25):7629–7634, 2015
2015
-
[15]
Lee, Dennis L
Jason D. Lee, Dennis L. Sun, Yuekai Sun, and Jonathan E. Taylor. Exact post-selection inference, with application to the lasso.The Annals of Statistics, 44(3):907–927, 2016
2016
-
[16]
Serfling.Approximation Theorems of Mathematical Statistics
Robert J. Serfling.Approximation Theorems of Mathematical Statistics. John Wiley & Sons, New York, 1980. 41
1980
-
[17]
van der Vaart.Asymptotic Statistics
Aad W. van der Vaart.Asymptotic Statistics. Cambridge University Press, Cambridge, 1998
1998
-
[18]
Newey and Daniel McFadden
Whitney K. Newey and Daniel McFadden. Large sample estimation and hypothesis testing. In Robert F. Engle and Daniel L. McFadden, editors,Handbook of Econometrics, volume 4, pages 2111–2245. Elsevier, Amsterdam, 1994
1994
-
[19]
Frank R. Hampel. The influence curve and its role in robust estimation.Journal of the American Statistical Association, 69(346):383–393, 1974
1974
-
[20]
Peter J. Huber. The behavior of maximum likelihood estimates under nonstandard con- ditions. InProceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, volume 1, pages 221–233. University of California Press, 1967
1967
-
[21]
A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity.Econometrica, 48(4):817–838, 1980
Halbert White. A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity.Econometrica, 48(4):817–838, 1980
1980
-
[22]
Marees, Hilde de Kluiver, Sven Stringer, Florence Vorspan, Emmanuel Curis, Cynthia Marie-Claire, and Eske M
Andries T. Marees, Hilde de Kluiver, Sven Stringer, Florence Vorspan, Emmanuel Curis, Cynthia Marie-Claire, and Eske M. Derks. A tutorial on conducting genome-wide associa- tion studies: Quality control and statistical analysis.International Journal of Methods in Psychiatric R...
2018
-
[23]
Martin, Hilary C
Emil Uffelmann, Qin Qin Huang, Nchangwi Syntia Munung, Jantina de Vries, Yukinori Okada, Alicia R. Martin, Hilary C. Martin, Tuuli Lappalainen, and Danielle Posthuma. Genome-wide association studies.Nature Reviews Methods Primers, 1:59, 2021
2021
-
[24]
E. C. Fieller. Some problems in interval estimation.Journal of the Royal Statistical Society. Series B (Methodological), 16(2):175–185, 1954
1954
-
[25]
Cochran.Sampling Techniques
William G. Cochran.Sampling Techniques. John Wiley & Sons, New York, 3 edition, 1977
1977
-
[26]
Gary W. Oehlert. A note on the delta method.The American Statistician, 46(1):27–29, 1992
1992
-
[27]
TwoSampleMR: Two sample mendelian randomiza- tion functions and interface to MR-Base
MRC Integrative Epidemiology Unit. TwoSampleMR: Two sample mendelian randomiza- tion functions and interface to MR-Base. R package documentation and source code, 2026. Accessed 2026-06-29
2026
-
[28]
TwoSampleMR package documentation, 2026
MRC Integrative Epidemiology Unit.mr wald ratio: Perform 2 Sample IV Using Wald Ratio. TwoSampleMR package documentation, 2026. Accessed 2026-07-02. 42
2026
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.