REVIEW 3 major objections 5 minor 39 references
Context-stratified Mendelian randomization: exploiting regional exposure variation to explore causal effect heterogeneity and non-linearity
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Context-stratified Mendelian randomization treats recruitment centre, region, or time period as an exogenous stratifier, so between-context differences in exposure become a test for effect heterogeneity and non-linearity.
desk verdict A simple, honest methods paper that formalizes context-stratified MR; the central interpretative claim rests on an exchangeability assumption the paper flags but does not probe. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the exogenous context variable, a pre-specified categorical variable (for instance, recruitment centre, region, or time period) that is not a function of the exposure, outcome, or instrument, and that induces differences in the average exposure level between subgroups. The method partitions the sample into $K$ contexts, computes a context-specific IV estimate $\hat\beta_k$ (e.g., by the ratio method), and then combines two summary statistics: Cochran's $Q$ statistic with either first-order or modified second-order weights to test whether the $\hat\beta_k$ vary beyond chance, and the slope from a meta-regression of $\hat\beta_k$ on the context-specific mean exposure $\bar{x}_k$ to test for a dose-response trend. Exogeneity of the context is what prevents collider bias and makes the context-specific estimates locally valid; a significant $Q$ or trend is then a signal of effect heterogeneity or non-linearity, provided other outcome-relevant factors are comparable across contexts.
What would settle it
Simulate the linear homogeneous scenario f(x) = 0.8x with a weak instrument (F ≈ 10) and use modified second-order Cochran's Q; if the rejection rate substantially exceeds 5%, the claimed false-positive control under homogeneity would be refuted.
Extended reading notes
Core claim
The central claim is that a Mendelian randomization estimate can be made context-specific by stratifying on an exogenous variable such as recruitment centre, geographic region, or time period, and that differences across these context-specific estimates provide evidence for effect heterogeneity or non-linearity. Each context-specific estimate is a valid local causal effect under the standard instrumental variable assumptions, without the constant-effect or rank-preserving assumptions needed by residual-based and doubly-ranked methods. The paper demonstrates in simulations that the approach detects quadratic and threshold effects with good power when between-context differences in exposure are large, and that the modified second-order Cochran's Q keeps false-positive rates near nominal in the linear homogeneous scenario. In the applied vitamin D example, no heterogeneity is found (Q p=0.28, trend p=0.76), and the paper notes that the narrow range of context-specific mean exposure limits the method's power and interpretability.
Load-bearing premise
The load-bearing premise is that the context variable is exogenous and that contexts are comparable in all outcome-relevant factors besides the exposure distribution; if a context also differs in a confounder, the between-context estimate differences cannot be attributed to the exposure.
Editorial extensions
If this is right
- If the central claim is right, any Mendelian randomization analysis with a natural context that has meaningful exposure variation can report a set of locally valid effects rather than a single population-averaged estimate, and a significant Q statistic or meta-regression trend reveals that the average hid heterogeneity or non-linearity.
- The method offers a checkable alternative to residual-based and doubly-ranked stratification: context-specific estimates are valid without constant-effect or rank-preserving assumptions, at the price of needing genuine between-context exposure variation.
- In the vitamin D example, the null findings are consistent with no causal effect, but the narrow 50-58 nmol/L range across centres means the analysis has low power to detect non-linearity; a context with wider exposure variation could change conclusions.
- Using modified second-order weights for Cochran's Q controls false-positive rates under homogeneity, whereas first-order weights over-reject with strong instruments and should not be used for null testing.
Reading between the lines
- A natural extension the paper leaves implicit: when centre-level exposure differences are small, stratifying by a temporal context such as season or year of recruitment could widen the exposure range and increase power.
- A useful falsification exercise would be to apply context-stratified MR to a negative-control outcome; if a trend appears there, it would indicate context-level confounding rather than genuine effect modification.
- If the method's power depends critically on between-context exposure variation, pooling multiple cohorts or countries into a single context-stratified analysis could make the trend test competitive with residual-based stratification.
- A significant Q in a context-stratified analysis could also arise from pleiotropy whose magnitude varies by context; comparing context-specific estimates with those from assumptions-free sensitivity analyses would help distinguish effect modification from assumption violation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes context-stratified Mendelian randomization (CS-MR), in which the study population is partitioned by an exogenous context variable (e.g., recruitment centre), separate instrumental-variable analyses are performed within each context, and the resulting context-specific estimates are examined for heterogeneity via Cochran's Q and for trend via meta-regression on the context-specific mean exposure. The method is illustrated with simulations covering linear, quadratic, and threshold exposure–response functions under two between-context exposure-difference scenarios, and with a UK Biobank application estimating the effect of 25-hydroxyvitamin D on coronary artery disease risk across 20 recruitment centres, where no causal effect or heterogeneity is found. The paper argues that CS-MR avoids the strong constant-effect or rank-preserving assumptions required by residual-based and doubly-ranked stratification methods, at the cost of requiring meaningful exogenous between-context exposure variation.
Significance. Strengths of the manuscript include a clearly described simulation design, fully provided R code, and a candid discussion of limitations, particularly the narrow exposure range in the applied example and the potential for context-level confounders. If the method works as claimed, it gives applied researchers a simple, transparent tool for exploratory investigation of effect heterogeneity and non-linearity. However, the simulation evidence only partially supports the abstract's claim of nominal false-positive control, and the absence of any sensitivity analysis for the exchangeability assumption leaves the central identification argument incomplete.
major comments (3)
- [Section 3.2 / Table 1 / Abstract] The abstract claims that "the approach detects heterogeneity when present while maintaining nominal false positive rates under homogeneity when appropriate methods are used." Under a linear homogeneous effect, Table 1 shows the first-order Q test rejects at 12.8% (larger differences) and 10.4% (smaller differences), while the modified second-order Q test rejects at 0.4% and 0.1%. Neither version maintains the nominal 5% level, and only the trend test (3.3% and 4.1%) is close to nominal. The phrase "appropriate methods" is undefined and the simulation does not identify a heterogeneity test with acceptable false-positive control; the abstract should be revised to describe the actual trade-off between the over-rejecting first-order Q and the under-rejecting modified second-order Q.
- [Section 2.2 and Section 5.3] The method's central claim that between-context differences in estimates indicate effect heterogeneity or non-linearity relies on contexts being comparable in all outcome-relevant factors other than the exposure distribution. The paper acknowledges this requirement but provides no sensitivity analyses, negative controls, or adjustment for context-level covariates. The simulation DGP in Section 3.1 contains no context-level variable other than α_k, so it cannot reveal whether a context-level effect modifier correlated with α_k would produce a spurious trend. In the UK Biobank example, centres differ in latitude, deprivation, and lifestyle; the null result cannot validate exchangeability, and a positive result would be ambiguous between effect modification by context and non-linearity in the exposure–response. Without such sensitivity analyses, the abstract's claim that the method can "investigate effect heterogeneity and non-linearity" is under-supported.
- [Supplementary Material A.2] The provided R code assigns `alpha` twice, first with the larger-difference sequence (from 8 by 0.2) and then immediately with the smaller-difference sequence (from 9 by 0.1). As printed, the simulation only runs the smaller-difference scenario, so the larger-difference rows of Table 1 cannot be reproduced from the code without manual editing. The code should be corrected to run both scenarios or clearly comment which line should be uncommented.
minor comments (5)
- [Section 3.2] The sentence "with elevated coverage rates for the heterogeneity test using first-order weights" should read "rejection rates" because the context is about false-positive proportions.
- [Figure 1] The left panel's y-axis label reads "Log odds ratio for coronary heart disease" but the outcome is coronary artery disease; the label should be consistent with the text and Table 2.
- [Section 4] The phrase "the outcome was also defined in the same way as in this paper" should refer to the previous publication [23] rather than "this paper".
- [Table 2] The centre name "Middlesborough" should be "Middlesbrough" to match standard spelling.
- [Section 2.3, Step 5] The meta-regression uses the observed context-specific mean exposure as a regressor, but this mean is estimated with error; the paper does not discuss the impact of this measurement error on the trend test's calibration, although the simulation results suggest the effect is modest.
Circularity Check
No significant circularity; the method's estimates come from independent subgroup IV analyses, and the simulation and applied example do not fit the target result into the inputs.
full rationale
The paper's derivation chain is self-contained in the relevant sense. Context-specific estimates are obtained by separate Mendelian randomization analyses within each subgroup (Section 2.3, Step 3), and heterogeneity is assessed with Cochran's Q and meta-regression of these independent estimates on subgroup mean exposure (Steps 4-5). Nothing in the method defines the target heterogeneity or trend in terms of the fitted quantities: the subgroup mean exposure is an observed summary used as a covariate, not a parameter fitted to the outcome. The simulation study sets the data-generating parameters (alpha_k, instrument effect 0.5, and the three true effect functions) by the authors, and then evaluates rejection rates; this is a validation exercise, not a prediction derived from fitted inputs. The applied example uses a genetic instrument previously published by the authors (reference [23]), but that is background evidence for the instrument, not a definitional source of the paper's claim; the context-specific estimates and their heterogeneity statistics are computed fresh from UK Biobank data. The main assumption that contexts be comparable in outcome-relevant factors other than exposure distribution is explicitly stated and caveated in Sections 2.2, 5.2, and 5.3, including the admission that 'differences between estimates may not be attributable to the exposure.' That is a limitation and correctness risk, not circularity. No equation in the paper reduces to a fit renamed as a prediction, no uniqueness theorem is imported from the authors, and no ansatz is smuggled in via citation. The comparison with residual-based and doubly-ranked methods uses previously published subgroup estimates only as illustration, not as the basis for the proposed method's validity.
Assumptions & free parameters
assumptions (4)
- domain assumption The standard IV assumptions (relevance, independence, exclusion restriction) hold within every context.
- domain assumption The context variable is exogenous, meaning it is not a function of any variable in the model.
- domain assumption For attribution of differences to the exposure, contexts must be comparable in other outcome-related factors.
- domain assumption The trend test assumes a linear relationship between context-specific estimates and context mean exposure.
Cite this review
Pith. "Pith review of Context-stratified Mendelian randomization: exploiting regional exposure variation to explore causal effect heterogeneity and non-linearity." pith.science (2026). https://pith.science/paper/6H3C7KWE
@misc{pith2026250711088,
author = {Pith},
title = {Pith review of: Context-stratified Mendelian randomization: exploiting regional exposure variation to explore causal effect heterogeneity and non-linearity},
year = {2026},
howpublished = {\url{https://pith.science/paper/6H3C7KWE}},
note = {Machine review of arXiv:2507.11088}
}
read the original abstract
Mendelian randomization (MR) uses genetic variants as instrumental variables to make causal claims. Standard MR approaches typically report a single population-averaged estimate, limiting their ability to explore effect heterogeneity or non-linear dose-response relationships. Existing stratification methods, such as residual-based and doubly-ranked stratified MR, attempt to overcome this but rely on strong and unverifiable assumptions. We propose an alternative, context-stratified Mendelian randomization, which exploits exogenous variation in the exposure across subgroups -- such as recruitment centres, geographic regions, or time periods -- to investigate effect heterogeneity and non-linearity. Separate MR analyses are performed within each context, and heterogeneity in the resulting estimates is assessed using Cochran's Q statistic and meta-regression. We demonstrate through simulations that the approach detects heterogeneity when present while maintaining nominal false positive rates under homogeneity when appropriate methods are used. In an applied example using UK Biobank data, we assess the effect of vitamin D levels on coronary artery disease risk across 20 recruitment centres. Despite some regional variation in vitamin D distributions, there is no evidence for a causal effect or heterogeneity in estimates. Compared to stratification methods requiring model-based assumptions, the context-stratified approach is simple to implement and robust to collider bias, provided the context variable is exogenous. However, the method's power and interpretability depend critically on meaningful exogenous variation in exposure distributions between contexts. In the example of vitamin D, subgroups from other stratification methods explored a much wider range of the exposure distribution.
Figures
Reference graph
Works this paper leans on
-
[1]
Mendelia n randomization: using genes as instruments for making causal inferences in epidemio logy
Lawlor D, Harbord R, Sterne J, Timpson N, Davey Smith G. Mendelia n randomization: using genes as instruments for making causal inferences in epidemio logy. Statistics in Medicine 2008; 27(8):1133–1163, doi:10.1002/sim.3034
-
[2]
Mendelian randomization: methods for causal inference usi ng genetic variants
Burgess S, Thompson SG. Mendelian randomization: methods for causal inference usi ng genetic variants . Chapman & Hall, Boca Raton, FL, 2021
work page 2021
-
[3]
Mendelian randomization as an instrumental variable approach to causal inference
Didelez V, Sheehan N. Mendelian randomization as an instrumental variable approach to causal inference. Statistical Methods in Medical Research 2007; 16(4):309–330, doi: 10.1177/0962280206077743
-
[4]
Identification of causal effects us ing instrumental variables
Angrist J, Imbens G, Rubin D. Identification of causal effects us ing instrumental variables. Journal of the American Statistical Association 1996; 91(434):444–455, doi: 10.2307/2291629
doi:10.2307/2291629 1996
-
[5]
Identification and estimation of local ave rage treatment effects
Imbens GW, Angrist JD. Identification and estimation of local ave rage treatment effects. Econometrica 1994; 62(2):467–475, doi:10.2307/2951620
doi:10.2307/2951620 1994
-
[6]
Tian H, Tom BD, Burgess S. A data-adaptive method for investiga ting effect heterogeneity with high-dimensional covariates in Mendelian randomization. BMC Medical Research Methodology 2024; 24(1):34
work page 2024
-
[7]
Identification of effect modifiers using a stratified 13 Mendelian randomization algorithmic framework
Man A, Kn¨ usel L, Graf J, Lali R, Le A, Di Scipio M, Mohammadi-She mirani P, Chong M, Pigeyre M, Kutalik Z, et al. . Identification of effect modifiers using a stratified 13 Mendelian randomization algorithmic framework. European Journal of Epidemiology 2025; 40(3):275–296
work page 2025
-
[8]
Instrumental variable analysis with a nonlinear exposure–outcome relationship
Burgess S, Davies NM, Thompson SG, EPIC-InterAct Consortiu m. Instrumental variable analysis with a nonlinear exposure–outcome relationship. Epidemiology 2014; 25(6):877– 885, doi:10.1097/ede.0000000000000161
Show all 39 references
-
[9]
Statistics in medicine – reporting of subgroup analyses in clinical trials
Wang R, Lagakos SW, Ware JH, Hunter DJ, Drazen JM. Statistics in medicine – reporting of subgroup analyses in clinical trials. New England Journal of Medicine 2007; 357(21):2189–2194
2007
-
[10]
Illustrating bias due to conditioning on a collider
Cole SR, Platt R W, Schisterman EF, Chu H, Westreich D, Richards on D, Poole C. Illustrating bias due to conditioning on a collider. International Journal of Epidemiology 2010; 39(2):417–420, doi:10.1093/ije/dyp334
2010 doi
-
[11]
Analysis and interp retation of treatment effects in subgroups of patients in randomized clinical trials
Yusuf S, Wittes J, Probstfield J, Tyroler HA. Analysis and interp retation of treatment effects in subgroups of patients in randomized clinical trials. JAMA 1991; 266(1):93–98, doi:10.1001/jama.1991.03470010097038
1991
-
[12]
C-reactive protein levels and risk of dementia
Burgess S. “C-reactive protein levels and risk of dementia”: Su bgroup analyses in Mendelian randomization are likely to be misleading. Alzheimer’s & Dementia 2022; 18(12):2732–2733
2022
-
[13]
Semiparametric methods for estimation o f a nonlinear exposure- outcome relationship using instrumental variables with application to Mendelian randomization
Staley JR, Burgess S. Semiparametric methods for estimation o f a nonlinear exposure- outcome relationship using instrumental variables with application to Mendelian randomization. Genetic Epidemiology 2017; 41(4):341–352, doi:10.1002/gepi.22041
2017 doi
-
[14]
Relaxing parametric assumpt ions for non-linear Mendelian randomization using a doubly-ranked stratification metho d
Tian H, Mason AM, Liu C, Burgess S. Relaxing parametric assumpt ions for non-linear Mendelian randomization using a doubly-ranked stratification metho d. PLOS Genetics 2023; 19(6):e1010 823
2023
-
[15]
Violation of the constant genetic effect assumption can result in biased estimates for non-linear Mendelian randomization
Burgess S. Violation of the constant genetic effect assumption can result in biased estimates for non-linear Mendelian randomization. Human Heredity 2023; 88(1):79–90
2023
-
[16]
Non-linear Mendelian randomization: evaluation of effe ct modification in the residual and doubly-ranked methods with simulated and empirica l examples
Hamilton FW, Hughes DA, Lu T, Kutalik Z, Gkatzionis A, Tilling K, Hart wig FP, Davey Smith G. Non-linear Mendelian randomization: evaluation of effe ct modification in the residual and doubly-ranked methods with simulated and empirica l examples. European journal of Epidemiology 2025
2025
-
[17]
Modeling site effec ts in the design and analysis of multi-site trials
Feaster DJ, Mikulich-Gilbertson S, Brincks AM. Modeling site effec ts in the design and analysis of multi-site trials. The American Journal of Drug and Alcohol Abuse 2011; 37(5):383–391
2011
-
[18]
How should meta-regression analyses be undertaken and interpreted? Statistics in medicine 2002; 21(11):1559–1573
Thompson SG, Higgins JP. How should meta-regression analyses be undertaken and interpreted? Statistics in medicine 2002; 21(11):1559–1573
2002
-
[19]
Combining information on multiple instrumental variables in Mendelian randomization: comparison of allele score and su mmarized data methods
Burgess S, Dudbridge F, Thompson SG. Combining information on multiple instrumental variables in Mendelian randomization: comparison of allele score and su mmarized data methods. Statistics in Medicine 2016; 35(11):1880–1906, doi:10.1002/sim.6835
2016 doi
-
[20]
Bowden J, Hemani G, Davey Smith G. Detecting individual and glob al horizontal pleiotropy in Mendelian randomization – a job for the humble heteroge neity statistic? American Journal of Epidemiology 2018; 187(12):2681–2685, doi:10.1093/aje/kwy185. 14
2018 doi
-
[21]
Improving the accuracy of two-sample summary data Mendelian ran domization: moving beyond the NOME assumption
Bowden J, Del Greco F, Minelli C, Lawlor D, Sheehan N, Thompson J, Davey Smith G. Improving the accuracy of two-sample summary data Mendelian ran domization: moving beyond the NOME assumption. International Journal of Epidemiology 2019; 48(3):728– 742, doi:10.1093/ije/dyy258
2019 doi
-
[22]
Detection of widespread ho rizontal pleiotropy in causal relationships inferred from Mendelian randomization betwee n complex traits and diseases
Verbanck M, Chen CY, Neale B, Do R. Detection of widespread ho rizontal pleiotropy in causal relationships inferred from Mendelian randomization betwee n complex traits and diseases. Nature Genetics 2018; 50(5):693–698, doi:10.1038/s41588-018-0099-7
2018 doi
-
[23]
Estimating dose-response relationships for vitamin d with coronary heart disease, stroke, and all-cause mortality: obs ervational and mendelian randomisation analyses
Sofianopoulou E, Kaptoge SK, Afzal S, Jiang T, Gill D, Gunderse n TE, Bolton TR, Allara E, Arnold MG, Mason AM, et al. . Estimating dose-response relationships for vitamin d with coronary heart disease, stroke, and all-cause mortality: obs ervational and mendelian randomisation...
2024
-
[24]
Towards more reliable non-linear mendelian randomiza tion investigations
Burgess S. Towards more reliable non-linear mendelian randomiza tion investigations. European Journal of Epidemiology 2024; 39(5):447–449
2024
-
[25]
The nonlinear two-stage least-squares estimator
Amemiya T. The nonlinear two-stage least-squares estimator. Journal of Econometrics 1974; 2(2):105–110, doi:10.1016/0304-4076(74)90033-5
1974 doi
-
[26]
Linearity in instrumental variables estimat ion: Problems and solutions
Mogstad M, Wiswall M. Linearity in instrumental variables estimat ion: Problems and solutions. Technical Report, Forschungsinstitut zur Zukunft der Arbeit. Bonn, Germany. 2010
2010
-
[27]
Deep iv: A flexib le approach for counterfactual prediction
Hartford J, Lewis G, Leyton-Brown K, Taddy M. Deep iv: A flexib le approach for counterfactual prediction. International Conference on Machine Learning , PMLR, 2017; 1414–1423
2017
-
[28]
Kernel instrumental variable re gression
Singh R, Sahani M, Gretton A. Kernel instrumental variable re gression. Advances in Neural Information Processing Systems 2019; 32:4593–4605
2019
-
[29]
Deep generalized method of mo ments for instrumental variable analysis
Bennett A, Kallus N, Schnabel T. Deep generalized method of mo ments for instrumental variable analysis. Advances in neural information processing systems 2019; 32:3564–3574
2019
-
[30]
Control function instrumental variable estima tion of nonlinear causal effect models
Guo Z, Small DS. Control function instrumental variable estima tion of nonlinear causal effect models. Journal of Machine Learning Research 2016; 17(1):3448–3482
2016
-
[31]
Polynomial Mendelian randomization r eveals non-linear causal effects for obesity-related traits
Sulc J, Sjaarda J, Kutalik Z. Polynomial Mendelian randomization r eveals non-linear causal effects for obesity-related traits. Human Genetics and Genomics Advances 2022; 3(3):100 124
2022
-
[32]
Non-linear Mendelian randomization: detection of biases using negative controls with a fo cus on BMI, Vitamin D and LDL cholesterol
Hamilton FW, Hughes DA, Spiller W, Tilling K, Davey Smith G. Non-linear Mendelian randomization: detection of biases using negative controls with a fo cus on BMI, Vitamin D and LDL cholesterol. European Journal of Epidemiology 2024; 39(5):451–465
2024
-
[33]
Conventional and genetic evidence on alcohol and vascular diseas e aetiology: a prospective study of 500 000 men and women in China
Millwood IY, Walters RG, Mei XW, Guo Y, Yang L, Bian Z, Bennett DA , Chen Y, Dong C, Hu R, et al. . Conventional and genetic evidence on alcohol and vascular diseas e aetiology: a prospective study of 500 000 men and women in China. The Lancet 2019; 393(10183):1831–1842
2019
-
[34]
Alcohol, ALDH2, and esophageal cancer : a meta-analysis which illustrates the potentials and limitations of a Mendelian randomization a pproach
Lewis S, Davey Smith G. Alcohol, ALDH2, and esophageal cancer : a meta-analysis which illustrates the potentials and limitations of a Mendelian randomization a pproach. Cancer Epidemiology Biomarkers & Prevention 2005; 14(8):1967–1971, doi:10.1158/1055-9965. epi-05-0196. 15
2005 doi
-
[35]
Pleiotropy-robust Mendelian rand omization
van Kippersluis H, Rietveld CA. Pleiotropy-robust Mendelian rand omization. Interna- tional Journal of Epidemiology 2018; 47(4):1279–1288, doi:10.1093/ije/dyx002
2018 doi
-
[36]
Detecting and corr ecting for bias in Mendelian randomization analyses using gene-by-environment inter actions
Spiller W, Slichter D, Bowden J, Davey Smith G. Detecting and corr ecting for bias in Mendelian randomization analyses using gene-by-environment inter actions. International Journal of Epidemiology 2019; doi:10.1093/ije/dyy202
2019 doi
-
[37]
The GENIUS approach t o robust Mendelian randomization inference
Tchetgen Tchetgen E, Sun B, Walter S. The GENIUS approach t o robust Mendelian randomization inference. Statistical Science 2021; 36(3):443–464
2021
-
[38]
Meta-regression of genome-wide association studies to estimate a ge-varying genetic effects
Pagoni P, Higgins JP, Lawlor DA, Stergiakouli E, Warrington NM, Morris TT, Tilling K. Meta-regression of genome-wide association studies to estimate a ge-varying genetic effects. European Journal of Epidemiology 2024; 39(3):257–270
2024
-
[39]
Commentary: Interpretation and sensitivity analysis for the localized average causal effect curve
Small DS. Commentary: Interpretation and sensitivity analysis for the localized average causal effect curve. Epidemiology 2014; 25(6):886–888, doi:10.1097/ede. 0000000000000187. 16 Supplementary Material A.1 Artificial intelligence prompt The prompt to write the initial draft of...
2014 doi
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.