{"id":"3da57f05-a9ff-4366-b0d4-ae531a6c2f52","arxiv_id":"2507.11088","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Context-stratified Mendelian randomization, splitting analysis by exogenous subgroups with differing exposure distributions, detects nonlinear causal effects in simulations while producing a null vitamin D result across UK Biobank centres.","lead":"This paper proposes a simple way to use Mendelian randomization to test whether an exposure's causal effect varies across subgroups, by running the analysis separately within regions or centres that have different average exposure levels. It shows in simulations the approach detects non-linear effects, and it reports a null application to vitamin D and heart disease risk in UK Biobank.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The method's central claim depends on contexts being exchangeable except for the exposure distribution; the paper provides no sensitivity analysis, so the trend test cannot distinguish non-linearity from effect modification by a context-level covariate correlated with exposure.","rationale":"We read the paper as a methods proposal whose central claim is that context-stratified MR can validly detect effect heterogeneity and non-linearity. The statistical machinery (ratio estimates per context, Q, meta-regression) is standard and the simulation is reproducible. The most load-bearing threat is not the test statistic's calibration—the paper is transparent about the conservative modified second-order weights—but the exchangeability of contexts. The simulation validates the pipeline only under a DGP where contexts differ solely by a location shift in the exposure; no context-level confounders or effect modifiers are present. In the applied example, centres differ in many ways beyond vitamin D, and the paper explicitly notes that an intrinsic correlation between exposure and geography makes causal attribution harder. We therefore propose a targeted simulation that adds a context-level effect modifier correlated with the exposure mean; this directly tests whether the meta-regression trend can be trusted when the 'other factors comparable' clause of Section 2.2 fails. If the trend test is inflated in this setting, the method's ability to 'investigate non-linearity' is compromised, and the abstract's claim needs qualification. This is the same weakness the reader identified, so we agree; the verdict remains CONDITIONAL pending the sensitivity analysis.","tokens_in":13706,"tokens_out":10785,"duration_ms":142947,"concrete_test":"Extend the simulation in Section 3.1 as follows. In each context k, draw a context-level covariate Z_k that is correlated with the exposure mean α_k (e.g., Z_k = α_k + N(0, 0.2)). Generate Y from a linear exposure effect that is modified by Z_k: y_ik = (0.8 + 0.3 Z_k) x_ik − u_ik + ε_Yik, while keeping the instrument G valid within each context. Apply the paper's heterogeneity Q (modified second-order) and meta-regression trend test across the 10 contexts. Repeat 1000 times and record rejection rates at the 5% level. If the trend test rejects in >5% of replications (or Q >5%), the method conflates context-level effect modification with exposure non-linearity, so the applied null result cannot be interpreted as evidence about the vitamin D dose–response. If the rejection rates stay near 5%, the concern is relieved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that differences in context-specific MR estimates indicate effect heterogeneity or non-linearity in the exposure–response. This requires that contexts are comparable in all outcome-relevant factors other than the exposure distribution (Section 2.2). The paper acknowledges this assumption but does not probe it: the simulation DGP in Section 3.1 has no context-level variable other than the exposure mean α_k, so the only possible source of between-context estimate differences is the exposure–response function itself. In real applications (e.g., UK Biobank centres), context is correlated with latitude, deprivation, lifestyle, and healthcare; any of these could modify the causal effect of the exposure. If a context-level effect modifier is correlated with the context-specific mean exposure, the meta-regression of estimates on mean exposure will produce a non-zero slope even when the exposure–response is linear within every context. The applied example's null result is then uninformative about non-linearity, and a positive result would be ambiguous between effect modification by the context and non-linearity in the exposure. Because the paper provides no negative controls, placebo contexts, or adjustment for context-level covariates, the central claim of the abstract is only supported under an exchangeability assumption that is plausible but unverified in the example.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes context-stratified Mendelian randomization (CS-MR), in which the study population is partitioned by an exogenous context variable (e.g., recruitment centre), separate instrumental-variable analyses are performed within each context, and the resulting context-specific estimates are examined for heterogeneity via Cochran's Q and for trend via meta-regression on the context-specific mean exposure. The method is illustrated with simulations covering linear, quadratic, and threshold exposure–response functions under two between-context exposure-difference scenarios, and with a UK Biobank application estimating the effect of 25-hydroxyvitamin D on coronary artery disease risk across 20 recruitment centres, where no causal effect or heterogeneity is found. The paper argues that CS-MR avoids the strong constant-effect or rank-preserving assumptions required by residual-based and doubly-ranked stratification methods, at the cost of requiring meaningful exogenous between-context exposure variation.","tokens_in":13937,"tokens_out":6852,"duration_ms":74768,"significance":"Strengths of the manuscript include a clearly described simulation design, fully provided R code, and a candid discussion of limitations, particularly the narrow exposure range in the applied example and the potential for context-level confounders. If the method works as claimed, it gives applied researchers a simple, transparent tool for exploratory investigation of effect heterogeneity and non-linearity. However, the simulation evidence only partially supports the abstract's claim of nominal false-positive control, and the absence of any sensitivity analysis for the exchangeability assumption leaves the central identification argument incomplete.","major_comments":[{"comment":"The abstract claims that \"the approach detects heterogeneity when present while maintaining nominal false positive rates under homogeneity when appropriate methods are used.\" Under a linear homogeneous effect, Table 1 shows the first-order Q test rejects at 12.8% (larger differences) and 10.4% (smaller differences), while the modified second-order Q test rejects at 0.4% and 0.1%. Neither version maintains the nominal 5% level, and only the trend test (3.3% and 4.1%) is close to nominal. The phrase \"appropriate methods\" is undefined and the simulation does not identify a heterogeneity test with acceptable false-positive control; the abstract should be revised to describe the actual trade-off between the over-rejecting first-order Q and the under-rejecting modified second-order Q.","section":"Section 3.2 / Table 1 / Abstract"},{"comment":"The method's central claim that between-context differences in estimates indicate effect heterogeneity or non-linearity relies on contexts being comparable in all outcome-relevant factors other than the exposure distribution. The paper acknowledges this requirement but provides no sensitivity analyses, negative controls, or adjustment for context-level covariates. The simulation DGP in Section 3.1 contains no context-level variable other than α_k, so it cannot reveal whether a context-level effect modifier correlated with α_k would produce a spurious trend. In the UK Biobank example, centres differ in latitude, deprivation, and lifestyle; the null result cannot validate exchangeability, and a positive result would be ambiguous between effect modification by context and non-linearity in the exposure–response. Without such sensitivity analyses, the abstract's claim that the method can \"investigate effect heterogeneity and non-linearity\" is under-supported.","section":"Section 2.2 and Section 5.3"},{"comment":"The provided R code assigns `alpha` twice, first with the larger-difference sequence (from 8 by 0.2) and then immediately with the smaller-difference sequence (from 9 by 0.1). As printed, the simulation only runs the smaller-difference scenario, so the larger-difference rows of Table 1 cannot be reproduced from the code without manual editing. The code should be corrected to run both scenarios or clearly comment which line should be uncommented.","section":"Supplementary Material A.2"}],"minor_comments":[{"comment":"The sentence \"with elevated coverage rates for the heterogeneity test using first-order weights\" should read \"rejection rates\" because the context is about false-positive proportions.","section":"Section 3.2"},{"comment":"The left panel's y-axis label reads \"Log odds ratio for coronary heart disease\" but the outcome is coronary artery disease; the label should be consistent with the text and Table 2.","section":"Figure 1"},{"comment":"The phrase \"the outcome was also defined in the same way as in this paper\" should refer to the previous publication [23] rather than \"this paper\".","section":"Section 4"},{"comment":"The centre name \"Middlesborough\" should be \"Middlesbrough\" to match standard spelling.","section":"Table 2"},{"comment":"The meta-regression uses the observed context-specific mean exposure as a regressor, but this mean is estimated with error; the paper does not discuss the impact of this measurement error on the trend test's calibration, although the simulation results suggest the effect is modest.","section":"Section 2.3, Step 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a relevant problem and is clearly written, but the abstract overstates the false-positive control, and the exchangeability assumption needs at least a targeted sensitivity analysis. The code bug in the supplementary is easy to fix. I recommend major revision rather than rejection because the core idea is defensible and the limitations are openly acknowledged; however, the authors must substantially revise the presentation of the simulation results and add a sensitivity analysis for the context-level confounding scenario."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper formalizes a straightforward idea: split the sample by an exogenous context (recruitment centre, region), run standard MR within each context, test heterogeneity with Cochran's Q, and meta-regress the context-specific estimates on mean exposure. Nothing fancy, but the formalization is new, and the simulations are reproducible with code in the supplement. The applied vitamin D example is a clean null, and the authors are honest that the exposure range across centres is narrow.\n\nWhat the paper does well: it states its assumptions clearly, explicitly acknowledges that differences between contexts are only attributable to the exposure if other factors are comparable, and it does not oversell the method. The comparison with residual-based and doubly-ranked stratification is fair, and the discussion of interpretational pitfalls is careful. Credit is due for shipping code and for being transparent about the limitations of the applied example.\n\nThe soft spots, in proportion: first, the stress-test concern lands. The central claim—that a trend in context estimates reflects non-linearity—requires contexts to be exchangeable except for the exposure distribution. The paper says this in Section 2.2 and returns to it in the discussion, but it never probes it. The simulation DGP has no context-level covariate other than the exposure mean, so a context-level effect modifier correlated with exposure would produce exactly the same pattern as non-linearity. The applied example is a null, so it cannot distinguish either. This is not fatal, but it is a real gap: a sensitivity analysis adjusting for context-level covariates, or a placebo context analysis, would make the method's interpretability claims credible. Second, the false-positive control is imperfect: first-order Q over-rejects (10-13% under the null), and modified second-order Q under-rejects substantially (0.1-0.4%). The paper reports these numbers honestly, but the abstract's 'nominal' claim is only true for the conservative choice. Minor, since the text is transparent about the trade-off. Third, the applied example is too under-powered to validate anything; it is an illustration of feasibility, not a demonstration of the method's value.\n\nWho this is for: applied epidemiologists and statisticians working on non-linear MR. It deserves a serious referee. The core idea is simple and reproducible, and the limitations are addressable. I would recommend sending it to peer review with a request for sensitivity analysis around context comparability and a more precise statement about the false-positive properties. Not a rejection; a conditional accept in spirit.","headline":"A simple, honest methods paper that formalizes context-stratified MR; the central interpretative claim rests on an exchangeability assumption the paper flags but does not probe.","tokens_in":14440,"tokens_out":2187,"would_cite":false,"duration_ms":29504,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Context-stratified Mendelian randomization treats recruitment centre, region, or time period as an exogenous stratifier, so between-context differences in exposure become a test for effect heterogeneity and non-linearity.","keywords":["Mendelian randomization","effect heterogeneity","non-linearity","meta-regression","Cochran's Q","context stratification","instrumental variables","collider bias"],"falsifier":"Simulate the linear homogeneous scenario f(x) = 0.8x with a weak instrument (F ≈ 10) and use modified second-order Cochran's Q; if the rejection rate substantially exceeds 5%, the claimed false-positive control under homogeneity would be refuted.","tokens_in":1883,"feed_emoji":"🧬","tokens_out":5337,"duration_ms":145395,"temperature":0.7,"pith_summary":"Context-stratified Mendelian randomization proposes a simple way to test for effect heterogeneity and non-linearity in instrumental variable analyses: split the study population by an exogenous context that shifts the average exposure, run Mendelian randomization separately in each context, then compare the context-specific estimates. The paper claims this avoids the strong unverifiable assumptions of residual-based and doubly-ranked stratification, and shows by simulation that Cochran's Q with modified second-order weights controls false positives while a meta-regression trend detects quadratic and threshold effects when between-context exposure differences are substantial. In the UK Biobank vitamin D example, the approach is feasible but underpowered because centre-specific mean levels range only from 50 to 58 nmol/L, which is why the null result should not be read as a strong no-effect conclusion.","feed_headline":"Regional splits reveal when causal effects vary","feed_subtitle":"Run Mendelian randomization inside each context, then meta-regress estimates on exposure means to spot non-linearity.","key_machinery":"The load-bearing object is the exogenous context variable, a pre-specified categorical variable (for instance, recruitment centre, region, or time period) that is not a function of the exposure, outcome, or instrument, and that induces differences in the average exposure level between subgroups. The method partitions the sample into $K$ contexts, computes a context-specific IV estimate $\\hat\\beta_k$ (e.g., by the ratio method), and then combines two summary statistics: Cochran's $Q$ statistic with either first-order or modified second-order weights to test whether the $\\hat\\beta_k$ vary beyond chance, and the slope from a meta-regression of $\\hat\\beta_k$ on the context-specific mean exposure $\\bar{x}_k$ to test for a dose-response trend. Exogeneity of the context is what prevents collider bias and makes the context-specific estimates locally valid; a significant $Q$ or trend is then a signal of effect heterogeneity or non-linearity, provided other outcome-relevant factors are comparable across contexts.","core_discovery":"The central claim is that a Mendelian randomization estimate can be made context-specific by stratifying on an exogenous variable such as recruitment centre, geographic region, or time period, and that differences across these context-specific estimates provide evidence for effect heterogeneity or non-linearity. Each context-specific estimate is a valid local causal effect under the standard instrumental variable assumptions, without the constant-effect or rank-preserving assumptions needed by residual-based and doubly-ranked methods. The paper demonstrates in simulations that the approach detects quadratic and threshold effects with good power when between-context differences in exposure are large, and that the modified second-order Cochran's Q keeps false-positive rates near nominal in the linear homogeneous scenario. In the applied vitamin D example, no heterogeneity is found (Q p=0.28, trend p=0.76), and the paper notes that the narrow range of context-specific mean exposure limits the method's power and interpretability.","pith_inferences":["A natural extension the paper leaves implicit: when centre-level exposure differences are small, stratifying by a temporal context such as season or year of recruitment could widen the exposure range and increase power.","A useful falsification exercise would be to apply context-stratified MR to a negative-control outcome; if a trend appears there, it would indicate context-level confounding rather than genuine effect modification.","If the method's power depends critically on between-context exposure variation, pooling multiple cohorts or countries into a single context-stratified analysis could make the trend test competitive with residual-based stratification.","A significant Q in a context-stratified analysis could also arise from pleiotropy whose magnitude varies by context; comparing context-specific estimates with those from assumptions-free sensitivity analyses would help distinguish effect modification from assumption violation."],"forward_implications":["If the central claim is right, any Mendelian randomization analysis with a natural context that has meaningful exposure variation can report a set of locally valid effects rather than a single population-averaged estimate, and a significant Q statistic or meta-regression trend reveals that the average hid heterogeneity or non-linearity.","The method offers a checkable alternative to residual-based and doubly-ranked stratification: context-specific estimates are valid without constant-effect or rank-preserving assumptions, at the price of needing genuine between-context exposure variation.","In the vitamin D example, the null findings are consistent with no causal effect, but the narrow 50-58 nmol/L range across centres means the analysis has low power to detect non-linearity; a context with wider exposure variation could change conclusions.","Using modified second-order weights for Cochran's Q controls false-positive rates under homogeneity, whereas first-order weights over-reject with strong instruments and should not be used for null testing."],"supporting_citations":[{"why":"Introduces the residual-based stratification method that context-stratified MR contrasts with.","marker":"[13]"},{"why":"Introduces the doubly-ranked stratification method, which relaxes the constant-effect assumption but still relies on a rank-preserving assumption.","marker":"[14]"},{"why":"Shows that violations of the constant genetic effect assumption can bias non-linear MR, motivating the exogenous-context alternative.","marker":"[15]"},{"why":"Evaluates residual and doubly-ranked methods and demonstrates that their assumptions can fail, supporting the case for context-stratified analysis.","marker":"[16]"},{"why":"Supplies the meta-regression framework used to test for trends in estimates across contexts.","marker":"[18]"},{"why":"Frames Cochran's Q as a heterogeneity statistic for Mendelian randomization, used here to compare context-specific estimates.","marker":"[20]"},{"why":"Gives the modified second-order weights that keep false-positive rates near nominal in the simulations.","marker":"[21]"},{"why":"Documents that heterogeneity tests over-reject with strong instruments, explaining why first-order weights are unsuitable for null testing.","marker":"[22]"},{"why":"Provides the vitamin D instrument, outcome definition, and previous stratified estimates used to compare and interpret the applied example.","marker":"[23]"}],"fun_headline_variants":["Stratify by region to reveal causal effect variation","Context-based MR uncovers effect heterogeneity","Regional split MR detects non-linear causal effects","Context stratification reveals causal effect variation"],"cache_read_input_tokens":16640,"weakest_assumption_plain":"The load-bearing premise is that the context variable is exogenous and that contexts are comparable in all outcome-relevant factors besides the exposure distribution; if a context also differs in a confounder, the between-context estimate differences cannot be attributed to the exposure.","fun_headline_variants_meta":{"raw":{"variants":["Stratify by region to reveal causal effect variation","Context-based MR uncovers effect heterogeneity","Regional split MR detects non-linear causal effects","Context stratification reveals causal effect variation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000523,"raw_usage":{"total_tokens":2555,"prompt_tokens":996,"completion_tokens":1559,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":1505}},"tokens_in":612,"tokens_out":1559,"duration_ms":12132,"temperature":1.0,"reasoning_tokens":1505,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:16:50.414542+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the linear homogeneous scenario f(x) = 0.8x with a weak instrument (F ≈ 10) and use modified second-order Cochran's Q; if the rejection rate substantially exceeds 5%, the claimed false-positive control under homogeneity would be refuted.","supporting_citations":[{"cited_title":"Semiparametric methods for estimation o f a nonlinear exposure- outcome relationship using instrumental variables with application to Mendelian randomization","cited_arxiv_id":null,"evidence_quote":"Introduces the residual-based stratification method that context-stratified MR contrasts with."},{"cited_title":"Relaxing parametric assumpt ions for non-linear Mendelian randomization using a doubly-ranked stratiﬁcation metho d","cited_arxiv_id":null,"evidence_quote":"Introduces the doubly-ranked stratification method, which relaxes the constant-effect assumption but still relies on a rank-preserving assumption."},{"cited_title":"Violation of the constant genetic eﬀect assumption can result in biased estimates for non-linear Mendelian randomization","cited_arxiv_id":null,"evidence_quote":"Shows that violations of the constant genetic effect assumption can bias non-linear MR, motivating the exogenous-context alternative."},{"cited_title":"Non-linear Mendelian randomization: evaluation of eﬀe ct modiﬁcation in the residual and doubly-ranked methods with simulated and empirica l examples","cited_arxiv_id":null,"evidence_quote":"Evaluates residual and doubly-ranked methods and demonstrates that their assumptions can fail, supporting the case for context-stratified analysis."},{"cited_title":"How should meta-regression analyses be undertaken and interpreted? Statistics in medicine 2002; 21(11):1559–1573","cited_arxiv_id":null,"evidence_quote":"Supplies the meta-regression framework used to test for trends in estimates across contexts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames Cochran's Q as a heterogeneity statistic for Mendelian randomization, used here to compare context-specific estimates."},{"cited_title":"Detection of widespread ho rizontal pleiotropy in causal relationships inferred from Mendelian randomization betwee n complex traits and diseases","cited_arxiv_id":null,"evidence_quote":"Documents that heterogeneity tests over-reject with strong instruments, explaining why first-order weights are unsuitable for null testing."},{"cited_title":"Estimating dose-response relationships for vitamin d with coronary heart disease, stroke, and all-cause mortality: obs ervational and mendelian randomisation analyses","cited_arxiv_id":null,"evidence_quote":"Provides the vitamin D instrument, outcome definition, and previous stratified estimates used to compare and interpret the applied example."}],"review_version":1}