{"id":"1415d8dd-d5c5-4205-9feb-3014e2f3e768","arxiv_id":"2411.10648","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A subsampling-based studentized Sobel test with a pivotal null distribution, combined across splits by a Cauchy combination, gives accurate false-positive control in mediation analysis.","lead":"This paper introduces a new test for whether a mediator carries an exposure's effect on an outcome, designed to work correctly under all three forms of the 'no mediation' null. It splits the data into pieces, computes a simple Sobel statistic on each piece, then combines them into a single well-behaved test statistic.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dependent Cauchy combination of split p-values is the soft spot: no theorem in the paper guarantees the final CSMT size, though Liu-Xie (2020) may fill the gap.","rationale":"The reader's weakest assumption identifies exactly the place where the paper's central claim is least secured. Theorem 2 is a correct and useful result for a single random split, but the proposed CSMT method goes beyond it by applying a Cauchy combination to M dependent split p-values. The paper does not provide a theorem for this final procedure, so a reader cannot verify from the manuscript alone that the standard Cauchy quantile controls size under the dependence induced by reusing the same data. This is a genuine proof gap, not merely a stylistic omission: the abstract claims accurate size for CSMT, and the only formal support in the text is the single-split result plus a citation to Liu-Xie. The cited result is likely to fill the gap, because Liu-Xie's Cauchy combination is designed for arbitrary dependence and requires only marginal validity of the p-values; Theorem 2 supplies asymptotic marginal validity. Thus the concern is real but probably addressable, which supports the reader's CONDITIONAL verdict rather than a rejection. I would not label the paper internally inconsistent, because nothing in Theorem 2 or the CSMT construction appears false; the issue is a missing formal link between the single-split theorem and the combined procedure. The proposed Monte Carlo check is a direct way to see whether the missing link actually breaks size control in the settings the paper recommends.","tokens_in":7103,"tokens_out":24218,"duration_ms":260319,"concrete_test":"Monte Carlo check: simulate 5,000 datasets from Model (2) under each null case (H00, H01, H10) with n=600, K=12, and M=500, using the paper's R code; compute C for each dataset and record the empirical rejection rate at alpha=0.05. If the empirical size is at most 0.05 and the empirical right-tail quantile of C does not exceed the standard Cauchy quantile at the corresponding level, the dependent-Cauchy combination is size-valid for the recommended settings. If it exceeds 0.05, the concern lands and the central size-control claim for CSMT fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Single-split Theorem 2 gives a pivotal t null for T_n, so the core construction is internally sound. But CSMT is the procedure users actually apply: it forms M dependent p-values from random splits of the same data and rejects using the standard Cauchy quantile on C = sum_m w_m tan(pi(0.5-p_m)). Section 4 asserts C follows a standard Cauchy and cites Liu-Xie, but it never states the Liu-Xie theorem or verifies its conditions in this setting. Liu-Xie's arbitrary-dependence guarantee is a tail bound, not an exact Cauchy law: it holds when each p_m is a valid (uniform or conservative) null p-value. Here p_m is only asymptotically t-calibrated by Theorem 2, and finite-sample validity under all three null cases is not proved. If the marginal p_m are not valid, or if the dependence violates a condition of Liu-Xie's theorem, the final test's size could deviate from nominal. Since the abstract's accuracy claim is about CSMT, this missing link is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a subsampling-based test for the composite null of no mediation effect (alpha*beta = 0). The idea is to split the data into K disjoint subsamples, compute the Sobel statistic in each subsample, and then studentize their mean to obtain a statistic T_n whose asymptotic null distribution is claimed to be t with K-1 degrees of freedom under all three null configurations (H00, H01, H10). This is the content of Theorem 2. To reduce variability across different random splits and improve power, the paper then combines p-values from M random splits using the Cauchy combination test, calling the resulting procedure CSMT. Section 5 recommends K = floor(0.5 sqrt(n)), and Section 6 reports simulation comparisons with Sobel, MaxP, and an adaptive bootstrap test, followed by a real-data analysis of the SPIRIT trial.","tokens_in":7295,"tokens_out":9094,"duration_ms":106991,"significance":"If the main claim is correct, the paper offers a simple, pivotal null distribution for a composite-null mediation problem, which is a genuinely useful contribution because standard Sobel and MaxP tests are conservative under H00. Theorem 2 is elementary and self-contained, with no fitted quantities entering the null distribution, which is a strength. The paper also provides an R implementation and reports extensive simulations. However, the actual procedure recommended to users, CSMT, relies on a Cauchy combination of dependent p-values for which no theorem is stated or proved; this gap is load-bearing for the paper's headline claim of accurate size control. The fixed-K theory also does not directly cover the recommended growing K.","major_comments":[{"comment":"The claim that the combined statistic C_m follows a standard Cauchy distribution, and hence that the p-value is 0.5 - arctan(C_m)/pi, is not justified. The p-values p_m are computed from M random splits of the same data and are therefore dependent. Liu and Xie (2020) is cited, but the paper neither states the precise theorem being used nor verifies its conditions. In particular, Theorem 2 gives only marginal asymptotic t-calibration of each p_m, not finite-sample validity, and arbitrary dependence among valid p-values does not make the weighted sum of their Cauchy transforms exactly Cauchy. Unless the authors state the relevant Liu-Xie result and prove that its conditions hold (or prove an asymptotic version for their setting), the size control of CSMT is not established by the theoretical part of the manuscript.","section":"Section 4"},{"comment":"Theorem 2 is stated and proved only for fixed K, but Section 5 recommends K = floor(0.5 sqrt(n)), which grows with n. The manuscript does not provide a theorem that covers K = K_n tending to infinity, even though in this regime the subsample sizes n/K_n also grow. The authors should either extend the asymptotic result to K_n -> infinity with n/K_n -> infinity, or explicitly limit the theoretical claim to a fixed K chosen by the user and describe the recommended K as a finite-sample heuristic. As written, the recommended implementation is not covered by the stated theorem.","section":"Theorem 2 and Section 5"},{"comment":"The proof of Theorem 2 says the result follows by applying Theorem 1 to each subsample and the continuous mapping theorem, but this omits an essential step: one must establish joint convergence of (S_{G_1}, ..., S_{G_K}) to independent normal variables, and then verify that the sample variance of the S_{G_i} is consistent. Since the subsamples are disjoint and K is fixed, this is straightforward, but it should be stated explicitly. Additionally, in the proof of Theorem 1 under H01/H10, the variance estimator is written as \\hat{\\alpha}^2 s_1^2 + \\hat{\\beta}^2 s_2^2, whereas the Sobel statistic in equation (1) uses \\hat{\\alpha}^2 s_2^2 + \\hat{\\beta}^2 s_1^2; the subscripts appear to be swapped.","section":"Proof of Theorem 2"}],"minor_comments":[{"comment":"The notation is inconsistent: the statistic is first written as C_m with index m ranging over splits, but later the same symbol C_m is used as the combined statistic; the index m is reused in both the denominator and the numerator of the weighted sum.","section":"Section 4"},{"comment":"The sentence 'where wms are non-negative weights satisfying that sum_{m=1}^M w_m = 1' should read 'non-negative weights w_m satisfying ...' to avoid a missing subscript.","section":"Section 4"},{"comment":"The recommendation K = floor(0.5 sqrt(n)) is based only on empirical evidence summarized in the supplementary materials; the main text should clarify whether this choice is meant to be applied as a deterministic function of n or as a practical guideline independent of the asymptotic theory.","section":"Section 5"},{"comment":"In Example 2, the text says 'the power of the ABtest decreases as the signal strength increases,' but the accompanying Figure 3 is not described in enough detail; a brief explanation of why this occurs would improve readability.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The core idea of studentizing subsample Sobel statistics is sound and clean, and Theorem 2 is correct as a fixed-K statement. The main obstacle to acceptance is the missing theoretical support for the Cauchy combination step in Section 4; this is a fixable gap if the authors provide a precise theorem and conditions, but it is not merely a presentation issue. The growing-K mismatch between Theorem 2 and Section 5 is a second gap that should also be addressed. I would be willing to see a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper has a correct and elegant core result, but the headline method (CSMT) relies on a step that is not fully proven. The idea: split data into K subsamples, compute the Sobel statistic in each, and studentize with the subsample variance to get T_n. Theorem 2 shows that under every composite-null case, T_n converges to a t_{K-1} distribution. That gives a pivotal null without knowing which of H00, H01, H10 generated the data—a genuine improvement over classical Sobel and MaxP. The proof is a clean application of Glonek's result and the continuous mapping theorem; I checked the steps and they hold. The studentization is the right fix for the different variance scaling under H00 vs H01/H10.\n\nThe soft spot is the final Cauchy combination. The paper takes M random splits, gets p-values p_m from the single-split test, and combines them via Liu-Xie's Cauchy method, treating C as standard Cauchy. The problem: p_m are dependent across splits—the subsamples overlap heavily—and the paper neither states Liu-Xie's theorem nor verifies its conditions. Liu-Xie's guarantee is a tail bound under arbitrary dependence, but it requires each p_m to be a valid null p-value. Here we only have asymptotic t-calibration, and the dependence persists as n grows. So the size of CSMT is not actually proven; it's supported by simulations. For a methods paper, that's a real gap, especially since CSMT is the procedure users would apply. I suspect it's fixable—maybe a direct proof that the combined statistic is asymptotically standard Cauchy, or a careful appeal to Liu-Xie with stronger finite-sample control—but as written, the abstract overstates what's established.\n\nMinor issues: a subscripts typo in the proof of Theorem 1 (the variance estimator is written with s1/s2 swapped), and the recommended K = floor(0.5 sqrt(n)) is justified only by simulation. The power plots show ABtest beating CSMT at weak signals; the authors note ABtest has inflated size, so the comparison is fair, but the abstract's 'higher power' claim should be read in that context. The real-data analysis is exploratory and doesn't add much.\n\nWhat's new: the studentized subsample construction for the mediation null. I don't think that specific trick appears in the cited literature. Code is provided.\n\nI'd engage with this. The core theorem is solid, and the missing Cauchy-link is a well-defined problem rather than a fundamental flaw. Worth a serious referee.","headline":"Correct and elegant subsample studentization for the mediation null, but the final Cauchy combination step is not fully justified and needs a theorem or a condition check before CSMT is presented as validated.","tokens_in":7796,"tokens_out":3135,"would_cite":false,"duration_ms":34306,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62F05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Splitting the sample into subsamples and studentizing per-subsample Sobel statistics yields a t-distribution null that is the same for all three composite-null cases in mediation analysis.","keywords":["mediation analysis","composite null hypothesis","Sobel test","subsampling","studentization","Cauchy combination","size control"],"falsifier":"Simulate data under H00 (both coefficients zero) with large n, fix one dataset, compute p-values from M = 500 random splits, form the Cauchy-combined statistic C_m, and repeat over many datasets; if the empirical distribution of C_m deviates substantially from the standard Cauchy in the tail, the pivotal rejection region is not valid.","tokens_in":6914,"feed_emoji":"📊","tokens_out":6319,"duration_ms":56876,"temperature":0.7,"pith_summary":"Testing for a mediation effect is hard because the null hypothesis of no mediation is a union of three distinct cases: the exposure–mediator effect is zero, the mediator–outcome effect is zero, or both are zero. Classical tests such as Sobel's are calibrated for only one case and turn conservative when both effects vanish. The paper shows that splitting the sample into K disjoint subsamples, computing a Sobel statistic in each, and studentizing these values gives a statistic that converges to a t-distribution with K−1 degrees of freedom under every null case, making the null distribution pivotal. Repeating this across many random splits and combining p-values with a Cauchy combination test yields a procedure, called CSMT, that controls size accurately and shows competitive power in simulations.","feed_headline":"Split the sample and every mediation null becomes a t-test","feed_subtitle":"Subsample Sobel statistics, studentized across splits, share one pivotal null distribution, restoring accurate size.","key_machinery":"The central object is the studentized subsample Sobel statistic, built by partitioning the data into K disjoint subsamples, computing the classical Sobel statistic S_{G_i} within each, and applying a one-sample t-statistic to the K resulting values. Because each S_{G_i} is asymptotically normal under every null scenario but with a variance that depends on which null case holds, the studentization cancels the unknown scale and leaves a pivotal t_{K−1} limit. The second piece is the Cauchy combination test, which aggregates p-values from M random splits using random positive weights; the paper rejects when the combined statistic exceeds the standard Cauchy quantile.","core_discovery":"The central claim is Theorem 2: under the composite null H0: αβ = 0, the studentized subsample statistic T_n = $K^{{1/2}}$ \\bar{S}_K / ( (1/(K−1)) \\sum (S_{G_i} − \\bar{S}_K)^2 )^{1/2} converges in distribution to a t random variable with K−1 degrees of freedom. Each subsample Sobel statistic S_{G_i} is asymptotically normal with mean zero and variance either 1 (when exactly one of α and β is zero) or 1/4 (when both are zero), so the sample variance across subsamples estimates and removes the scale. The null distribution is therefore the same in all three null cases, requiring a single cutoff. Repeating the split M times and combining the resulting p-values through a weighted Cauchy combination test gives CSMT, which the paper demonstrates by simulation to control size more accurately than the Sobel test, the MaxP test, and an adaptive bootstrap test while achieving higher power at stronger signals.","pith_inferences":["Editorial inference: the proof of Theorem 2 only uses asymptotic normality of the per-subsample statistics, so the same subsample-studentization device could pivotailize other test statistics whose normal limit has a null-dependent variance.","Editorial inference: the recommended choice K = ⌊0.5√n⌋ is based on empirical evidence; a data-driven or cross-validated choice of K has not been explored and could improve finite-sample behavior.","Editorial inference: because the Cauchy combination's null distribution is assumed rather than proved under the dependence across splits, practitioners using CSMT on smaller samples should verify the empirical null by permutation before trusting nominal p-values.","Editorial inference: the SPIRIT data analysis points to butyric, acetic, and valeric acids as candidate mediators of metformin's anti-inflammatory effect; a confirmatory study with multiple-testing adjustment would be needed to make this a robust scientific claim."],"forward_implications":["A single t_{K−1} cutoff can be used for mediation testing regardless of which of the three null cases holds, eliminating the need to know the null type.","Combining p-values from repeated random splits stabilizes the test results and raises power, as shown in the numerical studies.","CSMT controls size more accurately than the classical Sobel and MaxP tests, which are conservative when both coefficients are zero.","At stronger signal strengths, CSMT outperforms the adaptive bootstrap test while avoiding that test's inflated size.","The method remains valid for the general structural equation framework covered by the assumptions, not only for linear models."],"supporting_citations":[{"why":"Defines the original Sobel statistic that the paper recomputes inside each subsample.","marker":"[Sobel, 1982]"},{"why":"Proves the asymptotic normality of the Sobel statistic under the composite null, the foundation for Theorem 1 and hence Theorem 2.","marker":"[Glonek, 1993]"},{"why":"Supplies the MaxP test that the paper uses as a benchmark in size and power comparisons.","marker":"[MacKinnon et al., 2002]"},{"why":"Provides the Cauchy combination test used to aggregate p-values across random splits.","marker":"[Liu and Xie, 2020]"},{"why":"Supplies the adaptive bootstrap test benchmark and the power-simulation setup the paper adopts.","marker":"[He et al., 2023]"}],"fun_headline_variants":["Subsampling makes every mediation null a t-test","Cauchy-combined subsamples fix mediation tests","One t-distribution for all mediation nulls","Subsampling gives mediation tests a power boost","Composite null no match for subsample Sobel"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The procedure's size control ultimately depends on the Cauchy combination of p-values from different random splits following the standard Cauchy distribution even though those p-values are dependent, and the paper does not prove this.","fun_headline_variants_meta":{"raw":{"variants":["Subsampling makes every mediation null a t-test","Cauchy-combined subsamples fix mediation tests","One t-distribution for all mediation nulls","Subsampling gives mediation tests a power boost","Composite null no match for subsample Sobel"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000671,"raw_usage":{"total_tokens":3023,"prompt_tokens":875,"completion_tokens":2148,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":2085}},"tokens_in":491,"tokens_out":2148,"duration_ms":15958,"temperature":1.0,"reasoning_tokens":2085,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:28:30.741555+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate data under H00 (both coefficients zero) with large n, fix one dataset, compute p-values from M = 500 random splits, form the Cauchy-combined statistic C_m, and repeat over many datasets; if the empirical distribution of C_m deviates substantially from the standard Cauchy in the tail, the pivotal rejection region is not valid.","supporting_citations":[{"cited_title":"Asymptotic confidence intervals for indirect effects in structural equation models","cited_arxiv_id":null,"evidence_quote":"Defines the original Sobel statistic that the paper recomputes inside each subsample."},{"cited_title":"On the behaviour of wald statistics for the disjunction of two regular hypotheses","cited_arxiv_id":null,"evidence_quote":"Proves the asymptotic normality of the Sobel statistic under the composite null, the foundation for Theorem 1 and hence Theorem 2."},{"cited_title":"A comparison of methods to test mediation and other intervening variable effects","cited_arxiv_id":null,"evidence_quote":"Supplies the MaxP test that the paper uses as a benchmark in size and power comparisons."},{"cited_title":"Cauchy combination test: a powerful test with analytic p-value calculation under arbitrary dependency structures","cited_arxiv_id":null,"evidence_quote":"Provides the Cauchy combination test used to aggregate p-values across random splits."}],"review_version":1}