{"id":"01c45499-9503-4ccc-a555-e7ec9acf02c1","arxiv_id":"2607.17596","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In multiverse analyses, converting p-values to e-values via a mixture calibrator yields more powerful false-discovery control than universal mixture or soft-rank e-values, and it flags four internet/social-media–well-being associations in UK teen data.","lead":"This paper compares ways to correct for multiple testing when researchers run a \"multiverse\" of many treatment–outcome–subgroup analyses, using e-values to control false discoveries without assuming tests are independent. It finds that converting ordinary p-values into e-values gives more statistical power than two other e-value methods, and applies the approach to teenager technology use and mental well-being.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Finite-sample anti-conservatism in LRT p-values can break the e-variable property of Ec1/Ec2 and with it e-BH FDR control; Table 1 does not test the e-value expectation.","rationale":"I read the paper as making a two-part contribution: (i) in practical multiverse settings, p-to-e calibrated e-values, especially Ec1, dominate universal-mixture and soft-rank e-values in power; (ii) applying Ec1 + e-BH to the MCS data gives valid, dependence-robust FDR control and yields four practically relevant rejections. The weakest support is not the e-value algebra (Eqs. 3, 5, 7-8 are standard and correctly cited) but the conversion of asymptotic LRT p-values into exact e-variables in finite samples. This is exactly the reader's weakest assumption. The concern is real because e-BH's FDR guarantee is unconditional on E[E]≤1; a p-to-e calibrator can have controlled marginal rejection probability while violating the e-variable expectation, and Table 1's 100-repetition type I rates are too coarse to detect this in the tail region that drives e-BH. I do not think this warrants rejection: the authors disclose the caveat, the Gaussian case is exact, and the simulation evidence is broadly consistent with adequate control at n ≥ 100. But the claim as stated ('e-values... guarantee FDR control regardless of dependence') needs a finite-n qualification for non-Gaussian GLMs, and the application's four rejections should be backed by an FDR simulation in the actual 40-test/correlated design. The reader's CONDITIONAL verdict already asks for exactly this kind of evidence, so no change is needed.","tokens_in":12108,"tokens_out":11908,"duration_ms":103518,"concrete_test":"Simulate the Section 4 design under the global null with K = 40: logistic GLMs with n ∈ {100, 300, 500}, outcome prevalence around 5-10%, correlated treatments and outcomes as in the paper's design. For each of ≥1000 datasets, compute the LRT p-values, apply Ec1/Ec2, run e-BH at α = 0.05, and estimate FDR (with standard error) and average Ec1 for each hypothesis. Reject the paper's practical claim if FDR exceeds 0.05 by more than Monte Carlo error, or if average null Ec1 clearly exceeds 1. Also repeat with Firth-corrected or permutation p-values to isolate the chi-square approximation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline practical guarantee is FDR control via e-BH, which is valid only if every E_lj is a genuine e-variable: E_P[E_lj] ≤ 1 under each null. For p-to-e calibrators (7)-(8), this is equivalent to the input p-values being super-uniform under the null. In the non-Gaussian GLMs that drive Section 4, the p-values are LRT p-values based on chi-square asymptotics, not exact finite-n p-values. The authors explicitly flag this in Section 3.1 ('Since such a p-value is not exact for finite n, neither is the resulting calibrated e-value'), but that caveat sits on the central claim: if finite-n p-values are even mildly anti-conservative, Ec1 and Ec2 are not e-variables and the e-BH FDR guarantee collapses. Table 1 only reports per-test rejection proportions at fixed α; a p-value can have acceptable tail rejection rates yet still give E[Ec1(P)] > 1 because Ec1(p) ~ 1/(p log^2 p) as p→0, making the expectation highly sensitive to small-p behavior. The multiverse application involves 40 correlated hypotheses, so a few inflated null e-values could produce the four rejections even if every marginal type I rate is below 5%. Thus the practical finding depends on an unvalidated finite-sample property, not on the e-value mathematics.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using e-values to control the false discovery rate in GLM-based multiverse analyses. It reviews three e-value constructions: universal mixture e-variables (Eq. 3), soft-rank e-variables (Eq. 5), and p-to-e calibrators (Eqs. 7-8), proves the validity of the universal mixture, and compares power via simulation. Based on average log e-values in linear and logistic regression, it recommends the mixture calibrator Ec1 and applies it with e-BH to the Millennium Cohort Study teenage well-being data, reporting four significant outcome-treatment associations. The central claim is that simple p-to-e calibration dominates generic e-value strategies in this setting.","tokens_in":1364,"tokens_out":1558,"duration_ms":113182,"significance":"If established, the paper offers a practical, dependence-robust FDR control method for multiverse analyses, avoiding independence or positive-dependence assumptions and nuisance parameter estimation. Strengths include a correct e-variable validity proof in Section 3.1, explicit discussion of the finite-sample caveat, and reproducible code/data availability. The substantive application is a useful counterpoint to specification-curve averaging. However, the empirical support for the headline power comparison is thin, and the finite-sample validity of the recommended calibrator is not verified at the level needed for strict FDR guarantees.","major_comments":[{"comment":"The statement that p-to-e calibration 'significantly outperforms' the other approaches in statistical power is not supported by the reported evidence. The figures plot average log e-values from 100 repetitions with no error bars, confidence intervals, or paired comparisons, and an average log e-value is not a power curve. There are no simulations of e-BH FDR control in a multiverse setting with multiple correlated outcomes/treatments, which is the paper's stated target. Please add Monte Carlo intervals, formal comparisons, and FDR/power simulations with realistic multiverse dimensions, or weaken the claim.","section":"Section 3.2, Figures 1-2, Abstract"},{"comment":"p-to-e calibrators are e-variables only when input p-values are super-uniform under the null. For logistic regression, the LRT p-values are asymptotic only; the paper acknowledges this, yet relies on the e-BH finite-sample FDR guarantee. Table 1 reports only rejection proportions at alpha=0.01 and 0.05, which does not validate E[Ec1(P)] <= 1. Since Ec1(p) ~ 1/(p log^2 p) as p goes to 0, its expectation is sensitive to small-p behavior even if tail rejection rates are controlled. This is load-bearing for Section 4. Please prove or empirically validate finite-sample super-uniformity (e.g., report null means of calibrated e-values with Monte Carlo intervals), use exact/permutation p-values, or label the FDR control as asymptotic.","section":"Section 3.1, Eqs. (7)-(8), Table 1"},{"comment":"The displayed null H_lj includes 'beta_j' != 0 for j' != j', contradicting the following sentence that no restrictions are placed on remaining treatment effects. The global H_l (all effects nonzero) is not the relevant alternative for a single test H_lj. This affects the definition of q_l in Eq. (3) and the interpretation of all simulation and application results. Please clarify the hypothesis space and the alternative under which q_l is specified.","section":"Section 2, Eq. (2)"},{"comment":"The sentence 'there were six hypotheses for which 1/Ec1 > 20' appears inverted; the individual-test threshold at level 0.05 is Ec1 > 20 (equivalently 1/Ec1 < 0.05). In addition, the application's four rejections disappear under universal and soft-rank e-values; this is informative, but it means the headline result hinges entirely on the asymptotic validity of Ec1 and should be presented together with the caveat in the second major comment.","section":"Section 4, Table 2"}],"minor_comments":[{"comment":"Typos: 'self-steem' should be 'self-esteem'; 'sensitvy' should be 'sensitivity'.","section":"Abstract, Section 1"},{"comment":"The inverse-gamma notation IG(a/2,l/2) is confusing because l is also used to index outcomes. Use a different symbol such as b/2.","section":"Section 3.1"},{"comment":"The caption and text do not specify the null coefficient values used for the type I error rows. State explicitly which beta* are set to zero.","section":"Table 1"},{"comment":"Captions say 'average over 100 repetitions' but no uncertainty is displayed. Add error bars or confidence intervals, or state that none are shown.","section":"Figures 1-2"},{"comment":"The closed-form marginal likelihood uses inconsistent notation for sigma^2_l and sigma-hat^2_l. Clarify the parameterization of the inverse-gamma prior.","section":"Section 3.1"},{"comment":"The claim that odds ratios in [2,4] are 'practically relevant' could use a sentence on effect-size calibration or external benchmarks, since practical relevance is a subjective judgment.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is in scope and the mathematical core is sound, but the headline empirical claim and the finite-sample validity of the recommended calibrator need work before publication. The issues appear fixable: add proper Monte Carlo uncertainty and FDR simulations, clarify Equation (2), and either validate E[Ec1(P)] <= 1 under the null or reframe the FDR guarantee as asymptotic. I would not reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent, useful comparison paper, not a breakthrough. The genuinely new bit is the simulation-based comparison showing p-to-e calibration (especially the mixture calibrator Ec1) beats universal mixture and soft-rank e-values for GLM multiverse testing, plus the MCS application that flips the earlier 'negligible' finding. The math for e-value validity is standard and correct, and the authors explicitly flag the finite-sample caveat for LRT p-values. Worth a proper look, but the authors oversell 'significantly outperforms' relative to the evidence, and the FDR guarantee rests on a property the simulations don't actually check.\n\nWhat's good: the framing is honest, the comparison is fair, and the practical recommendation (Ec1 + e-BH for multiverse GLM analyses) is concrete. The application is a nice demonstration that e-value FDR control can change conclusions from SCA. The authors also clearly state that e-values are conservative and that power can be moderate—rare candor.\n\nSoft spots, in order of severity. First, the stress-test concern lands: Table 1 reports rejection proportions at alpha=1% and 5%, but the e-BH guarantee requires E[E]<=1, not just tail control. A p-value can be slightly anti-conservative and still give rejection rates below alpha, while E[Ec1(P)]>1. So the finite-n validity of Ec1/Ec2 is not demonstrated; the authors' acknowledgment is honest but the simulation doesn't close the gap. For the MCS application n is large, so this is minor for the headline application, but for the general recommendation it matters. Second, 'significantly outperforms' is not supported: Figures 1-2 are averages without error bars and there are no MC intervals in Table 1. The dominance is visually clear in many scenarios, so the conclusion is probably right, but the wording is stronger than the numbers. Third, the code/data statement says 'dedicated Github repository' with no URL or commit—unverifiable as is. Fourth, the pre-screening before e-BH in Section 4 is ambiguous: six hypotheses with 1/Ec1>20, then refined e-BH; it's likely fine but should be spelled out.\n\nAll fixable. The central comparative finding holds up for the investigated regimes. This deserves a serious referee; I'd want a revision that adds proper Monte Carlo error bars or a direct check of e-variable validity, and a real code link. For readers in multiverse analysis or applied e-value methodology, this is a useful addition. I'd cite the mixture-calibrator recommendation if I worked in that area.","headline":"Useful comparison paper; the mixture p-to-e calibrator recommendation is plausible but oversold by 'significant' without error bars, and finite-n e-variable validity is asserted rather than checked.","tokens_in":12941,"tokens_out":2626,"would_cite":true,"duration_ms":23163,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62J12","62H15"],"pacs":[],"model":"deepseek-v4-flash","headline":"A simple p-to-e calibration rule can outperform generic e-value strategies for false-discovery control in multiverse analysis, and reveals previously missed associations between teen technology use and mental well-being.","keywords":["multiverse analysis","e-values","false discovery rate","multiple testing","generalized linear models","p-to-e calibration","teenager well-being"],"falsifier":"Run the paper's logistic-regression setting with n = 50 and a rare binary outcome under the null, applying Ec1 to the asymptotic likelihood-ratio p-value; if the fraction of rejected hypotheses at level α exceeds α over many replications, the p-value is anti-conservative and the e-variable guarantee fails in that regime.","tokens_in":11930,"feed_emoji":"📊","tokens_out":5438,"duration_ms":46409,"temperature":0.7,"pith_summary":"This paper argues that multiverse analyses—testing many treatment–outcome pairs at once—can be made statistically honest at modest cost by using e-values instead of p-values. It shows that a simple conversion of ordinary likelihood-ratio p-values into e-values, called p-to-e calibration, delivers higher statistical power than generic universal or soft-rank e-value constructions in generalized-linear-model settings. The paper's core claim is that this calibration rule, combined with the e-BH procedure, controls the false discovery rate regardless of how the individual tests depend on each other. Applied to a large survey of teenagers, the method finds four significant technology–wellbeing associations that a previous aggregate specification-curve analysis had obscured. The sympathetic reader should take the central message to be that e-values can make multiverse analysis inferentially valid without needing unrealistic assumptions about test dependence.","feed_headline":"P-to-e calibration beats generic e-values in multiverse scans","feed_subtitle":"A simple p-value-to-e-value conversion controls false discoveries under arbitrary dependence and finds teen tech–wellbeing links a prior ana","key_machinery":"The central object is the mixture p-to-e calibrator Ec1(π) = (1 − π + π log π) / [π (log π)²], which maps a p-value π into an e-value—a test statistic whose expectation under the null is at most 1. The calibrator does the power work by blending a family of simple calibrator functions. With e-values in hand, the e-BH procedure rejects hypotheses with the largest e-values, up to a threshold that uses the order statistics of the e-values, and this controls the FDR without requiring independence, positive dependence, or estimation of nuisance parameters.","core_discovery":"The paper claims that, for GLM-based multiverse analyses, p-to-e calibration—especially the mixture calibrator Ec1—has higher statistical power than universal mixture e-values and soft-rank e-values, while retaining the e-value property that guarantees FDR control under arbitrary dependence via the e-BH procedure. In a cohort of 11,884 UK teenagers, applying Ec1 with refinement to e-BH rejects four outcome–treatment pairs: internet use with low self-esteem, internet use with depressive symptoms, internet use with high peer problems, and social media use with depressive symptoms, with estimated odds ratios between roughly 2 and 4. The paper interprets these as practically relevant association","pith_inferences":["The paper does not say this, but the same calibration recipe would extend to permutation-based p-values, which are super-uniform by construction and would remove the asymptotic caveat outside Gaussian models.","Because e-BH is dependence-robust, the framework could be used to compare results across different control-covariate subsets without modeling their correlation; whether the four findings survive such specification choices remains open.","E-values compose across independent datasets, so the approach could be extended to anytime-valid pooling of subpopulation or cohort analyses, which would be useful for ongoing surveillance of technology–wellbeing associations."],"forward_implications":["Multiverse analyses can achieve finite-sample FDR control without assuming independence or positive dependence among the multiple tests.","The mixture calibrator Ec1 is preferable to universal and soft-rank e-values for well-specified GLM-based multiverse testing.","E-values remain a conservative tool: statistical power may be moderate unless sample sizes and effect sizes are sufficiently large.","The teenager well-being application indicates that averaging technology effects across treatments, as specification-curve analysis does, can hide strong and practically relevant associations."],"fun_headline_variants":["P-to-e calibration boosts power in multiverse analysis","Multiverse scans: p-to-e calibration controls false discoveries","E-value method reveals teen tech use tied to depression, self-esteem","P-to-e beats generic e-values in multiverse FDR control"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The FDR guarantee rests on the calibrated e-values being true e-values, which holds only if the likelihood-ratio p-values fed into the calibrator are super-uniform under the null—an assumption the paper verifies by simulation for n ≥ 100 but does not prove for finite samples in non-Gaussian models.","fun_headline_variants_meta":{"raw":{"variants":["P-to-e calibration boosts power in multiverse analysis","Multiverse scans: p-to-e calibration controls false discoveries","E-value method reveals teen tech use tied to depression, self-esteem","P-to-e beats generic e-values in multiverse FDR control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000598,"raw_usage":{"total_tokens":2627,"prompt_tokens":733,"completion_tokens":1894,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":1836}},"tokens_in":477,"tokens_out":1894,"duration_ms":11364,"temperature":1.0,"reasoning_tokens":1836,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T17:32:48.174096+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's logistic-regression setting with n = 50 and a rare binary outcome under the null, applying Ec1 to the asymptotic likelihood-ratio p-value; if the fraction of rejected hypotheses at level α exceeds α over many replications, the p-value is anti-conservative and the e-variable guarantee fails in that regime.","supporting_citations":[],"review_version":1}