REVIEW 4 major objections 6 minor 24 references
A simple p-to-e calibration rule can outperform generic e-value strategies for false-discovery control in multiverse analysis, and reveals previously missed associations between teen technology use and mental well-being.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 17:32 UTC pith:RA5MRP4T
load-bearing objection Useful comparison paper; the mixture p-to-e calibrator recommendation is plausible but oversold by 'significant' without error bars, and finite-n e-variable validity is asserted rather than checked. the 4 major comments →
E-Values For Multiplicity Control In Multiverse Analysis
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that, for GLM-based multiverse analyses, p-to-e calibration—especially the mixture calibrator Ec1—has higher statistical power than universal mixture e-values and soft-rank e-values, while retaining the e-value property that guarantees FDR control under arbitrary dependence via the e-BH procedure. In a cohort of 11,884 UK teenagers, applying Ec1 with refinement to e-BH rejects four outcome–treatment pairs: internet use with low self-esteem, internet use with depressive symptoms, internet use with high peer problems, and social media use with depressive symptoms, with estimated odds ratios between roughly 2 and 4. The paper interprets these as practically relevant association
What carries the argument
The central object is the mixture p-to-e calibrator Ec1(π) = (1 − π + π log π) / [π (log π)²], which maps a p-value π into an e-value—a test statistic whose expectation under the null is at most 1. The calibrator does the power work by blending a family of simple calibrator functions. With e-values in hand, the e-BH procedure rejects hypotheses with the largest e-values, up to a threshold that uses the order statistics of the e-values, and this controls the FDR without requiring independence, positive dependence, or estimation of nuisance parameters.
Load-bearing premise
The FDR guarantee rests on the calibrated e-values being true e-values, which holds only if the likelihood-ratio p-values fed into the calibrator are super-uniform under the null—an assumption the paper verifies by simulation for n ≥ 100 but does not prove for finite samples in non-Gaussian models.
What would settle it
Run the paper's logistic-regression setting with n = 50 and a rare binary outcome under the null, applying Ec1 to the asymptotic likelihood-ratio p-value; if the fraction of rejected hypotheses at level α exceeds α over many replications, the p-value is anti-conservative and the e-variable guarantee fails in that regime.
If this is right
- Multiverse analyses can achieve finite-sample FDR control without assuming independence or positive dependence among the multiple tests.
- The mixture calibrator Ec1 is preferable to universal and soft-rank e-values for well-specified GLM-based multiverse testing.
- E-values remain a conservative tool: statistical power may be moderate unless sample sizes and effect sizes are sufficiently large.
- The teenager well-being application indicates that averaging technology effects across treatments, as specification-curve analysis does, can hide strong and practically relevant associations.
Where Pith is reading between the lines
- The paper does not say this, but the same calibration recipe would extend to permutation-based p-values, which are super-uniform by construction and would remove the asymptotic caveat outside Gaussian models.
- Because e-BH is dependence-robust, the framework could be used to compare results across different control-covariate subsets without modeling their correlation; whether the four findings survive such specification choices remains open.
- E-values compose across independent datasets, so the approach could be extended to anytime-valid pooling of subpopulation or cohort analyses, which would be useful for ongoing surveillance of technology–wellbeing associations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes using e-values to control the false discovery rate in GLM-based multiverse analyses. It reviews three e-value constructions: universal mixture e-variables (Eq. 3), soft-rank e-variables (Eq. 5), and p-to-e calibrators (Eqs. 7-8), proves the validity of the universal mixture, and compares power via simulation. Based on average log e-values in linear and logistic regression, it recommends the mixture calibrator Ec1 and applies it with e-BH to the Millennium Cohort Study teenage well-being data, reporting four significant outcome-treatment associations. The central claim is that simple p-to-e calibration dominates generic e-value strategies in this setting.
Significance. If established, the paper offers a practical, dependence-robust FDR control method for multiverse analyses, avoiding independence or positive-dependence assumptions and nuisance parameter estimation. Strengths include a correct e-variable validity proof in Section 3.1, explicit discussion of the finite-sample caveat, and reproducible code/data availability. The substantive application is a useful counterpoint to specification-curve averaging. However, the empirical support for the headline power comparison is thin, and the finite-sample validity of the recommended calibrator is not verified at the level needed for strict FDR guarantees.
major comments (4)
- [Section 3.2, Figures 1-2, Abstract] The statement that p-to-e calibration 'significantly outperforms' the other approaches in statistical power is not supported by the reported evidence. The figures plot average log e-values from 100 repetitions with no error bars, confidence intervals, or paired comparisons, and an average log e-value is not a power curve. There are no simulations of e-BH FDR control in a multiverse setting with multiple correlated outcomes/treatments, which is the paper's stated target. Please add Monte Carlo intervals, formal comparisons, and FDR/power simulations with realistic multiverse dimensions, or weaken the claim.
- [Section 3.1, Eqs. (7)-(8), Table 1] p-to-e calibrators are e-variables only when input p-values are super-uniform under the null. For logistic regression, the LRT p-values are asymptotic only; the paper acknowledges this, yet relies on the e-BH finite-sample FDR guarantee. Table 1 reports only rejection proportions at alpha=0.01 and 0.05, which does not validate E[Ec1(P)] <= 1. Since Ec1(p) ~ 1/(p log^2 p) as p goes to 0, its expectation is sensitive to small-p behavior even if tail rejection rates are controlled. This is load-bearing for Section 4. Please prove or empirically validate finite-sample super-uniformity (e.g., report null means of calibrated e-values with Monte Carlo intervals), use exact/permutation p-values, or label the FDR control as asymptotic.
- [Section 2, Eq. (2)] The displayed null H_lj includes 'beta_j' != 0 for j' != j', contradicting the following sentence that no restrictions are placed on remaining treatment effects. The global H_l (all effects nonzero) is not the relevant alternative for a single test H_lj. This affects the definition of q_l in Eq. (3) and the interpretation of all simulation and application results. Please clarify the hypothesis space and the alternative under which q_l is specified.
- [Section 4, Table 2] The sentence 'there were six hypotheses for which 1/Ec1 > 20' appears inverted; the individual-test threshold at level 0.05 is Ec1 > 20 (equivalently 1/Ec1 < 0.05). In addition, the application's four rejections disappear under universal and soft-rank e-values; this is informative, but it means the headline result hinges entirely on the asymptotic validity of Ec1 and should be presented together with the caveat in the second major comment.
minor comments (6)
- [Abstract, Section 1] Typos: 'self-steem' should be 'self-esteem'; 'sensitvy' should be 'sensitivity'.
- [Section 3.1] The inverse-gamma notation IG(a/2,l/2) is confusing because l is also used to index outcomes. Use a different symbol such as b/2.
- [Table 1] The caption and text do not specify the null coefficient values used for the type I error rows. State explicitly which beta* are set to zero.
- [Figures 1-2] Captions say 'average over 100 repetitions' but no uncertainty is displayed. Add error bars or confidence intervals, or state that none are shown.
- [Section 3.1] The closed-form marginal likelihood uses inconsistent notation for sigma^2_l and sigma-hat^2_l. Clarify the parameterization of the inverse-gamma prior.
- [Section 4] The claim that odds ratios in [2,4] are 'practically relevant' could use a sentence on effect-size calibration or external benchmarks, since practical relevance is a subjective judgment.
Circularity Check
No significant circularity; the only self-citation ([2]) is minor and not load-bearing.
full rationale
The paper's central derivation is self-contained. Section 3.1 proves the universal mixture e-variable property in-line (E[E^um] <= 1 follows from the pointwise bound p(y) <= \hat p_lj(y) and \int q_l = 1). The p-to-e calibrators (7)-(8) are standard transforms whose e-variable property is inherited from super-uniformity of the input p-values, with the finite-n caveat for LRT p-values stated explicitly in Section 3.1. Soft-rank e-values are justified by exchangeability ([18]) and e-BH FDR control by [5]; neither is imported from the authors' prior work. The simulation comparison of average log e-values and type I error rates is an empirical evaluation, not a fitted quantity renamed as a prediction. The MCS application applies the pre-specified calibrator to LRT p-values and reports e-BH rejections; no parameter fitted to the data is called a prediction. The only self-citation is [2], co-authored by Rossell, used for data preparation and for motivating the critique of SCA; that citation is not the load-bearing step for any e-value proof or for the comparative power claim. The paper's explicit finite-sample caveat ('Since such a p-value is not exact for finite n, neither is the resulting calibrated e-value') is a validity limitation to verify in applications, not a circularity. Hence score 2 only for the minor, non-load-bearing self-citation.
Axiom & Free-Parameter Ledger
free parameters (2)
- Simulation design: effect-size grid, sample sizes, covariate count, pairwise covariance =
beta* in {0, 0.15, 0.3, 0.5}; n in 100-10000; 10 covariates; pairwise covariance 0.5
- Inverse-gamma hyperparameters for sigma^2 in the universal mixture e-variable =
not stated
axioms (5)
- domain assumption The GLM in Equation (1) with link function F is correctly specified for every outcome-treatment-subgroup combination, and the weighted sum-to-zero coding makes beta_j^(l) the average treatment effect.
- domain assumption For p-to-e calibrators, the likelihood-ratio p-values are super-uniform under the null, or the chi-square approximation is adequate at the sample sizes used.
- standard math The e-BH procedure of Wang and Ramdas (2022) controls FDR at alpha K0/K <= alpha under arbitrary dependence.
- ad hoc to paper Zellner's unit information prior and the inverse-gamma prior on sigma^2 yield a proper marginal likelihood for the universal mixture e-variable.
- domain assumption The permutation scheme (regressing the treatment on other covariates and permuting residuals) makes T0 exchangeable with T1,...,TB under the null.
read the original abstract
Multiverse analysis refers to a common situation where one wishes to assess the association between multiple possible treatment definitions and multiple possible outcome definitions, potentially within multiple sub-populations, among other possible analysis specifications. Multiverse analysis is a useful exploratory tool to assess heterogeneity across the considered specifications, but it is sometimes also used to assess statistical significance. In the latter case, it is critical to acknowledge that multiple comparisons are being performed, and to ensure a valid statistical control of false positive findings. We study the use of e-values within generalized linear models as a tool to control the false discovery rate regardless of the dependence structure of the multiple analyses being performed, while accounting for confounding covariates. We compare the performance of several approaches: universal e-values, soft-rank e-values, and p-to-e calibration. We find that, for problem characteristics typically encountered in multiverse analyses, p-to-e calibration significantly outperforms the other two approaches in terms of statistical power, but said power may be moderate unless the sample size or effect sizes are large enough. An application studying the association between teenager technology use and mental well-being reveals association between depression, low self-steem and peer problems with internet and social media usage.
Reference graph
Works this paper leans on
-
[1]
Nature human behaviour4(11), 1208–1214 (2020)
Simonsohn, U., Simmons, J.P., Nelson, L.D.: Specification curve analysis. Nature human behaviour4(11), 1208–1214 (2020)
2020
-
[2]
Journal of the Royal Statistical Society C71(5), 1330–1355 (2022)
Semken, C., Rossell, D.: Specification analysis for technology use and teenager well-being: Statistical validity and a bayesian proposal. Journal of the Royal Statistical Society C71(5), 1330–1355 (2022)
2022
-
[3]
Nature human behaviour3(2), 173–182 (2019)
Orben, A., Przybylski, A.K.: The association between adolescent well-being and digital technology use. Nature human behaviour3(2), 173–182 (2019)
2019
-
[4]
Journal of the american statistical association90(430), 773–795 (1995)
Kass, R.E., Raftery, A.E.: Bayes factors. Journal of the american statistical association90(430), 773–795 (1995)
1995
-
[5]
Journal of the Royal Statistical Society B84(3), 822–852 (2022)
Wang, R., Ramdas, A.: False discovery rate control with e-values. Journal of the Royal Statistical Society B84(3), 822–852 (2022)
2022
-
[6]
The Annals of Statistics49(3), 1736–1754 (2021) 16
Vovk, V., Wang, R.: E-values: Calibration, combination and applications. The Annals of Statistics49(3), 1736–1754 (2021) 16
2021
-
[7]
Journal of the Royal Statistical Society B86(5), 1091–1128 (2024)
Gr¨ unwald, P., Heide, R., Koolen, W.: Safe testing. Journal of the Royal Statistical Society B86(5), 1091–1128 (2024)
2024
-
[8]
Chugg, B., Ramdas, A., Gr¨ unwald, P.: E-values as statistical evidence: A com- parison to bayes factors, likelihoods, and p-values. arXiv2603.24421, 1–34 (2026)
arXiv 2026
-
[9]
Foundations and Trends®in Statistics1(1-2), 1–390 (2025)
Ramdas, A., Wang, R.: Hypothesis testing with e-values. Foundations and Trends®in Statistics1(1-2), 1–390 (2025)
2025
-
[10]
Journal of the Royal Statistical Society B 57(1), 289–300 (1995)
Benjamini, Y., Hochberg, Y.: Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society B 57(1), 289–300 (1995)
1995
-
[11]
The Annals of Statistics29, 1165–1188 (2001)
Benjamini, Y., Yekutieli, D.: The control of the false discovery rate in multiple testing under dependency. The Annals of Statistics29, 1165–1188 (2001)
2001
-
[12]
The Annals of Statistics31, 2013–2035 (2003)
Storey, J.D.: The positive false discovery rate: a Bayesian interpretation and the q-value. The Annals of Statistics31, 2013–2035 (2003)
2013
-
[13]
Journal of the Americal Statistical Association96, 1151– 1160 (2001)
Efron, B., Tibshirani, R., Storey, J.D., Tusher, V.: Empirical Bayes analysis of a microarray experiment. Journal of the Americal Statistical Association96, 1151– 1160 (2001)
2001
-
[14]
Journal of the Royal Statistical Society B69, 347–368 (2007)
Storey, J.D.: The optimal discovery procedure: A new approach to simultaneous significance testing. Journal of the Royal Statistical Society B69, 347–368 (2007)
2007
-
[15]
Journal of the Royal Statistical Society, Series B71, 905–925 (2009)
Guindani, M., M¨ uller, P., Zhang, S.: A Bayesian discovery procedure. Journal of the Royal Statistical Society, Series B71, 905–925 (2009)
2009
-
[16]
Efron, B.: Large-scale Inference: Empirical Bayes Methods for Estimation, Test- ing, and Prediction vol. 1. Cambridge University Press, Cambridge, United Kingdom (2012)
2012
-
[17]
Proceedings of the National Academy of Sciences117(29), 16880–16890 (2020)
Wasserman, L., Ramdas, A., Balakrishnan, S.: Universal inference. Proceedings of the National Academy of Sciences117(29), 16880–16890 (2020)
2020
-
[18]
Biometrika111(4), 1405–1412 (2024) https://doi.org/10
Koning, N.W.: More power by using fewer permutations. Biometrika111(4), 1405–1412 (2024) https://doi.org/10. 1093/biomet/asae031 https://academic.oup.com/biomet/article- pdf/111/4/1405/60783709/asae031.pdf
2024
-
[19]
Statistical Science41(1), 121–142 (2026) https://doi.org/10.1214/24-STS952
Ramdas, A., Manole, T.: Randomized and Exchangeable Improvements of Markov’s, Chebyshev’s and Chernoff’s Inequalities. Statistical Science41(1), 121–142 (2026) https://doi.org/10.1214/24-STS952
-
[20]
In: Bayesian Inference and Decision Techniques: Essays 17 in Honor of Bruno de Finetti
Zellner, A.: On assessing prior distributions and bayesian regression analysis with g-prior distributions. In: Bayesian Inference and Decision Techniques: Essays 17 in Honor of Bruno de Finetti. North-Holland/Elsevier, Amsterdam; New York (1986)
1986
-
[21]
The Annals of Statistics6, 461–464 (1978)
Schwarz, G.: Estimating the dimension of a model. The Annals of Statistics6, 461–464 (1978)
1978
-
[22]
Bayesian and likelihood methods in statistics and econometrics7, 473–488 (1990)
Kass, R.E., Tierney, L., Kadane, J.B.: The validity of posterior expansions based on Laplace’s method. Bayesian and likelihood methods in statistics and econometrics7, 473–488 (1990)
1990
-
[23]
Statistical Science38(1), 13–29 (2021)
Rossell, D., Rubio, F.J.: Additive Bayesian variable selection under censoring and misspecification. Statistical Science38(1), 13–29 (2021)
2021
-
[24]
Shafer, G., Shen, A., Vereshchagin, N., Vovk, V.: Test Martingales, Bayes Factors and p-Values. Statistical Science26(1), 84–101 (2011) https://doi.org/10.1214/ 10-STS347 18 Depressed Low self-esteem High total difficulties High emotional problems High conduct problems High hyperactivity/inattention High peer problems Low pro-sociality Fig. 3Estimated tre...
2011
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.