REVIEW 2 major objections 5 minor 47 references
Post-Selection Inference for Multiverse Analysis in Mixed-Effects Models (PIMAX)
T0 review · 2 major / 5 minor · reviewed 2026-07-12 · grok-4.5
Pith's one-line read PIMAX gives valid multiverse inference for clustered data without needing a fully specified random-effects structure.
desk verdict Clean, usable synthesis that actually solves multiverse inference under clustering without forcing a random-effects covariance; theory is short and rests on known sign-flip results, simulations show the type-I win over GLMM+Holm. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Shared cluster-level sign flips applied to standardized second-stage score contributions, combined by a non-decreasing function (mean or max) and embedded in closed testing: the common flips preserve dependence across the multiverse while the two-stage reduction moves dependence from the observation level to the cluster level.
What would settle it
Generate clustered binary data with a correctly specified fixed-effects mean but a deliberately misspecified second-stage working model for the cluster summaries, then check whether the empirical type I error of the global PIMAX test stays at the nominal level as the number of clusters grows.
Extended reading notes
Core claim
Under mild regularity conditions on cluster independence, second-stage mean structure, and score moments, the joint sign-flipping test of the combined standardized scores is asymptotically valid for the intersection null that the effect is zero in every candidate specification; closed testing on the same flips then yields strong FWER control and simultaneous lower confidence bounds on the number of true discoveries, all without specifying a random-effects covariance.
Load-bearing premise
The second-stage working model must correctly capture the mean of the cluster-level summary statistics; if that mean structure is wrong, the test no longer targets the original effect of interest.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PIMAX, which embeds the flip2sss two-stage cluster-summary construction inside the PIMA multiverse/post-selection framework. For a multiverse of fixed-effects specifications on clustered data, it uses common cluster-level sign-flips of standardized scores to obtain (i) a global test of the intersection null with weak FWER control, (ii) simultaneous lower confidence bounds on the number of true discoveries, and (iii) multiplicity-adjusted p-values with strong FWER control via closed testing (or the maxT shortcut). Validity is asymptotic in the number of clusters under Assumptions 1–3 (correct second-stage mean of the summaries, cluster independence, and Lindeberg-type regularity). Simulations for binary outcomes show type-I control for PIMAX while GLMM+Holm inflates, and a SHARE application illustrates the procedure.
Significance. If the asymptotic claims hold, PIMAX fills a genuine gap: multiverse/post-selection inference for clustered data without committing to a fully specified random-effects covariance. That is practically important, because random-effects misspecification is a well-known source of type-I inflation in GLMMs, and multiverse analysis is increasingly used in the social and biomedical sciences. Strengths include short, transparent proofs that reduce to established sign-flipping results (Appendix A), an explicit bias-control condition (Proposition 2), reproducible code, and simulations that cleanly separate validity from power. The contribution is a careful synthesis rather than a wholly new theory, but the synthesis is load-bearing for the intended applications.
major comments (2)
- Assumption 1 (correct conditional mean of the first-stage summary tj under the second-stage working model (5)) is the load-bearing premise for Proposition 1 and thus for Theorems 1–2. The paper states it clearly, but the manuscript would be stronger if §5 included at least one design in which the second-stage mean is mildly misspecified (e.g., omitted between-cluster covariate or wrong within/between coding of the target) so that readers can see how type-I error degrades. Without that, the practical scope of the asymptotic guarantee remains hard to judge from the current figures alone.
- The simulations (Figures 1–4) use J ∈ {20,30,40} and balanced nj ∈ {10,20} with a single random-effects structure. The abstract and introduction emphasize unbalanced designs and heteroscedasticity; those features are not exercised in the Monte Carlo study. A short additional panel or appendix table with unbalanced nj and/or cluster-level variance heterogeneity would make the empirical support match the claimed robustness more closely.
minor comments (5)
- Figure 4 omits GLMM entirely (correctly, given type-I failure), but the caption and surrounding text could state more explicitly that power is reported only for methods that control type I, to avoid a casual reader comparing apples to oranges with Figures 3.
- In Definition 5 and the subsequent global test, the two-sided critical-value indexing uses ⌈αB/2⌉ and ⌈(1−α/2)B⌉; a one-line remark that B must be large enough for these order statistics to be well-defined (already implied by B ≥ 1/α) would help implementers.
- Table 1 is helpful; a parallel one-line reminder in the text of §3.1 that a is a fixed contrast (not estimated) would reduce any ambiguity about the first-stage reduction.
- SHARE analysis: the multiverse of 48 models / 408 tests is fine for illustration, but a sentence on how the maxT shortcut scales for larger K would be useful for readers planning bigger multiverses.
- Minor typos: “eH0” vs “˜H0” notation in the appendix proofs; “nJ” axis labels in Figures 1–4 should be “nj” for consistency with the text.
Circularity Check
No significant circularity: PIMAX is a transparent synthesis of flip2sss + PIMA whose validity theorems reduce to independent sign-flipping results applied at cluster level, not to self-defined quantities or fitted predictions.
-
self citation load bearing
[Section 1 (Introduction) and Section 4 (Inference in a multiverse...); also Abstract]
"Sign-flipping score tests ... form the basis of two recent inferential frameworks: post-selection inference in multiverse analysis (PIMA) and the sign-flipping score-based two-stage summary-statistics approach (flip2sss). ... In this paper, we combine these two approaches to develop PIMAX"
The paper’s framing and the name PIMAX rest on the authors’ own prior works (Girardi et al. 2024 with Vesely; Andreella et al. 2025 with Andreella). This is ordinary synthesis self-citation and is not load-bearing for the new validity theorems, which reduce instead to independent results of De Santis et al. (2025a,b). Included only as the single minor self-citation that justifies score 1 rather than 0.
full rationale
The derivation chain is: (i) first-stage cluster summaries + second-stage working model (Defs 1–3, Assump 1) yield equivalence of original H0 and second-stage null (Prop 1); (ii) cluster-wise scores + sign-flips give an asymptotically valid test under Assumps 1–3 (Thm 1, by reduction to De Santis et al. 2025b Thm 2); (iii) common flips across the multiverse + non-decreasing combiner give a valid global test for the intersection null (Thm 2, by reduction to De Santis et al. 2025a Thm 3.2); (iv) closed testing on the same flips supplies strong FWER and simultaneous true-discovery bounds. Props 1–2 and the common-flip construction are new content; the load-bearing asymptotic validity statements are external theorems applied to the new objects. Self-citations to PIMA (Girardi et al. 2024, overlapping author) and flip2sss (Andreella et al. 2025, overlapping author) merely identify the two components being combined; they do not force the multiverse-level claims by construction, nor do they import an unverified uniqueness theorem. No fitted constants are re-labeled as predictions, no ansatz is smuggled, and no known empirical pattern is merely renamed. The paper is therefore self-contained as a methods synthesis; the only residual is ordinary self-citation of the authors’ prior building blocks, which does not raise the score above 1.
Assumptions & free parameters
free parameters (2)
- number of sign-flip transformations B =
1000 / 5000
- combining function ψ =
mean / max
assumptions (5)
- domain assumption Assumption 1: conditional mean of cluster summary tj correctly specified by second-stage model (5)
- domain assumption Assumption 2: clusters are mutually independent
- domain assumption Assumption 3: Lindeberg condition and non-degenerate asymptotic variance of oracle cluster scores
- standard math Closed testing principle yields strong FWER control when local tests are valid
- standard math Sign-flipping score tests remain asymptotically valid under variance misspecification
invented entities (1)
-
PIMAX procedure
Cite this review
Pith. "Pith review of Post-Selection Inference for Multiverse Analysis in Mixed-Effects Models (PIMAX)." pith.science (2026). https://pith.science/paper/WBN6CXWW
@misc{pith2026260703225,
author = {Pith},
title = {Pith review of: Post-Selection Inference for Multiverse Analysis in Mixed-Effects Models (PIMAX)},
year = {2026},
howpublished = {\url{https://pith.science/paper/WBN6CXWW}},
note = {Machine review of arXiv:2607.03225}
}
read the original abstract
Sign-flipping score tests provide robust inference in generalized linear models under variance misspecification and form the basis of two recent inferential frameworks: post-selection inference in multiverse analysis (PIMA) and the sign-flipping score-based two-stage summary-statistics approach (flip2sss). PIMA provides asymptotically valid inference across a multiverse of model specifications, whereas flip2sss extends sign-flipping score testing to longitudinal and hierarchical data through cluster-level summary statistics. In this paper, we combine these two approaches to develop PIMAX, a multiverse inferential framework for clustered observations. The resulting method extends post-selection inference to clustered-data settings, accommodating heteroscedasticity, unbalanced designs, and within-cluster dependence. Given a multiverse of candidate specifications, PIMAX provides a global p-value for testing whether any specification exhibits a non-zero effect (weak control of the family-wise error rate, FWER), lower confidence bounds on the number of true discoveries, and multiplicity-adjusted p-values for identifying the specific contributing specifications (strong FWER control). By avoiding inference based on a fully specified random-effects covariance structure, PIMAX solves a key source of type I error inflation due to random-effects misspecification while enabling inference across a multiverse of fixed-effects specifications.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
two-stage summary statistics
Robust inference for generalized linear mixed models: a “two-stage summary statistics” approach based on score sign flipping , author=. Psychometrika , volume=. 2025 , publisher=
2025
-
[2]
Statistical Science , pages=
Multiple testing for exploratory research , author=. Statistical Science , pages=. 2011 , publisher=
2011
-
[3]
and Young, S
Westfall, Peter H. and Young, S. Stanley , title =. 1993 , isbn =
1993
-
[4]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
Permutation-based true discovery guarantee by sum tests , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2023 , publisher=
2023
-
[5]
Trends in ecology & evolution , volume=
Generalized linear mixed models: a practical guide for ecology and evolution , author=. Trends in ecology & evolution , volume=. 2009 , publisher=
2009
-
[6]
The Annals of Statistics , volume=
Only closed testing procedures are admissible for controlling false discovery proportions , author=. The Annals of Statistics , volume=. 2021 , publisher=
2021
-
[7]
Biometrika , volume=
On the relative efficiency of using summary statistics versus individual-level data in meta-analysis , author=. Biometrika , volume=. 2010 , publisher=
2010
-
[8]
BMJ: British Medical Journal , volume=
Analysis of serial measurements in medical research , author=. BMJ: British Medical Journal , volume=
Show all 47 references
-
[9]
Scientific Meeting of the Italian Statistical Society , pages=
Blockwise Resampling for Robust Fixed Effects Inference in Linear Mixed Models , author=. Scientific Meeting of the Italian Statistical Society , pages=. 2025 , organization=
2025
-
[10]
Biometrika , volume=
Misspecified maximum likelihood estimates and generalised linear mixed models , author=. Biometrika , volume=. 2001 , publisher=
2001
-
[11]
Advances in Cross-National Comparison: A European Working Book for Demographic and Socio-Economic Variables , editor =
-
[12]
2021 , address =
Le condizioni di salute della popolazione anziana in. 2021 , address =
2021
-
[13]
Measuring disability: a systematic review of the validity and reliability of the Global Activity Limitations Indicator (
Van Oyen, Herman and Bogaert, Petronille and Yokota, Renata TC and Berger, Nicolas , journal=. Measuring disability: a systematic review of the validity and reliability of the Global Activity Limitations Indicator (. 2018 , publisher=
2018
-
[14]
Electronic Journal of Statistics , volume=
Permutation-based multiple testing when fitting many generalized linear models , author=. Electronic Journal of Statistics , volume=. 2025 , publisher=
2025
-
[15]
Journal of the American Statistical Association , volume=
Inference in generalized linear models with robustness to misspecified variances , author=. Journal of the American Statistical Association , volume=. 2025 , publisher=
2025
-
[16]
Biometrika , volume=
On closed testing procedures with special reference to ordered analysis of variance , author=. Biometrika , volume=. 1976 , publisher=
1976
-
[17]
2001 , isbn =
Pesarin, Fortunato , title =. 2001 , isbn =
2001
-
[18]
Data resource profile: the
B. Data resource profile: the. International journal of epidemiology , volume=. 2013 , publisher=
2013
-
[19]
Box, George E. P. , title =. Journal of the American Statistical Association , year =
-
[20]
and Clayton, David G
Breslow, Norman E. and Clayton, David G. , title =. Journal of the American Statistical Association , year =
-
[21]
and Searle, Shayle R
McCulloch, Charles E. and Searle, Shayle R. and Neuhaus, John M. , title =. 2008 , isbn =
2008
-
[22]
The best writing on mathematics (Pitici M, ed) , volume=
The statistical crisis in science , author=. The best writing on mathematics (Pitici M, ed) , volume=
-
[23]
American Scientist , year =
Gelman, Andrew and Loken, Eric , title =. American Scientist , year =
-
[24]
Perspectives on Psychological Science , year =
Steegen, Sara and Tuerlinckx, Francis and Gelman, Andrew and Vanpaemel, Wolf , title =. Perspectives on Psychological Science , year =
-
[25]
, title =
Sterling, Theodore D. , title =. Journal of the American Statistical Association , year =
-
[26]
, title =
Greenwald, Anthony G. , title =. Psychological Bulletin , year =
-
[27]
and Berlin, Jesse A
Begg, Colin B. and Berlin, Jesse A. , title =. Journal of the Royal Statistical Society: Series A , year =
-
[28]
and Nelson, Leif D
Simmons, Joseph P. and Nelson, Leif D. and Simonsohn, Uri , title =. Psychological Science , year =
-
[29]
Scientometrics , year =
Fanelli, Daniele , title =. Scientometrics , year =
-
[30]
and Lakens, Dani\"el , title =
Nosek, Brian A. and Lakens, Dani\"el , title =. Social Psychology , year =
-
[31]
2015 , volume =
Estimating the Reproducibility of Psychological Science , journal =. 2015 , volume =
2015
-
[32]
Psychometrika , year =
Girardi, Paolo and Vesely, Anna and Lakens, Dani\"el and Alto\`e, Gianmarco and Pastore, Massimiliano and Calcagn\`i, Antonio and Finos, Livio , title =. Psychometrika , year =
-
[33]
and Finos, Livio , title =
Hemerik, Jesse and Goeman, Jelle J. and Finos, Livio , title =. Journal of the Royal Statistical Society: Series B , year =
-
[34]
Statistics in Medicine , year =
Cnaan, Avital and Laird, Nan and Slasor, Peter , title =. Statistics in Medicine , year =
-
[35]
2022 , url =
pima: Post-selection Inference in Multiverse Analysis , author =. 2022 , url =
2022
-
[36]
Scandinavian Journal of Statistics , pages=
A simple sequentially rejective multiple test procedure , author=. Scandinavian Journal of Statistics , pages=. 1979 , publisher=
1979
-
[37]
Harvard Data Science Review , volume=
Selective inference: The silent killer of replicability , author=. Harvard Data Science Review , volume=. 2020 , publisher=
2020
-
[38]
arXiv preprint arXiv:2604.27907 , year=
Multivariate mixed models with model-free random effects , author=. arXiv preprint arXiv:2604.27907 , year=
-
[39]
The Annals of Statistics , year =
Post hoc confidence bounds on false positives using reference families , author =. The Annals of Statistics , year =
-
[40]
NeuroImage , volume=
Notip: Non-parametric true discovery proportion control for brain imaging , author=. NeuroImage , volume=. 2022 , publisher=
2022
-
[41]
Statistics in Medicine , volume=
Permutation-based true discovery proportions for functional magnetic resonance imaging cluster analysis , author=. Statistics in Medicine , volume=. 2023 , publisher=
2023
-
[42]
Biometrika , pages=
Bias reduction of maximum likelihood estimates , author=. Biometrika , pages=. 1993 , publisher=
1993
-
[43]
and Levy, Roger and Scheepers, Christoph and Tily, Harry J
Barr, Dale J. and Levy, Roger and Scheepers, Christoph and Tily, Harry J. , title =. Journal of Memory and Language , year =
-
[44]
arXiv preprint arXiv:1506.04967 , year =
Bates, Douglas and Kliegl, Reinhold and Vasishth, Shravan and Baayen, Harald , title =. arXiv preprint arXiv:1506.04967 , year =
-
[45]
Journal of Memory and Language , year =
Matuschek, Hannes and Kliegl, Reinhold and Vasishth, Shravan and Baayen, Harald and Bates, Douglas , title =. Journal of Memory and Language , year =
-
[46]
Nature Human Behaviour , volume=
Specification curve analysis , author=. Nature Human Behaviour , volume=. 2020 , publisher=
2020
-
[47]
Statistics in Medicine , volume =
Heinze, Georg , title =. Statistics in Medicine , volume =
Reviewed July 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.