{"id":"2c644bba-f730-41d7-bb83-4ba2a0f1f6d9","arxiv_id":"2607.03225","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"PIMAX delivers asymptotically valid multiverse inference for clustered data by embedding flip2sss cluster scores inside PIMA closed testing, without specifying random-effects covariances.","lead":"PIMAX combines multiverse post-selection inference with cluster-level sign-flipping scores so researchers can test many fixed-effects models on longitudinal or hierarchical data without fixing a random-effects covariance. It gives a global test, discovery counts, and adjusted p-values while avoiding a common source of type-I inflation.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The strongest claim is Theorem 2 (asymptotic validity of the global sign-flipping test under Assumptions 1–3) together with the closed-testing consequences for strong FWER and true-discovery bounds. That claim rests squarely on the second-stage mean correctness stated in Assumption 1 and on the cluster-level independence and Lindeberg conditions of Assumptions 2–3. All three are explicit; the proofs in the appendix reduce the multiverse result to the already-established single-model flip2sss theory via common flips. The simulations and the SHARE illustration supply the expected empirical corroboration without introducing contradictions. Because the reader already isolated the same load-bearing assumption and correctly judged the overall risk low, no adjustment to the ACCEPT verdict is warranted.","tokens_in":17057,"tokens_out":516,"duration_ms":5467,"concrete_test":"Re-run the between-cluster simulation of Section 5 with J=20, nj=10 under deliberate second-stage mean misspecification (omit the between-cluster nuisance wj from model (5) while it remains in the data-generating process). If the empirical type-I error of PIMAX-mean stays inside the 95% Monte-Carlo band around 0.05, the practical robustness of the procedure exceeds the formal assumption; if it exceeds 0.08, Assumption 1 is confirmed as the operative binding constraint.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's identification of Assumption 1 (correct second-stage mean of the cluster summaries) as the weakest link is accurate and already fully disclosed by the paper. Proposition 1 makes the equivalence of H0 and the second-stage null rest explicitly on that assumption; Theorems 1–2 inherit it. The paper further supplies the supporting bias-control condition (Proposition 2) and states the asymptotic-in-J character of the guarantees. Simulations (Figures 1–2) show that, under the designs examined, type-I error stays inside the Monte-Carlo band while GLMM+Holm inflates. No hidden inconsistency, circularity, or unstated premise undermines the central claim that the common-flip construction yields asymptotically valid weak and strong FWER control across a multiverse of fixed-effects specifications without a fully specified random-effects covariance. The remaining practical caveats (small-J behavior, first-stage stability) are already flagged by the authors and do not invalidate the asymptotic argument.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes PIMAX, which embeds the flip2sss two-stage cluster-summary construction inside the PIMA multiverse/post-selection framework. For a multiverse of fixed-effects specifications on clustered data, it uses common cluster-level sign-flips of standardized scores to obtain (i) a global test of the intersection null with weak FWER control, (ii) simultaneous lower confidence bounds on the number of true discoveries, and (iii) multiplicity-adjusted p-values with strong FWER control via closed testing (or the maxT shortcut). Validity is asymptotic in the number of clusters under Assumptions 1–3 (correct second-stage mean of the summaries, cluster independence, and Lindeberg-type regularity). Simulations for binary outcomes show type-I control for PIMAX while GLMM+Holm inflates, and a SHARE application illustrates the procedure.","tokens_in":17279,"tokens_out":1001,"duration_ms":8135,"significance":"If the asymptotic claims hold, PIMAX fills a genuine gap: multiverse/post-selection inference for clustered data without committing to a fully specified random-effects covariance. That is practically important, because random-effects misspecification is a well-known source of type-I inflation in GLMMs, and multiverse analysis is increasingly used in the social and biomedical sciences. Strengths include short, transparent proofs that reduce to established sign-flipping results (Appendix A), an explicit bias-control condition (Proposition 2), reproducible code, and simulations that cleanly separate validity from power. The contribution is a careful synthesis rather than a wholly new theory, but the synthesis is load-bearing for the intended applications.","major_comments":[{"comment":"Assumption 1 (correct conditional mean of the first-stage summary tj under the second-stage working model (5)) is the load-bearing premise for Proposition 1 and thus for Theorems 1–2. The paper states it clearly, but the manuscript would be stronger if §5 included at least one design in which the second-stage mean is mildly misspecified (e.g., omitted between-cluster covariate or wrong within/between coding of the target) so that readers can see how type-I error degrades. Without that, the practical scope of the asymptotic guarantee remains hard to judge from the current figures alone.","section":null},{"comment":"The simulations (Figures 1–4) use J ∈ {20,30,40} and balanced nj ∈ {10,20} with a single random-effects structure. The abstract and introduction emphasize unbalanced designs and heteroscedasticity; those features are not exercised in the Monte Carlo study. A short additional panel or appendix table with unbalanced nj and/or cluster-level variance heterogeneity would make the empirical support match the claimed robustness more closely.","section":null}],"minor_comments":[{"comment":"Figure 4 omits GLMM entirely (correctly, given type-I failure), but the caption and surrounding text could state more explicitly that power is reported only for methods that control type I, to avoid a casual reader comparing apples to oranges with Figures 3.","section":null},{"comment":"In Definition 5 and the subsequent global test, the two-sided critical-value indexing uses ⌈αB/2⌉ and ⌈(1−α/2)B⌉; a one-line remark that B must be large enough for these order statistics to be well-defined (already implied by B ≥ 1/α) would help implementers.","section":null},{"comment":"Table 1 is helpful; a parallel one-line reminder in the text of §3.1 that a is a fixed contrast (not estimated) would reduce any ambiguity about the first-stage reduction.","section":null},{"comment":"SHARE analysis: the multiverse of 48 models / 408 tests is fine for illustration, but a sentence on how the maxT shortcut scales for larger K would be useful for readers planning bigger multiverses.","section":null},{"comment":"Minor typos: “eH0” vs “˜H0” notation in the appendix proofs; “nJ” axis labels in Figures 1–4 should be “nj” for consistency with the text.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a clean, well-executed synthesis of two recent methods by overlapping author groups. Novelty is real but incremental; for a methods journal that values usable, theoretically grounded tools for multiverse analysis, that is sufficient. I see no circularity or hidden premise beyond the already-disclosed Assumption 1. Minor revision is appropriate; I would not send it back for major rework."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a practical methods paper that does what it claims. Vesely and Andreella take PIMA (multiverse closed testing via common sign-flips) and flip2sss (cluster-level score summaries that dodge random-effects misspecification) and glue them into PIMAX. The new pieces are the cluster-level closed-testing guarantees, the two short propositions that link the original null to the second-stage working model, and the finite-sample evidence that the combination keeps type I inside the Monte-Carlo band while GLMM+Holm inflates, especially in the between-cluster binary designs.\n\nWhat works: the construction is transparent. Common flips preserve dependence across the multiverse; maxT and mean combiners give both strong FWER p-values and simultaneous lower bounds on true discoveries. They avoid specifying the random-effects covariance entirely, which is the whole point for applied multilevel work where that structure is rarely unique or stable. Proofs in the appendix are short and lean on already-published Lindeberg/sign-flip arguments applied at the cluster level; no circularity. Code and SHARE data are public. The SHARE illustration is modest but honest about how much the chronic-morbidity coefficient can flip sign across plausible codings.\n\nSoft spots, in proportion: everything is asymptotic in the number of clusters J, and Assumption 1 (correct second-stage mean of the cluster summaries) is load-bearing for the equivalence of the original and working nulls. Both are stated clearly, and Proposition 2 gives the bias-control condition. First-stage Firth logistic helps with sparse binary clusters, but very small or sparse nj will still hurt. Free choices (B, combining function) are ordinary for resampling methods and must be fixed a priori. None of this undercuts the central claim.\n\nThis is for people who already run multiverse or multilevel analyses and need error control that does not collapse when the random structure is wrong. It deserves a serious referee; I would send it out and expect light-to-moderate revision on exposition of the assumptions and maybe one more simulation with unbalanced or crossed designs. Worth engaging.","headline":"Clean, usable synthesis that actually solves multiverse inference under clustering without forcing a random-effects covariance; theory is short and rests on known sign-flip results, simulations show the type-I win over GLMM+Holm.","tokens_in":17876,"tokens_out":529,"would_cite":true,"duration_ms":5532,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J12","62F03","62H15"],"pacs":[],"model":"grok-4.5","headline":"PIMAX gives valid multiverse inference for clustered data without needing a fully specified random-effects structure.","keywords":["clustered data","generalized linear mixed models","multiverse analysis","selective inference","sign-flipping score test","two-stage summary statistics","family-wise error rate","post-selection inference"],"falsifier":"Generate clustered binary data with a correctly specified fixed-effects mean but a deliberately misspecified second-stage working model for the cluster summaries, then check whether the empirical type I error of the global PIMAX test stays at the nominal level as the number of clusters grows.","tokens_in":17947,"feed_emoji":"📊","tokens_out":634,"duration_ms":5652,"temperature":0.7,"pith_summary":"When data come in clusters or panels, researchers face two hard problems at once: within-cluster dependence and a long list of defensible fixed-effects specifications. Standard mixed-model tests often inflate type I error if the random-effects covariance is misspecified, and they do not account for the multiplicity of model choices. PIMAX combines a two-stage cluster-summary construction with joint sign-flipping of scores across all candidate models. The result is a single procedure that supplies a global test for the presence of any effect, simultaneous lower bounds on how many specifications support an effect, and multiplicity-adjusted p-values that identify which ones. Because inference never relies on a fully parametric random-effects covariance, the method removes a common source of invalid inference while still letting analysts explore a multiverse of fixed-effects models.","feed_headline":"Multiverse tests for clustered data without random-effects guesswork","feed_subtitle":"PIMAX controls error across model choices while sidestepping misspecified mixed-model covariances","key_machinery":"Shared cluster-level sign flips applied to standardized second-stage score contributions, combined by a non-decreasing function (mean or max) and embedded in closed testing: the common flips preserve dependence across the multiverse while the two-stage reduction moves dependence from the observation level to the cluster level.","core_discovery":"Under mild regularity conditions on cluster independence, second-stage mean structure, and score moments, the joint sign-flipping test of the combined standardized scores is asymptotically valid for the intersection null that the effect is zero in every candidate specification; closed testing on the same flips then yields strong FWER control and simultaneous lower confidence bounds on the number of true discoveries, all without specifying a random-effects covariance.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["PIMAX brings multiverse inference to clustered data without covariance guesswork","Sign-flipping multiverse tests for mixed models that skip random-effects specs","Post-selection inference across model choices in hierarchical data via PIMAX","Valid FWER control for multiverse analysis of clustered observations","Joint sign-flips yield true-discovery bounds without random-effects structure"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The second-stage working model must correctly capture the mean of the cluster-level summary statistics; if that mean structure is wrong, the test no longer targets the original effect of interest.","fun_headline_variants_meta":{"raw":{"variants":["PIMAX brings multiverse inference to clustered data without covariance guesswork","Sign-flipping multiverse tests for mixed models that skip random-effects specs","Post-selection inference across model choices in hierarchical data via PIMAX","Valid FWER control for multiverse analysis of clustered observations","Joint sign-flips yield true-discovery bounds without random-effects structure"]},"model":"grok-4.5","effort":"low","cost_usd":0.0038,"raw_usage":{"total_tokens":1230,"prompt_tokens":799,"num_sources_used":0,"completion_tokens":78,"cost_in_usd_ticks":38000000,"prompt_tokens_details":{"text_tokens":799,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":353,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":799,"tokens_out":78,"duration_ms":3104,"temperature":1.0,"reasoning_tokens":353,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T03:59:29.817395+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Generate clustered binary data with a correctly specified fixed-effects mean but a deliberately misspecified second-stage working model for the cluster summaries, then check whether the empirical type I error of the global PIMAX test stays at the nominal level as the number of clusters grows.","supporting_citations":[],"review_version":1}