{"id":"c0ef48f0-4f00-47c8-be7c-a19be9793d80","arxiv_id":"2411.13764","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A p-value whose null density is non-decreasing remains valid after arbitrary data-dependent selection, and all standard p-values satisfy this condition.","lead":"This paper introduces selectively dominant p-values, a class of p-values that remain valid for inference even after the hypothesis being tested is chosen based on the data. It shows that most standard p-values have this property, which simplifies and unifies many selective inference procedures.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the central theorems are correct, and the known-selection-function assumption is explicitly scoped rather than a hidden flaw.","rationale":"The reader's weakest_assumption is accurate: the framework requires a known, parameter-free selection function s(x,z). However, this is a scoping condition, not a mathematical gap. The paper explicitly defines the framework for known selection procedures in the introduction and in Definition 1, and all applications provide explicit s or a way to compute it (e.g., the polyhedral lemma in the LASSO example). The central theorems are proved carefully, and the characterization in Theorem 2 is correct. The reader's verdict of ACCEPT with high confidence is consistent with the paper's rigor. My stress-test found no hidden flaw, no circular reasoning, and no unsupported leap in the main derivations. The only mild concern is that the practical feasibility of the method depends on computability of s, but the paper acknowledges this and gives examples where s is simple or simulable. Therefore the verdict should remain UNCHANGED. I partially agree with the reader: the known-s assumption is the natural weak point, but it is not a defect in the argument.","tokens_in":51753,"tokens_out":16864,"duration_ms":162201,"concrete_test":"Numerically verify the central inequality behind Theorem 2's sufficiency direction: for f(x)=x^k with k=0,1,2, all interval selection functions s(x)=1_{x∈[a,b]} on a fine grid, and a fine grid of t, confirm that ∫_0^t s(x)f(x)dx / ∫_0^1 s(x)f(x)dx ≤ ∫_0^t s(x)dx / ∫_0^1 s(x)dx. Separately, re-run the Section 5.1 OSC analysis with a deliberately misspecified selection function (e.g., using s(x)=1_{x≤α'} for α'≠α) to confirm that the corrected p-values no longer control Type I error, demonstrating that the known-s assumption is essential to the method's validity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No significant objection identified. The central claim is that selectively dominant p-values (equivalently, p-values with non-decreasing conditional density under the null) yield valid post-selection inference via the correction psel = ∫_0^p s(x,Z)dx / ∫_0^1 s(x,Z)dx. The proofs of Theorem 1 and Theorem 2 are sound: the key stochastic-order inequality for non-decreasing densities holds for arbitrary selection functions, and the converse construction is valid. The most load-bearing assumption is that the selection function s(x,z) is known and parameter-free, but the paper states this explicitly ('arbitrary, albeit known, selection procedure') and provides several settings (data-carving, LASSO, winner-selection) where s is exactly computable. A limitation of scope is not a flaw in the argument. The examples and corollaries follow from the theorems as claimed, and the corrections to prior work (e.g., Hung and Fithian's tie handling) are well-founded. No critical error or circular step was found.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces selectively dominant p-values, defined as p-values whose conditional density under the null is non-decreasing on [0,1]. It proves (Theorem 2) that this density condition is equivalent to selective dominance: after any selection process with known selection probability s(x,z), the conditional distribution of the p-value given selection stochastically dominates that of a uniform random variable subjected to the same selection. Theorem 1 then shows that the adjusted p-value p_sel = ∫_0^p s(x,Z)dx / ∫_0^1 s(x,Z)dx is uniformly valid post-selection. The theory is applied to inference on winners (Section 3), rank verification in exponential families (Section 4), and combination tests including data carving and publication-bias-aware meta-analysis (Section 5). The paper also corrects a tie-handling error in Hung and Fithian (2019) and gives a counterexample to their claimed level inflation.","tokens_in":51958,"tokens_out":6269,"duration_ms":64075,"significance":"If the results hold, this is a substantial unification of selective inference: instead of deriving conditional procedures case by case, one only needs a valid p-value, a non-decreasing null density, and a known selection function. The appendices provide careful measure-theoretic proofs of the main theorems, a verified counterexample to the rank-verification level claim, and a transparent coupling argument for the data-carving example. The paper also makes its code available. The central assumption—that the selection function is known and parameter-free—is explicitly stated and scoped; it limits applicability but does not affect the internal validity of the theorems. My reading confirms the soundness of the main derivations.","major_comments":[],"minor_comments":[{"comment":"The phrase 'all commonly used p-values' overstates the support provided by Examples 2–6, which cover important parametric families, permutation tests, and F-tests but not a mathematically defined class of 'commonly used' p-values; moreover, Theorem 2 requires a conditional density, so discrete p-values are excluded unless randomized. Suggest qualifying the wording.","section":"Abstract and Section 1.2"},{"comment":"The statement that the selection function is 'always accessible to us via extensive simulations' deserves a caveat: a Monte Carlo estimate of s(x,z) does not by itself provide exact finite-sample Type I error control unless the estimation error is incorporated; the paper currently treats s as known.","section":"Section 2.3, Example 9"},{"comment":"The statement that pO/α remains valid under 'p-hacking' is heuristic, based on modeling null p-hacked p-values as having increasing density on [0,α]; this is a reasonable empirical model but should be labeled as a heuristic rather than a theorem.","section":"Section 5.1"},{"comment":"The LASSO post-selection inference example assumes the noise level σ is known; the paper could note that unknown σ requires additional treatment (e.g., via the square-root LASSO) so that readers do not infer that the framework resolves all LASSO settings without further conditions.","section":"Appendix A.5"},{"comment":"In the displayed re-expression of psel, the denominator is written as q+(Z) + (1/N(Z))(q+(Z) - q(Z)), which has the sign reversed relative to the correct expression q+(Z) + (1/N(Z))(q(Z) - q+(Z)) shown in the main text around equation (23); the surrounding narrative and Example 12 use the correct form, so this appears to be a typographical error.","section":"Appendix A.12"}],"recommendation":"accept","confidential_remarks":"The paper is a strong contribution that delivers on its central claim. The only caveat worth noting to the editor is that the breadth claim 'all commonly used p-values' is informal, but the concrete examples and theorems are correctly scoped. I recommend acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers a genuinely new idea: selectively dominant p-values, characterized by a non-decreasing conditional density under the null, and a single integral correction that makes post-selection inference valid for any known selection function. That is a real step forward. It subsumes the threshold-selection case in Hung and Fithian and extends cleanly to winners, rank verification, and meta-analysis. The correction to Hung and Fithian's tie handling is valid and well demonstrated. The new procedures (top-k Fisher, tie-corrected rank verification, hybrid inference generalized beyond Gaussian) are derived from the theorems, and the proofs in the appendices are careful. I checked the key stochastic-order argument in Theorem 1 and the converse in Theorem 2; both hold. The paper also does something rare: it reports when its own methods give little gain, e.g., hybrid inference over its union-bound competitor, and when the conditional LCB is vacuous for exponential data. That honesty earns trust.\n\nThe soft spots are real but mostly minor. The framework requires the selection function to be known and parameter-free. That is a genuine limitation, not a hidden flaw. The paper says so explicitly and gives several settings where it holds exactly, but the claim of handling 'arbitrary selection' should be read as 'arbitrary, if known and computable.' Also, the assertion that 'all commonly used p-values' are selectively dominant is supported by a list of examples rather than an exhaustive characterization. I would have liked a more precise statement about which families the increasing-density condition covers. Both of these are scope limits, not errors.\n\nThe paper is long and appendix-heavy, but the main theorems are easy to isolate. The code is available, and the empirical claims match the theory. If I worked in selective inference, I would cite this. It deserves a serious referee: the concept is novel, the proofs are sound, and several existing results are sharpened. My recommendation is to send it out for review, with the expectation that the authors will clarify the scope of the selection-function assumption.\n\nBring it to reading group; it will generate good discussion about what post-selection inference is really doing.","headline":"A genuinely new and mostly correct unification of selective inference via p-values, with an honest mapping of its limitations; deserves serious peer review.","tokens_in":52420,"tokens_out":1192,"would_cite":true,"duration_ms":14611,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62J05"],"pacs":[],"model":"deepseek-v4-flash","headline":"A p-value with non-decreasing null density can be adjusted for any known selection procedure by one integral, making selective inference a plug-in operation.","keywords":["selective inference","post-selection inference","p-values","selective dominance","inference on winners","rank verification","data carving","Fisher combination test"],"falsifier":"A decisive check: simulate any standard test statistic under its null, transform to a p-value, and look at the histogram conditional on the auxiliary variable $Z$. If the density decreases on $[0,1]$ for a set of $Z$ with positive probability, Theorem 2 says the p-value is not selectively dominant, so a single such example among commonly used tests would break the paper's blanket claim. A toy calculation shows the failure mode is real: with $p=1-\\sqrt{1-U}$ (a valid decreasing-density p-value) and selection $s(x)=\\mathbf{1}\\{x\\le0.1\\}$, the corrected test rejects with probability $P(p\\le0.01\\mid p\\le0.1)\\approx0.105$, above the nominal $0.1$.","tokens_in":51590,"feed_emoji":"📊","tokens_out":7005,"duration_ms":58957,"temperature":0.7,"pith_summary":"Selective inference is the problem of doing valid tests after the data have already picked the question. This paper proposes a general solution: any p-value whose null density is non-decreasing, which it calls selectively dominant, can be corrected for an arbitrary known selection procedure by a single integral transform. If true, standard p-values—from two-sided parametric tests, one-sided tests in monotone likelihood-ratio and exponential families, F-tests, and permutation tests—can all be used for selective inference, and several known selective methods become simple consequences rather than bespoke derivations. The paper demonstrates this by re-deriving conditional inference on winners, hybrid inference, rank verification, and new selective versions of Fisher's combination test.","feed_headline":"One integral fix makes p-values valid after data snooping","feed_subtitle":"Common p-values survive arbitrary known selection once corrected by the selection function's integral.","key_machinery":"The central object is the selection function $s(x,z)=P(S=1\\mid p=x,Z=z)$, assumed known, which governs how the analyst decides to test a null after seeing the p-value. For a selectively dominant p-value—defined as one whose post-selection distribution stochastically dominates that of a uniform p-value subjected to the same selection—Theorem 1 says the selectively adjusted p-value $p_{\\mathrm{sel}} = \\int_0^p s(x,Z)\\,dx/\\int_0^1 s(x,Z)\\,dx$ is stochastically uniform after selection, so rejecting when $p_{\\mathrm{sel}}\\le\\alpha$ controls selective error. Theorem 2 identifies selective dominance with the condition that the null conditional density of $p$ given $Z$ is non-decreasing, which is verified for UMP and UMPU p-values, permutation p-values, and $F$-test p-values.","core_discovery":"The central discovery is that selective dominance holds exactly when the null conditional density of the p-value $p$ given $Z$ is non-decreasing, and that this condition turns the selection-adjusted p-value $p_{\\mathrm{sel}} = \\int_0^p s(x,Z)\\,dx \\,/\\, \\int_0^1 s(x,Z)\\,dx$ into a valid p-value after selection, where $s(x,z)=P(S=1\\mid p=x,Z=z)$ is the known selection function. The paper shows that all the p-values practitioners commonly use satisfy this condition, and then uses the corrected p-value to give short derivations of conditional and hybrid inference on winners, rank verification in exponential families, data carving, and selective variants of Fisher's combination test. It also shows that the correction is tight when the original p-value is exactly uniform under the null.","pith_inferences":["The paper treats the selection function as known; a natural extension, not pursued in the paper, is a plugin or conservative-envelope version that uses an estimated or upper-bounded $s(x,z)$ and studies how misspecification degrades selective error.","The explicit tie-handling correction in rank verification suggests that other selective methods that ignore ties in discrete or grouped data may be either conservative or anticonservative, and the same $1/N$ tie-breaking adjustment could be ported to them.","Practically, the framework turns common heuristics like 'we only publish $p\\le0.05$' into a transparent correction $p/0.05$ that applies to almost any standard p-value, without deriving a truncated-normal distribution for each new test statistic."],"forward_implications":["A researcher who sees a p-value only when $p\\le\\alpha$ can correct for this publication bias by rejecting when $p\\le\\alpha^2$, for any selectively dominant p-value.","For independent selectively dominant p-values, the winning null is rejected when $p_{(1)}\\le\\alpha p_{(2)}$, and the closed version of this test makes sequential discoveries while controlling family-wise error.","In the Gaussian rank-verification problem, a two-sided rejection at level $\\alpha$ justifies the statement that the winner is strictly larger, while verifying that the winner is at least as large requires no selection correction at all.","Fisher's top-$k$ and truncated combination tests remain valid even when some null p-values are conservative (super-uniform), whereas earlier versions required exact uniformity.","Data carving becomes a general tool: when data fission or thinning makes the conditional distribution of the selection statistic given the p-value parameter-free, the selection function is known and the integral correction can be computed, at least numerically."],"supporting_citations":[{"why":"Supplies the conditional selective-inference recipe that the paper recasts in p-value form, including the winner problem and data carving.","marker":"[Fithian et al., 2017]"},{"why":"Provides hybrid inference on winners, which the paper re-derives, generalizes beyond Gaussian data, and compares to a union-bound competitor.","marker":"[Andrews et al., 2023]"},{"why":"Proposes rank verification in exponential families, whose tie handling the paper corrects and extends.","marker":"[Hung and Fithian, 2019]"},{"why":"Gives the earlier threshold-selection p-value correction for $t$- and $z$-tests, which Example 7 generalizes to all selectively dominant p-values.","marker":"[Hung and Fithian, 2020]"},{"why":"Provides the polyhedral lemma used in Appendix A.5 to map LASSO post-selection inference into the selective-dominance framework.","marker":"[Lee et al., 2016]"},{"why":"Supplies the proof strategy in Appendix B for establishing non-decreasing densities of UMP and UMPU p-values.","marker":"[Lei and Fithian, 2018]"},{"why":"Gives the randomized permutation p-value that is exactly uniform under the null, used as Example 5.","marker":"[Hemerik and Goeman, 2018]"},{"why":"Supplies the 92-pair original-replication psychology dataset used in the publication-bias meta-analysis in Section 5.1.","marker":"[Collaboration, 2015]"},{"why":"Introduces the truncated Fisher combination test that Corollary 6 generalizes to super-uniform p-values.","marker":"[Zaykin et al., 2002]"},{"why":"Provides the UMP and UMPU testing theory that underpins Examples 3 and 4 and the proofs in Appendix B.","marker":"[Lehmann et al., 1986]"}],"fun_headline_variants":["Integral correction makes p-values valid after selection","Common p-values survive arbitrary selection with one fix","A single integral restores p-value validity after snooping","Selective inference gets a simple p-value fix","Data snooping: your p-values are rescued by an integral"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the selection function $s(x,z)=P(S=1\\mid p=x,Z=z)$ is known exactly and depends on no unknown parameters; whenever selection probabilities must be estimated or depend on the effect size being tested, the adjusted p-value cannot be computed from the formula.","fun_headline_variants_meta":{"raw":{"variants":["Integral correction makes p-values valid after selection","Common p-values survive arbitrary selection with one fix","A single integral restores p-value validity after snooping","Selective inference gets a simple p-value fix","Data snooping: your p-values are rescued by an integral"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000697,"raw_usage":{"total_tokens":3132,"prompt_tokens":909,"completion_tokens":2223,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":2147}},"tokens_in":525,"tokens_out":2223,"duration_ms":910997,"temperature":1.0,"reasoning_tokens":2147,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:55:07.840565+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check: simulate any standard test statistic under its null, transform to a p-value, and look at the histogram conditional on the auxiliary variable $Z$. If the density decreases on $[0,1]$ for a set of $Z$ with positive probability, Theorem 2 says the p-value is not selectively dominant, so a single such example among commonly used tests would break the paper's blanket claim. A toy calculation shows the failure mode is real: with $p=1-\\sqrt{1-U}$ (a valid decreasing-density p-value) and selection $s(x)=\\mathbf{1}\\{x\\le0.1\\}$, the corrected test rejects with probability $P(p\\le0.01\\mid p\\le0.1)\\approx0.105$, above the nominal $0.1$.","supporting_citations":[{"cited_title":"Rank verification for exponential families","cited_arxiv_id":null,"evidence_quote":"Proposes rank verification in exponential families, whose tie handling the paper corrects and extends."},{"cited_title":"Exact post-selection inference, with application to the lasso","cited_arxiv_id":null,"evidence_quote":"Provides the polyhedral lemma used in Appendix A.5 to map LASSO post-selection inference into the selective-dominance framework."},{"cited_title":"AdaPT: An Interactive Procedure for Multiple Testing with Side Information","cited_arxiv_id":null,"evidence_quote":"Supplies the proof strategy in Appendix B for establishing non-decreasing densities of UMP and UMPU p-values."},{"cited_title":"Truncated product method for combining p-values","cited_arxiv_id":null,"evidence_quote":"Introduces the truncated Fisher combination test that Corollary 6 generalizes to super-uniform p-values."}],"review_version":1}