{"id":"ddf99e94-dfc8-4b0b-aed5-602afc114d16","arxiv_id":"2412.01008","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Generalized universal inference e-values can be plugged into e-BH to control FDR for multiple tests on risk minimizers, with a quantile regression application.","lead":"This paper combines generalized universal inference with the e-BH multiple testing procedure, creating e-values for testing many hypotheses at once when the quantities of interest are minimizers of a risk function. This gives finite-sample error control without distributional assumptions, illustrated on quantile regression with public code.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The implemented learning-rate calibration is not proved to satisfy the strong central condition, so the finite-sample FDR guarantee may not cover the simulations.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing concern: the data-dependent learning-rate selection is not shown to produce e-values, so the theoretical guarantee is conditional in a way that the simulations do not close. I considered whether Theorem 2's uniform-convergence assumption is more central, but for the bounded-data quantile-regression examples the condition is standard and not the main risk. I also checked whether the merging step in (1) could invalidate the e-BH argument; it is the standard weighted average suggested by Vovk and Wang (2021), so that is not the issue. The paper itself flags the reliance on Algorithm 1 in Section 2 and in the simulations, which supports the concern without implying any misrepresentation. A conditional acceptance remains appropriate because the theoretical framework is sound under the stated SCC hypothesis and the empirical FDR at the global null (0.04) is below the nominal level, albeit with only 100 iterations. The proposed Monte Carlo check directly tests whether the implemented calibration preserves the e-value property, which is the condition on which both Theorem 1 and e-BH FDR control rest. If the check passes, the gap is empirical rather than observed; if it fails, the simulations' finite-sample validity claim would need to be weakened or the calibration procedure revised.","tokens_in":7684,"tokens_out":4491,"duration_ms":40660,"concrete_test":"Re-run the paper's Example 1 at Δ=0 (all nulls true) with n=50, M=49, and for a fixed τ (e.g., τ=0.5) compute G_n(Θ0) using the learning rate selected by Algorithm 1 on each replicate. Estimate E[G_n(Θ0)] over at least 10^4 Monte Carlo replicates and construct a standard error. If the estimate exceeds 1 by more than 3 standard errors, the calibrated GUe-value is not an e-value, and Theorem 1 cannot be invoked for the implemented procedure; if it is within Monte Carlo error of 1 for all tested τ, the primary gap is not empirically visible in this setting and the paper should still state the missing condition explicitly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 1 assumes every GUe-value uses a learning rate satisfying the strong central condition (SCC), which makes G_n(Θ0) an e-value and lets e-BH control FDR. In Section 4, however, learning rates are chosen by Algorithm 1 of Dey et al. (2024), a bootstrap calibration procedure that targets coverage of confidence sets, not the SCC inequality E exp[-ω(ℓ(θ;Z)-ℓ(θ*;Z))] ≤ 1. The paper provides no theorem that the calibrated ω satisfies SCC or is independent of the validation split S2, so the resulting G_n(Θ0) need not be an e-value. The FDR guarantee in the e-BH theorem therefore does not automatically cover the implemented method. The authors are transparent about relying on an algorithm that 'empirically maintains frequentist validity,' but the simulations are thin: 100 Monte Carlo iterations, only global-null FDR reported, no partial-null FDR, and no comparison baselines. The uniform-convergence assumption in Theorem 2 is a secondary concern; for the bounded data in the simulations it is quite plausible and not the main gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a multiple-testing method that combines generalized universal inference (GUe) e-values with the e-BH procedure. For hypotheses about minimizers of risk functions, each test uses a GUe-value G_n(Θ0)=inf_{θ∈Θ0} exp[-ω|S2|(R̂_S2(θ̂_S1)-R̂_S2(θ))]. The authors prove in Theorem 1 that when every GUe-value uses a learning rate satisfying the strong central condition and all nulls are true, the averaged e-BH-transformed e-value GM in Eq. (1) satisfies Pr(GM≥α^{-1})≤α; Theorem 2 gives consistency when at least one null is false, under a uniform-convergence condition on the empirical risk. Simulations for quantile regression assess Type II error and report FDR at the global null, with learning rates selected by Algorithm 1 of Dey et al. (2024).","tokens_in":7876,"tokens_out":15916,"duration_ms":147574,"significance":"If the guarantees are extended to the implemented learning-rate calibration, the paper offers a useful finite-sample-error-control framework for multiple testing on risk minimizers without likelihood assumptions, and it identifies quantile regression as a natural application. Feeding GUe-values into e-BH is a straightforward but potentially valuable combination because e-BH controls FDR under arbitrary dependence among e-values. The paper is also transparent about relying on bootstrap calibration for the learning rate; however, the current manuscript does not establish that the simulations are covered by the theorems, so the advertised finite-sample validity is not yet fully supported.","major_comments":[{"comment":"The finite-sample e-value guarantee in Theorem 1 is conditional on the learning rate satisfying the strong central condition and, implicitly in the GUe construction, being chosen independently of the validation sample S2. In the simulations, learning rates are chosen by Algorithm 1 of Dey et al. (2024), but the manuscript neither states whether that algorithm uses only S1 nor proves that the bootstrap-selected ω satisfies the strong central condition or is independent of S2. Consequently, the implemented procedure is not covered by the stated guarantees, and the abstract's claim of finite-sample valid error control is too strong as written. Please either prove a preservation result for the calibrated learning rate, or cleanly separate the validity simulations using fixed ω known to satisfy the strong central condition from the heuristic calibrated-ω results.","section":"Section 4 (Simulations), with Sections 2 and 3"},{"comment":"The empirical FDR evidence is limited to the all-null case: at Δ=0 and Γ=0 the FDR of the unmerged GUe-values is reported as 0.04, based on 100 Monte Carlo iterations. FDR control is a property that must hold under arbitrary mixtures of true and false nulls; the all-null FDR is only the FWER special case and cannot detect inflation caused by the learning-rate calibration in partial-null configurations. Please add partial-null FDR simulations with standard errors, and preferably report the distribution of calibrated learning rates to show that their data dependence does not invalidate the e-value property.","section":"Section 4 (Examples 1 and 2)"},{"comment":"The proof of Theorem 1 is not 'virtually identical' to Theorem 2 of Wang and Ramdas (2022), which is a statement about FDR control of the e-BH procedure, not a probability bound for the particular average in Eq. (1). Because this is the paper's main finite-sample validity result, please include the short derivation; for example, for nonnegative e-values with expectation at most 1, E[Σ_{m=1}^M m G_(m)] ≤ M(M+1)/2, so E[GM] ≤ (M+1)/(2M), and Markov's inequality gives the stated bound.","section":"Theorem 1, proof"}],"minor_comments":[{"comment":"The paper repeatedly claims FDR control, but Theorem 1 is a global Type I error bound for GM and no theorem explicitly states the e-BH FDR guarantee when applied to GUe-values; please state this as a corollary so that the headline claim is directly supported by a displayed result.","section":"Abstract and Introduction"},{"comment":"The statement that 'the strong central condition easily holds because the data are bounded' is too quick: boundedness alone does not imply E exp[-ω(ℓ(θ;Z)-ℓ(θ*;Z))] ≤ 1 for all θ and all ω in [0,ω̄), which also requires identifiability or a suitable separation property for the risk minimizer. Please give the specific argument or cite the relevant condition from Dey et al. (2024).","section":"Section 4 (Examples 1 and 2)"},{"comment":"At 100 Monte Carlo iterations, the standard error of an estimated proportion near 0.05 is about 0.02, so the reported 0.04 FDR should be accompanied by binomial confidence intervals or standard errors; the same applies to the Type II error curves.","section":"Section 4 (Examples 1 and 2)"},{"comment":"The uniform-convergence assumption sup_{ϑ∈Θ_m}|R̂_n^{(m)}(ϑ)-R^{(m)}(ϑ)|=o_p(1) is stated but not verified for the quantile-regression loss; please state compactness and Lipschitz conditions under which it holds, or cite a standard uniform law of large numbers.","section":"Theorem 2"},{"comment":"The superscripts (m) denote sorted GUe-values, but the null sets Θ0^m are then relabeled by the same sorted index; this relabeling should be stated explicitly to avoid ambiguity about which hypothesis corresponds to which entry of the sorted list.","section":"Eq. (1)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a short communication whose main ingredients come from Dey et al. (2024) and Wang and Ramdas (2022); its incremental contribution is the pairing of GUe-values with e-BH and the quantile-regression illustration. I would consider publication if the authors either prove that the calibrated learning rate preserves the e-value property under stated conditions or explicitly weaken the validity claims for the implemented procedure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a clean, competent extension, not a landmark. The authors take their earlier GUe-values from generalized universal inference and plug them into the e-BH procedure for FDR control, with a proof-of-concept on quantile regression. Theorems 1 and 2 are correct, but they are near-immediate corollaries of the cited GUe and e-BH results; the real content is in the application and simulation.\n\nWhat is done well: the writing is clear, the assumptions are stated explicitly, and the authors are transparent that the learning rate is selected by Algorithm 1 from their prior paper, an empirical calibration procedure. They do not oversell Theorem 2; they even call it 'fairly weak' and discuss limitations. The code is public, which helps reproducibility.\n\nThe soft spot is the gap between the theorem's hypothesis and the implementation. Theorem 1 assumes each GUe-value uses a learning rate satisfying the strong central condition. In the simulations, the learning rate comes from a bootstrap calibration that targets coverage, not the SCC inequality. No proof shows the calibrated omega preserves the e-value property. For bounded data the SCC plausibly holds for small omega, and the reported null FDR of 0.04 suggests the calibration does not break things in practice, but the finite-sample guarantee does not formally cover the implemented procedure. That is a genuine limitation, not a deal-breaker, but it should be fixed in revision.\n\nThe simulations are thin: 100 iterations, only global-null FDR, no partial-null FDR, and no comparison to any baseline. The Type II error curves are fine but do not establish much beyond a proof-of-concept.\n\nThe citation pattern is appropriate: the paper leans on the authors' own prior GUe paper and on Wang and Ramdas; that is fair since the results are built directly on those. Self-citation here is not a red flag.\n\nWho this is for: researchers who need finite-sample valid multiple testing for risk minimizers without likelihood assumptions. That is a real niche, and this paper is a usable step. I would want a revision that closes the calibration gap and beefs up the simulations before accepting, but it deserves serious refereeing.","headline":"A workmanlike extension of GUe-values to e-BH, correct under stated assumptions but with a real gap between theorem and implementation.","tokens_in":8429,"tokens_out":3223,"would_cite":false,"duration_ms":28172,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Generalized universal inference e-values can be plugged into the e-BH procedure to control false discovery rates for risk minimizers, with finite-sample validity and no likelihood assumptions.","keywords":["e-value","false discovery rate","generalized universal inference","empirical risk minimization","quantile regression","e-BH procedure","multiple testing","learning rate"],"falsifier":"Run the implemented procedure with bootstrap-selected learning rates under a bounded-loss global null where one can verify a fixed learning rate satisfies the strong central condition, with $M=49$, $n=50$, and $\\alpha=0.1$; if the empirical rejection frequency of $G_M$ exceeds $\\alpha$ by more than Monte Carlo error, Theorem 1 does not cover the data-selected learning rate and the finite-sample claim needs an added argument.","tokens_in":7466,"feed_emoji":"📊","tokens_out":12719,"duration_ms":98811,"temperature":0.7,"pith_summary":"This paper aims to show that e-values built through generalized universal inference (GUe-values) can be plugged into the e-BH multiple-testing procedure, yielding finite-sample control of the false discovery rate when the hypotheses concern the minimizer of a risk function. The central theoretical result is that if each GUe-value uses a learning rate satisfying the strong central condition and all null hypotheses are true, the combined e-value $G_M$ has type I error at most $\\alpha$; if at least one null is false and the empirical risks converge uniformly, the test is consistent. This matters because ordinary e-value constructions require a correctly specified statistical model, whereas quantile regression and other risk-minimization problems do not come with a trustworthy likelihood. The paper demonstrates the procedure on quantile regression, where it checks whether a covariate matters at any of several quantiles, and reports simulation power rising as the signal or sample size grows.","feed_headline":"e-values for risk minimizers control FDR without model assumptions","feed_subtitle":"Pairing GUe-values with e-BH gives finite-sample FDR control for risk minimizers, no likelihood needed.","key_machinery":"The central object is the GUe-value, which converts a risk-minimization problem into an e-value by exponentiating the gap between the validation-set empirical risk at a candidate or null region and at the training-set empirical risk minimizer, scaled by a learning rate $\\omega$. The proof engine is the strong central condition, $\\mathbb{E}\\exp[-\\omega(\\ell(\\theta;Z)-\\ell(\\theta^*;Z))] \\le 1$ for all $\\theta\\in\\Theta$ and $\\omega\\in[0,\\bar\\omega)$, which is exactly the inequality that makes the expectation of each GUe-value at most one. On top of that, the e-BH transformation $e_m^* = m e_m/M$ and the averaged combination $G_M$ turn a set of possibly dependent GUe-values into a single test with the stated error control; in the simulations, the learning rates are selected by the bootstrap-calibration algorithm suggested in the GUe framework.","core_discovery":"The paper's claim is that the GUe-value, $G_n(\\theta) = \\exp[-\\omega |S_2|(\\hat R_{S_2}(\\hat\\theta_1) - \\hat R_{S_2}(\\theta))]$, is a valid e-value (expectation at most one under the null) for a risk-minimizer null whenever the strong central condition holds, and that such e-values are exactly the right input for the e-BH procedure. Sorting the GUe-values and working with $e_m^* = m e_m / M$ therefore controls the false discovery rate under arbitrary dependence. For the meta-hypothesis that none of the individual nulls is false, the paper combines the sorted GUe-values into $G_M = M^{-1}\\sum_{m=1}^M (m/M)\\,G_{(m)}^{n_m}(\\Theta_0^{(m)})$ and proves that under the complete null $\\Pr(G_M \\ge \\alpha^{-1}) \\le \\alpha$, while under uniform convergence of empirical risk and at least one false null the rejection probability tends to 1. The intended application is quantile regression, where the target coefficients are minimizers of expected check loss rather than parameters of a likelihood.","pith_inferences":["A consequence the paper leaves implicit is that the e-BH layer also protects mixtures of true and false nulls at the individual level; the reported FDR is only checked under the global null, so a simulation with a sparse set of false nulls would directly test that part of the claim.","The theorems cover fixed learning rates, but the implementation chooses learning rates from the same data via bootstrap calibration; proving that the selected rate still satisfies the strong central condition would close the gap between the theoretical guarantee and the simulated procedure.","The average-with-weight $(m/M)$ is one valid symmetric merge; other merging functions from the e-value literature could give higher power in particular signal patterns, and the paper does not compare them.","For the full quantile-regression question, a natural extension is to let the number of tested quantiles grow with $n$ or to cover all of $[0,1]$ using the online version of the GUe-value; the paper names this as future work, but it is a direct testable next step."],"forward_implications":["Any ERM-based testing problem can inherit the recipe: if the strong central condition holds for each learning rate, a single combined GUe-value gives a finite-sample $\\alpha$-level test without specifying a likelihood.","The e-BH step controls FDR at level $\\alpha$ even when the individual GUe-values are arbitrarily dependent, so strongly correlated quantile-specific tests do not need an independence assumption.","In quantile regression, the meta-test 'this predictor matters at some quantile' is consistent: whenever the predictor has nonzero coefficient at a tested quantile, the rejection probability of $G_M$ tends to 1 as $n\\to\\infty$.","Because the construction only needs a loss function and the strong central condition, the same multiple-testing scheme applies beyond quantile regression to any target defined as a risk minimizer, such as robust M-estimates or the minimum clinically important difference.","The simulation results indicate that the combined test distinguishes all-null settings, where its type II error is near $1-\\alpha$, from settings with genuine signals, where the type II error drops as the signal strength or sample size grows."],"supporting_citations":[{"why":"Supplies the generalized universal inference construction, the strong central condition, the single-hypothesis validity theorem, and the bootstrap learning-rate algorithm used in the simulations.","marker":"Dey et al. (2024)"},{"why":"Gives the e-BH procedure and the FDR-control theorem for arbitrarily dependent e-values that the paper adapts.","marker":"Wang and Ramdas (2022)"},{"why":"Introduces universal inference via sample splitting, the e-value construction that GUe-values generalize.","marker":"Wasserman et al. (2020)"},{"why":"Justifies averaging as a dominating symmetric merging function for e-values, used to define the combined $G_M$.","marker":"Vovk and Wang (2021)"},{"why":"Defines quantile regression through check-loss minimization, the risk-minimization problem used in the application.","marker":"Koenker and Bassett (1978)"},{"why":"Provides the Gibbs-posterior calibration principle behind the learning-rate selection algorithm used in the simulation study.","marker":"Syring and Martin (2019)"},{"why":"Extends the calibration idea for Gibbs posteriors that the GUe learning-rate algorithm adapts.","marker":"Martin and Syring (2022)"}],"fun_headline_variants":["GUe-values with e-BH tame FDR for risk minimizers","No likelihood needed: FDR control for risk-minimizer tests","Universal inference e-values fix multiple testing without assumptions","Risk-minimizer e-values: multiple testing without model assumptions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on each learning rate being small enough that the strong central condition holds and on that rate being chosen before the validation data are seen, while the simulations instead pick the rate from the data by bootstrap calibration and no theorem shows the selected rate still satisfies the condition.","fun_headline_variants_meta":{"raw":{"variants":["GUe-values with e-BH tame FDR for risk minimizers","No likelihood needed: FDR control for risk-minimizer tests","Universal inference e-values fix multiple testing without assumptions","Risk-minimizer e-values: multiple testing without model assumptions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1290,"prompt_tokens":906,"completion_tokens":384,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":313}},"tokens_in":522,"tokens_out":384,"duration_ms":3812,"temperature":1.0,"reasoning_tokens":313,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:46:51.346499+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the implemented procedure with bootstrap-selected learning rates under a bounded-loss global null where one can verify a fixed learning rate satisfies the strong central condition, with $M=49$, $n=50$, and $\\alpha=0.1$; if the empirical rejection frequency of $G_M$ exceeds $\\alpha$ by more than Monte Carlo error, Theorem 1 does not cover the data-selected learning rate and the finite-sample claim needs an added argument.","supporting_citations":[],"review_version":1}