{"id":"e582b174-feaf-4b9c-81d9-a4ced2405fa7","arxiv_id":"2411.10647","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"This is a survey of false discovery rate control methods, organized into ranking, FDP estimation, and thresholding steps.","lead":"This paper is a review of methods that control the false discovery rate, the expected share of false positives among rejected hypotheses. It organizes the field into a three-step framework and surveys major procedures from the Benjamini-Hochberg method to Bayesian local FDR and e-value approaches.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3.2 defines the adaptive-BH threshold range as t∈(0,1−λ], which is internally inconsistent and appears to be a typo for t∈(0,λ]; this propagates into the derivation of the mirror-sequence threshold in Section 3.3.","rationale":"I read the paper in good faith as an expository survey whose value depends on the accuracy of its formal descriptions and theorem attributions. Most of the review is consistent with the standard FDR literature: the BH theorem, the PRDS statement, the SC oracle result, the e-BH arbitrary-dependence claim, and the broad three-step framework are all recognizable and correctly cited. The reader's strongest claim, that the paper can be used as an organized and reliable map of the area, is largely supported. However, the formal definition of the Storey-type adaptive BH threshold in Section 3.2 contains a concrete mathematical error that is not merely typographical in effect: it makes the derivation of the mirror-sequence threshold in Section 3.3 self-referential. The correction t∈(0,λ] is small but substantive, because the disjointness of the rejection set {P_i≤t} and the π0-estimation set {P_i≥λ} is exactly what makes the FDP estimator conservative. Since the reader's own weakest assumption flagged the correct representation of Theorem 2 and its supporting calculations, this is a partial agreement rather than a new unrelated objection. A clean revision that fixes the interval and re-verifies the mirror-sequence derivation would restore the survey's reliability; hence I would not reject the paper, but I would make acceptance conditional on that correction.","tokens_in":14240,"tokens_out":31220,"duration_ms":336494,"concrete_test":"Analytic check: verify the required disjointness condition for \\hat{FDP}_λ(t) in Section 3.2. The rejection set is {P_i≤t} and the π0-estimation set is {P_i≥λ}; these are disjoint exactly when t≤λ. Substitute λ=1−t into the printed bound t≤1−λ to obtain t≤t, and into the corrected bound t≤λ to obtain t≤1−t, i.e. t≤0.5, matching Eq. (9). If the authors confirm the printed bound is a typo and the corrected bound restores consistency, the survey's theorem statements are otherwise in line with the literature.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The survey's central claim is that its conceptual framework and descriptions of the latest FDR methodologies are reliable. The formal definition of the Storey-type adaptive BH procedure in Section 3.2 is therefore load-bearing, and it is wrong as printed. For a fixed λ, the estimator \\hat{FDP}_λ(t) = \\hatπ_0(λ)·mt / \\sum_i 1(P_i≤t) is intended to be conservative, with \\hatπ_0(λ) estimated from the upper-tail p-values {P_i ≥ λ}. Conservative validity requires the rejection set {P_i ≤ t} and the estimation set {P_i ≥ λ} to be disjoint, which requires t ≤ λ, not t ≤ 1−λ. The printed range (0,1−λ] is also internally inconsistent: substituting λ=1−t, as Section 3.3 does, gives t∈(0,t], a self-referential bound, whereas the corrected range t≤λ gives t≤1−t ⇔ t≤0.5, which is exactly the range used in Eq. (9). Thus the error propagates into the statement and derivation of Theorem 2, the mirror-sequence procedure attributed to Leung and Sun (2022). A reader using the printed definition with λ<1/2 would search for thresholds up to 1−λ, where the π0-estimation and rejection regions overlap and the conservative interpretation of the FDP estimator no longer applies. This is not merely cosmetic: the printed interval changes the procedure, and the derivation of the headline mirror-sequence threshold is incoherent without the correction.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a review article on false discovery rate (FDR) control in offline multiple testing. It proposes a three-step framework (ranking, FDP estimation, thresholding) and uses it to organize a large literature: the Benjamini-Hochberg procedure and its variants (adaptive/π0 procedures, mirror sequences, weighted p-values), the Sun-Cai local FDR procedure and its side-information extensions, FDR control under dependence (PRDS, weak dependence, general dependence via the BY correction, dBH, and e-values), and pointers to online testing. The paper is explicitly non-exhaustive and directs readers to original sources; its own contribution is an expository conceptual map rather than new theorems.","tokens_in":14593,"tokens_out":14262,"duration_ms":142907,"significance":"If corrected, the survey would be a useful and generally accurate compact map of a large and rapidly evolving literature. Its strengths are the clear three-step organizing principle, the broadly correct statements of the central results (BH, PRDS, e-BH), and the up-to-date coverage of mirror sequences, knockoffs, conformal p-values, and generalized e-values. The paper is also commendably explicit about its selective scope and cites accessible proofs for several key results. However, because the value of the paper is precisely that readers can rely on its formulas and conditions, a correctness error in the adaptive-BH threshold interval and its propagation into the mirror-sequence derivation are not merely cosmetic; they need to be fixed before the review can serve as a dependable reference.","major_comments":[{"comment":"The definition of t_λ_α as sup{t ∈ (0, 1−λ] : \\hat FDP_λ(t) ≤ α} is incorrect as printed and should be sup{t ∈ (0, λ] : ...}. Since \\hat π0(λ) is estimated from the upper tail {P_i ≥ λ}, the conservative interpretation of \\hat FDP_λ(t) requires the rejection set {P_i ≤ t} and the estimation set {P_i ≥ λ} to be disjoint, which is guaranteed by t ≤ λ. The printed range permits t > λ, where the estimator can be anti-conservative. For example, under the global null with m = 2, λ = 0.2, and t = 0.8, the event that both p-values are below 0.5 gives \\hat FDP_λ(t) = 0.5 while the true FDP is 1. This error propagates to Section 3.3: substituting λ = 1 − t into the corrected inequality t ≤ λ yields t ≤ 0.5, which is exactly Eq. (9), whereas substituting into the printed inequality t ≤ 1 − λ gives only the tautology t ≤ t and cannot justify the 0.5 bound. Theorem 2 itself appears correct under the corrected range, but the derivation as printed is incoherent.","section":"Section 3.2 and Section 3.3 (Eq. (7)-(9))"}],"minor_comments":[{"comment":"The sentence 'The ≥ in (6) is sharper when λ is bigger' is unclear and seems backwards, since the lower bound involves P(P_i ≥ λ), which decreases as λ increases; please rephrase or correct.","section":"Section 3.2, Eq. (5)-(6)"},{"comment":"Unlike Theorem 1, Theorem 2 states no explicit assumptions on the p-values; please state the required conditions (e.g., independence of null p-values, or the exact conditions in Leung and Sun 2022) so that the theorem statement is self-contained.","section":"Section 3.3, Theorem 2"},{"comment":"There is a notational slip in the sentence defining the weights: 'independent of P_1, ..., P_n' should be 'independent of P_1, ..., P_m'.","section":"Section 3.4"},{"comment":"The PRDS characterization for the multivariate normal is garbled: it should be stated in terms of the index set I0 (not H0) and, in the standard form, nonnegative covariances Σij ≥ 0 rather than strict positivity.","section":"Section 5.1, Example 1"},{"comment":"Please copyedit for typos such as 'acros s', 'Univeristy', 'is is', 'weighed' in Section 3.4, and 'a overview' in Section 5.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript cites several arXiv preprints by the authors themselves (Banerjee et al. 2023; Fu et al. 2022; Gang et al. 2023a,b). I have no evidence of any problem with these citations, but because the review is intended to be a reliable map of the field, the handling editor may wish to ensure that these preprints are not given disproportionate weight and that their results have been appropriately vetted by the authors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"About arXiv:2411.10647: it's a review of FDR control methods, and it's a good one for what it is. The three-step framework (rank, estimate FDP, threshold) is a genuinely useful organizing device, and the survey covers BH, Storey's null-proportion estimation, mirror sequences, weighted BH, Sun-Cai, dependence (PRDS, dBH, e-values), and online testing with mostly accurate statements and honest citations. No new results, which is expected for a review.\n\nThe one real issue I see is in Section 3.2: the adaptive-BH threshold is defined as sup{t ∈ (0,1−λ] : \\hat{FDP}_λ(t) ≤ α}. That's a typo. For the Storey estimator, the rejection set {p_i ≤ t} and the null-proportion estimation set {p_i ≥ λ} need to be disjoint, which requires t ≤ λ. The printed (0,1−λ] is wrong and will mislead a reader who tries to use the procedure with λ < 1/2. The stress-test note says this propagates into Section 3.3; actually, the mirror-sequence threshold is stated correctly as (0,0.5], so the propagation is more of a sloppy derivation than an incoherent result. Still, it needs fixing.\n\nOther caveats are routine: the coverage is selective (the authors admit this), asymptotic statements are compressed, and some theorems are stated without their regularity conditions. The self-citations are used as literature references, not as support for new claims, so no circularity concern. The theorem statements I checked align with the standard literature.\n\nOverall, I'd send this to referees. It deserves serious refereeing, especially because it is likely to be read by newcomers and practitioners. The fix is minor, so I'd expect a quick acceptance after revision.","headline":"A solid, useful FDR-control review with one clear typo in the adaptive-BH threshold range; publish after a minor fix.","tokens_in":15053,"tokens_out":4463,"would_cite":false,"duration_ms":41021,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H15","62J15","62F03"],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims that the vast majority of offline FDR control methods fit a three-step pattern — rank, estimate the false discovery proportion, then threshold — and organizes the literature around that pattern.","keywords":["false discovery rate","multiple testing","multiple comparisons","type 1 error rate","Benjamini-Hochberg procedure","local false discovery rate","e-value","mirror sequence"],"falsifier":"One concrete test is a literature count: collect the offline FDR-controlling procedures from a fixed set of recent papers and textbooks and check whether each one explicitly ranks hypotheses and defines an estimated FDP function before thresholding. If a substantial share, say a quarter or more, do not, then the 'vast majority' claim in Section 2.1 fails.","tokens_in":14064,"feed_emoji":"📊","tokens_out":10996,"duration_ms":92360,"temperature":0.7,"pith_summary":"This paper is a review of methods that control the false discovery rate (FDR), the expected proportion of false positives among rejected hypotheses. Its central organizing claim is that almost all offline FDR methods fit a three-step pattern: rank the hypotheses with a summary statistic, estimate the false discovery proportion (FDP) at each possible threshold, then reject everything whose statistic falls below the largest threshold with estimated FDP at or below the target level $\\alpha$. The authors demonstrate the pattern on the Benjamini-Hochberg procedure, null-proportion adaptive versions, mirror-sequence and knockoff methods, weighted p-values, the Bayesian Sun-Cai local-FDR procedure, and dependence-robust methods including e-BH. A reader who follows the review gets a common vocabulary for comparing methods, because most innovation in the field happens in either the ranking step or the FDP-estimation step.","feed_headline":"Most FDR control methods fit one three-step recipe","feed_subtitle":"A new survey maps ranking, false-discovery-proportion estimation, and thresholding onto the field's main procedures.","key_machinery":"The load-bearing object is the three-step framework, with the FDP estimators that make it concrete. For any threshold $t$, an FDR method must produce a ranking statistic $T_i$, an estimator $\\widehat{\\mathrm{FDP}}(t)$ for the proportion of false rejections among tests with $T_i \\le t$, and a threshold $t_\\alpha = \\sup\\{t : \\widehat{\\mathrm{FDP}}(t) \\le \\alpha\\}$. The paper supplies explicit estimators that all turn on the same identity: the expected number of null p-values below $t$ divided by the observed number of rejections. The BH estimator $mt/\\sum_i \\mathbf{1}(P_i \\le t)$, the Storey estimator $\\hat\\pi_0(\\lambda) mt / \\sum_i \\mathbf{1}(P_i \\le t)$, the mirror-sequence estimator $(1 + \\sum_i \\mathbf{1}(1-P_i \\le t))/\\sum_i \\mathbf{1}(P_i \\le t)$, and the weighted-BH estimator $\\sum_i w_i t / \\sum_i \\mathbf{1}(P_i/w_i \\le t)$ are all variants of this comparison, as is the Sun-Cai estimator that replaces counts by sums of local FDR values. The framework does the work of unifying frequentist and Bayesian methods under one decision rule.","core_discovery":"The paper's central claim, stated in Section 2.1, is that the vast majority of methodologies developed to control FDR in offline analyses adhere to the three-step framework of ranking, FDP estimation, and thresholding. It then recasts the field's main procedures as instances of that recipe: the BH procedure uses p-values for ranking and the plug-in estimator $\\widehat{\\mathrm{FDP}}(t) = mt/\\sum_i \\mathbf{1}(P_i \\le t)$; Storey's adaptive version multiplies that estimator by an estimated null proportion $\\hat\\pi_0(\\lambda)$; mirror-sequence methods use the null symmetry of $P_i$ and $1-P_i$ to estimate FDP; weighted BH uses weights in both ranking and FDP estimation; and the Sun-Cai procedure replaces p-values by the local false discovery rate $\\mathrm{Lfdr}(X_i) = P(\\theta_i=0 \\mid X_i)$, the posterior probability that a hypothesis is null given its statistic, which the paper argues carries more information about the alternative distribution. The paper also reviews dependence-robust methods, including the Benjamini-Yekutieli correction, conditional calibration, and e-BH, framed as ways to keep FDP estimation valid when independence fails. The claim is that the differences among methods are largely differences in Step 1 and Step 2, not in the overall decision structure.","pith_inferences":["The three-step decomposition is implicitly a design recipe: a methodologist can invent a new FDR procedure by swapping in a better ranking statistic or a sharper FDP estimator while leaving the thresholding rule untouched.","Because the review is explicitly limited to offline testing, a natural extension of the framework is to online FDR procedures; the paper's discussion of e-values and martingales suggests the same ranking-estimation-thresholding logic may carry over in a time-indexed form.","The paper's framing predicts that methods will be compared mainly by which step they improve, so a reader can classify any newly published FDR method by asking whether it contributes a new statistic, a new FDP estimator, or a new thresholding rule.","A testable consequence of the e-value material is that in strongly dependent settings e-BH will typically reject less than BH or adaptive BH; the paper cites comparisons but does not itself run one, so a simulation study could quantify the power gap."],"forward_implications":["Because the BH procedure controls FDR at $\\alpha |H_0|/m$ under independence and PRDS dependence, practitioners can expect it to be conservative when many nulls are true, and can sharpen it by estimating the null proportion.","Ranking by local FDR instead of p-value can improve power, because the local FDR uses information about the alternative distribution; the paper states this is why the Sun-Cai procedure can beat optimal p-value procedures.","Mirror-sequence and mirror-statistic constructions, including knockoffs, control FDR in finite samples by exploiting symmetry under the null, without requiring a user-chosen tuning parameter like $\\lambda$.","Under arbitrary dependence, e-BH controls FDR by turning e-values into super-uniform p-values, and e-values remain valid under optional stopping, making them attractive for sequential and aggregate evidence.","The three-step framework gives practitioners a direct way to compare methods: any new offline procedure can be described by what statistic it ranks on and how it estimates FDP before thresholding."],"supporting_citations":[{"why":"Defines the false discovery rate and introduces the BH procedure that anchors the review's three-step framework.","marker":"Benjamini and Hochberg, 1995"},{"why":"Introduces mFDR and the null-proportion estimation logic behind the adaptive BH procedures in Section 3.2.","marker":"Storey, 2002"},{"why":"Establishes the oracle local-FDR procedure and its adaptive version reviewed in Section 4.","marker":"Sun and Cai, 2007"},{"why":"Supplies the PRDS dependence result and the BY correction for general dependence used in Section 5.","marker":"Benjamini and Yekutieli, 2001"},{"why":"Proves that the e-BH procedure controls FDR under arbitrary dependence among e-values.","marker":"Wang and Ramdas, 2022"},{"why":"Provides the mirror-sequence FDP estimator and the ZAP method for side information used in Sections 3.3 and 4.2.","marker":"Leung and Sun, 2022"},{"why":"Introduces weighted p-values and the weighted BH procedure reviewed in Section 3.4.","marker":"Genovese et al., 2006"},{"why":"Introduces the knockoff mirror-statistic construction cited in Section 3.3.","marker":"Barber and Candès, 2015"},{"why":"Develops conditional FDR calibration, the dependence-adjusted BH procedure in Section 5.3.","marker":"Fithian and Lei, 2022"}],"fun_headline_variants":["Most FDR control methods share a 3-step recipe","Rank, estimate, threshold: the FDR control playbook","Three steps unify most false discovery rate methods","FDR control is a 3-step recipe for most methods"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim stands or falls on the accuracy of the paper's restatements of other people's theorems and on whether its selection of methods fairly represents the offline FDR literature.","fun_headline_variants_meta":{"raw":{"variants":["Most FDR control methods share a 3-step recipe","Rank, estimate, threshold: the FDR control playbook","Three steps unify most false discovery rate methods","FDR control is a 3-step recipe for most methods"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000247,"raw_usage":{"total_tokens":1546,"prompt_tokens":954,"completion_tokens":592,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":525}},"tokens_in":570,"tokens_out":592,"duration_ms":5815,"temperature":1.0,"reasoning_tokens":525,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:27:56.524529+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One concrete test is a literature count: collect the offline FDR-controlling procedures from a fixed set of recent papers and textbooks and check whether each one explicitly ranks hypotheses and defines an estimated FDP function before thresholding. If a substantial share, say a quarter or more, do not, then the 'vast majority' claim in Section 2.1 fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces mFDR and the null-proportion estimation logic behind the adaptive BH procedures in Section 3.2."},{"cited_title":"and Sun, W","cited_arxiv_id":null,"evidence_quote":"Provides the mirror-sequence FDP estimator and the ZAP method for side information used in Sections 3.3 and 4.2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the knockoff mirror-statistic construction cited in Section 3.3."},{"cited_title":"and Lei, L","cited_arxiv_id":null,"evidence_quote":"Develops conditional FDR calibration, the dependence-adjusted BH procedure in Section 5.3."}],"review_version":1}