{"id":"e1c632dd-9cc2-4c6d-ac59-06f8ec70ac8d","arxiv_id":"1908.00583","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The adaptively weighted Fisher method's 0/1 study weights consistently recover the true nonzero-effect studies, and the test has the same asymptotic Bahadur efficiency as Fisher's method.","lead":"This paper analyzes the adaptively weighted Fisher's method for combining p-values across studies, proving that its zero/one study weights converge to the true set of studies with nonzero effects and that the test is asymptotically Bahadur optimal. The results give formal statistical backing to a meta-analysis tool that uses those weights to reveal which studies contribute to a gene's significance.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1 is false as stated when a nonzero-effect study has n_k=o(n); the proof's tail-ratio expansion reverses sign in that regime, so λ_k>0 (or a condition like n_k/log n→∞) must be added.","rationale":"The reader's weakest assumption is exactly the proportional-sample-size condition, and my stress-test confirms that this is the load-bearing point. The concern is sharper than a mere proof gap: for a nonzero study with n_k=o(n), the paper's own tail-ratio algebra gives A_{i+ℓ}→0 rather than →∞, and an explicit log-n counterexample shows the stated Theorem 1 is false. However, the intended theorem is recoverable by adding the standard assumption λ_k>0 for all k (or at least n_k/log n→∞ for every nonzero-effect study), which is natural in meta-analysis settings where all studies grow comparably. The ABO result in Theorem 2 is plausibly correct once Theorem 1 is repaired, and the proof of Theorem 2 is essentially a Bonferroni sandwich around L_obs. I therefore keep the reader's CONDITIONAL verdict rather than moving to REJECT: the central claims are salvageable with a clarified hypothesis and corrected algebra, but the manuscript is not correct as written. The same concern was identified by the reader, so agreement is 'agree'.","tokens_in":10969,"tokens_out":18832,"duration_ms":190959,"concrete_test":"Settle it analytically or by simulation: set K=2 with θ1=θ2=0.5, n1=n-⌊log n⌋, n2=⌊log n⌋, generate the two-sample p-values as in Section 5, and for n=10^3,...,10^7 compute the AW argmin by comparing L(1,0)=1-F_{χ²_2}(-2log p1) with L(1,1)=1-F_{χ²_4}(-2(log p1+log p2)). If P(ŵ=(1,0)) does not tend to 0—indeed tends to 1 because L(1,0)/L(1,1) ∼ n^{θ²/8-1} → 0—then Theorem 1 is false under the stated λ2=0 regime. A positive control with n2=cn (λ2>0) should show P(ŵ=(1,1))→1, isolating the missing assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weak point is the proportional-sample-size assumption in Section 3: the paper only states lim n_k/n = λ_k, never λ_k>0. In the proof of Theorem 1, the tail ratio A_{i+ℓ} is expanded by replacing -2∑log p_j with n C_j. If an added study has a nonzero effect but λ_{i+ℓ}=0, then C_{i+ℓ}=C_i, the exponential factor exp{n/2(C_{i+ℓ}-C_i)} equals 1, and the displayed polynomial factor (n/2)^{-ℓ} drives A_{i+ℓ} to 0—the opposite of the claimed divergence to ∞. The proof also silently drops the λ_j in the exponent exp{n/2 ∑ c_{i+j}}. So the argument does not merely leave the o(n) regime unproved; its own asymptotic expansion predicts exclusion of such studies. The conclusion is actually false there. Take K=2, θ1=θ2=0.5, n1=n-⌊log n⌋, n2=⌊log n⌋. Then -2log p2 ≍ (θ²/4)log n, and the tail ratio L(1,0)/L(1,1) behaves like n^{θ²/8-1}=n^{-0.96875} → 0, so AW-Fisher asymptotically selects weight vector (1,0) instead of w*=(1,1). Thus Theorem 1 as stated is false unless every nonzero-effect study satisfies λ_k>0 (or at least n_k/log n → ∞). Because Theorem 2's proof uses weight consistency to identify the limiting -2/n log L_obs, the ABO claim is not established on the stated basis either.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript studies the adaptively weighted Fisher (AW-Fisher) method for combining p-values from independent studies, where binary weights are used to indicate which studies contribute to the overall signal. The paper claims two main results: Theorem 1, that the estimated AW weight vector converges to the true vector marking exactly the studies with nonzero effects; and Theorem 2, that the AW-Fisher test has the same exact Bahadur slope as Fisher's method and is therefore asymptotically Bahadur optimal. The proofs use exact-slope asymptotics and chi-square tail expansions, and simulations are provided for the convergence rates of the weight estimates.","tokens_in":11282,"tokens_out":7726,"duration_ms":79883,"significance":"If the main results were correct, they would substantially strengthen the theoretical justification of AW-Fisher, a method already used in genomic meta-analysis because its estimated weights provide biological information about study-level heterogeneity. Equally important, the claimed asymptotic Bahadur optimality would place AW-Fisher on par with Fisher's method in terms of asymptotic efficiency while preserving the interpretability advantage. The paper also addresses a practical need, and the simulation setup is natural. However, the central consistency theorem is false as stated, and the proof of the Bahadur optimality theorem contains serious algebraic and logical gaps. These issues need to be resolved before the paper can be considered acceptable.","major_comments":[{"comment":"Theorem 1 is false as stated because the assumption lim n_k/n = lambda_k does not require lambda_k > 0. For a study with nonzero effect but lambda_k = 0, the claimed weight consistency can fail. A concrete counterexample is K = 2, theta_1 = theta_2 = 0.5, n_1 = n - floor(log n), n_2 = floor(log n). Here lambda_1 = 1 and lambda_2 = 0, yet theta_2 is nonzero. Since -2 log p_2 is of order (theta_2^2/4) log n, the tail-ratio L({1,0})/L({1,1}) behaves like n^{theta_2^2/8 - 1} = n^{-0.96875}, which tends to 0, so AW-Fisher asymptotically selects the weight vector (1,0) rather than w* = (1,1). Thus the statement 'w_hat -> w*' requires an additional condition such as lambda_k > 0 for every k with theta_k nonzero, or at least n_k / log n -> infinity for such studies. This is a load-bearing error because it invalidates the stated consistency result.","section":"Section 3, proof of Theorem 1"},{"comment":"The proof of Theorem 1 contains an algebraic error in the tail-ratio expansion. The authors replace -2 sum_{j=1}^k log p_j by n C_k with C_k = sum_{j=1}^k lambda_j c_j(theta), but in the exponential factor they write exp{n/2 (C_{i+ell} - C_i)} and then replace this by exp{n/2 sum_{j=1}^ell c_{i+j}}, dropping the lambda_j factors. The correct exponent should be exp{n/2 sum_{j=1}^ell lambda_{i+j} c_{i+j}}. This matters because if lambda_{i+ell} = 0 for a nonzero-effect study, the exponent vanishes and the polynomial factor (n/2)^{-ell} makes the ratio tend to 0, not infinity. The proof's conclusion that the ratio diverges is therefore not supported by its own expansion, and the claimed convergence rate O(n^ell exp(-n/2 sum c_{i+j})) should be O(n^ell exp(-n/2 sum lambda_{i+j} c_{i+j})) when lambda factors are included.","section":"Section 3, proof of Theorem 1"},{"comment":"The proof of Theorem 2 has a serious algebraic mistake in the displayed chain after the Bonferroni bound. The text writes 'lim -2/n log(P(S > s_obs)) > lim -2/n {L_obs + log(2K-1)} = sum w*_i lambda_i c_i(theta)', but the logarithm must apply to L_obs: the correct inequality is -2/n log(P(S > s_obs)) >= -2/n(log L_obs + log(2K-1)). As written, the expression -2/n{L_obs + log(2K-1)} has the wrong scale and does not converge to the claimed slope. This is not a typographical nuisance; it breaks the derivation of the exact slope. In addition, the lower-bound part invokes the random w_hat where a fixed subset w* is required: because w_hat is a data-dependent minimizer, one cannot simply replace it by w* inside the limit without an independent argument. The consistency of the weights, even if proven, does not directly justify the interchange of the random weight sequence with the deterministic limit in the log-tail expression. Thus the theorem's conclusion c_AW = c_Fisher is not established on the basis given in the manuscript.","section":"Section 4, proof of Theorem 2"}],"minor_comments":[{"comment":"The phrase 'Asymptotic Barhadur optimality' contains a misspelling of Bahadur; the correct spelling is used elsewhere in the paper.","section":"Section 1"},{"comment":"There are repeated typos such as 'statisﬁes' for 'satisfies' and 'eﬀective size' for 'effect size'.","section":"Section 2 and Section 5"},{"comment":"The fitted parameter values in the first simulation are missing: the text reads 'The estimates are â = and b̂ = .' The authors should provide the actual numerical values.","section":"Section 5"},{"comment":"The caption says 'Comparing accuracy of the new approach and permutation approach to obtain the AW p-values,' but the figure actually displays convergence rates of the weight estimates. The caption should be corrected to match the displayed content.","section":"Figure 2 caption"},{"comment":"The submission header states 'Received October 2007. Revised February 2008. Accepted March 2008.' These dates are inconsistent with the arXiv submission date of August 2019 and should be corrected.","section":"Header, page 14"},{"comment":"The notation is confusing: F_chi_k(x) is used as the CDF of the chi distribution, since x appears inside log(1 - F_chi_k(x)) ~ -x^2/2, but the sentence says 'chi^2_k follows chi-square distribution.' Please clarify that the variable x is the square root of a chi-square variable.","section":"Lemma 2"}],"recommendation":"major_revision","confidential_remarks":"The counterexample in my first major comment shows that Theorem 1 is false under the stated assumptions, so the paper cannot be accepted in its present form. However, the error is repairable: adding lambda_k > 0 for all nonzero-effect studies (or equivalently n_k / log n -> infinity) should restore the consistency claim, and the proof needs to keep the lambda_j factors in the exponent. The ABO proof also needs substantial repair. Given that the AW-Fisher method is already used in practice and the asymptotic questions are of genuine interest, I believe a major revision is appropriate rather than outright rejection, provided the authors can supply correct proofs with the added conditions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the paper's new result, consistency of the AW-Fisher weights, is false as stated, and the ABO claim falls with it because it rests on that consistency. The manuscript also has presentation problems that would need a real revision.\n\nWhat's genuinely new: Li et al. (2011) established admissibility of AW-Fisher, not weight consistency. Proving that the 0/1 weights converge to the true support is a sensible and useful contribution, and the proof strategy of comparing chi-square tail ratios is the right idea. The ABO statement is then a fairly direct corollary of consistency plus Littell–Folks, so the independent new content is modest but present.\n\nThe problem is in Theorem 1. The proof assumes lim n_k/n = λ_k but never requires λ_k > 0. The tail-ratio expansion silently replaces -2 log p_j by n λ_j c_j(θ) and then drops the λ_j in the exponent. If a study with a true effect has λ_k = 0, its log p_k is o(n), so the exponential factor exp(n/2 (C_{i+ℓ} - C_i)) collapses to 1, and the polynomial factor (n/2)^{-ℓ} drives the ratio to 0, not ∞. The counterexample is K=2, θ1=θ2=0.5, n1=n-⌊log n⌋, n2=⌊log n⌋. Then L(1,0)/L(1,1) behaves like n^{θ²/4 - 1} → 0, so AW-Fisher asymptotically selects weight vector (1,0) instead of w*=(1,1). The theorem's conclusion is actually false there, not just unproved.\n\nThe fix is natural: require λ_k > 0 for studies with nonzero effects, or at least n_k/log n → ∞. Under such a condition the exponential separation argument goes through. But the statement as written is wrong, and this is a load-bearing flaw, not a cosmetic gap.\n\nThe ABO proof has additional algebra issues: the line with -2/n {Lobs + log(2K-1)} should put the log inside with Lobs, and the lower bound uses the random w_hat where a fixed subset w* is needed. The simulation section has blank fitted constants, and the Figure 2 caption does not match the plot. All fixable, but this is a rough draft.\n\nWould I cite it? No, not in current form. If I were the editor, I would desk reject with an invitation to resubmit after correcting the theorem's assumptions and cleaning up the proofs. The intended result, under a stated proportionality condition, is likely correct and worth publishing, but this manuscript is not there yet.","headline":"The paper's new consistency theorem is false as stated when a nonzero-effect study has o(n) sample size; the intended result likely holds under a proportional-sample-size assumption that the authors forgot to state.","tokens_in":856,"tokens_out":946,"would_cite":false,"duration_ms":72552,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62F05","62F12"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that the adaptively weighted Fisher's method assigns weight 1 exactly to studies with nonzero effects as sample sizes grow, and that it has the same asymptotic Bahadur slope as Fisher's method.","keywords":["adaptively weighted Fisher's method","combining p-values","meta-analysis","consistency","asymptotic Bahadur optimality","exact slope","binary adaptive weights","heterogeneity"],"falsifier":"Take $K=2$, put a nonzero effect in study 1 only, and let $n_1$ grow like $\\log n_2$ so that $n_1/n_2 \\to 0$; this regime is excluded by the theorem, and if simulation shows $P(\\hat{w}_1 = 1)$ does not approach 1, then the proportional-growth assumption is not merely technical but necessary. Conversely, under equal effects one can estimate $-2\\log P(S > s_{obs})/n$ by simulation and check whether it approaches $\\sum_k \\lambda_k c_k(\\theta)$, as Theorem 2 requires.","tokens_in":1696,"feed_emoji":"📊","tokens_out":2626,"duration_ms":89898,"temperature":0.7,"pith_summary":"Meta-analyses that combine p-values from several studies usually gain power but say nothing about which studies actually carry the signal. This paper studies the adaptively weighted Fisher (AW-Fisher) method, which assigns each study a binary weight of 0 or 1, and establishes two asymptotic facts: the estimated weights converge to the true set of studies with nonzero effects as sample sizes grow, and the test's exact Bahadur slope equals that of Fisher's method. A sympathetic reader cares because the weights are used in genomics to group genes by expression patterns across tissues; knowing the weights are consistent means that grouping is trustworthy in large samples, and knowing the slope matches Fisher's means the extra interpretability costs no asymptotic efficiency.","feed_headline":"Adaptive Fisher weights provably find the true studies","feed_subtitle":"New proofs show the weighted test is consistent and as efficient as Fisher's classic method.","key_machinery":"The proof is carried by two asymptotic identities. First, if the p-value of study $k$ has exact slope $c_k(\\theta)$, meaning $-2\\log p_k/n_k \\to c_k(\\theta)$, then with proportional sample sizes $-2\\log p_k/n \\to \\lambda_k c_k(\\theta)$. Second, the chi distribution's survival function satisfies $\\log(1-F_{\\chi_m}(x)) \\sim -x^2/2$ as $x \\to \\infty$, so the AW-Fisher loss $L(T(w;P))$ behaves exponentially with rate $\\sum_k w_k \\lambda_k c_k(\\theta)$. Comparing the tail ratio of the true subset with any other subset exposes an exponential factor $\\exp(n/2 \\sum_j c_j)$ when real studies are dropped, and a factor $O(1/n^{\\ell'})$ when null studies are included; these factors drive the probability of any wrong weight vector to zero. The same tail comparison sandwiches the AW-Fisher p-value between $L_{obs}$ and $(2^K - 1)L_{obs}$, transferring the consistency into equality of exact slopes.","core_discovery":"Let $p_k$ be the p-value from study $k$ and let $T(w;P) = -2\\sum_k w_k \\log p_k$ be Fisher's statistic on a weighted subset. The AW-Fisher statistic is $s(P) = -\\log(\\min_w L(T(w;P)))$, where the minimum is over all nonzero binary weight vectors and $L$ is the chi-square survival probability; the selected weight vector $\\hat{w}$ minimizes $L$. Theorem 1 states that if study sample sizes grow proportionally, $n_k/n \\to \\lambda_k$, then $\\hat{w} \\to w^*$ with probability tending to 1, where $w^*_k = 1$ exactly when $\\theta_k \\neq 0$. Theorem 2 states that, when all nonzero effects take a common value $\\theta \\neq 0$, the exact slope satisfies $c_{AW}(\\theta) = c_{Fisher}(\\theta) = \\sum_k \\lambda_k c_k(\\theta)$, so AW-Fisher is asymptotically Bahadur optimal alongside Fisher's method.","pith_inferences":["If the proportional-sample-size condition fails, for example one study's sample size is $o(n)$ while another grows linearly, the exponential separation in the consistency proof becomes polynomial and weight consistency may fail; a testable extension is to identify the largest unbalanced growth under which $\\hat{w}$ still recovers $w^*$.","The tail-ratio comparison suggests a finite-sample formula for misclassification probabilities, so fitting that formula to simulated or real data could give practical uncertainty estimates for learned weights.","Because asymptotic Bahadur optimality is proven under a common nonzero effect $\\theta_k \\equiv \\theta$, a natural follow-up is to derive the exact slope of AW-Fisher under mixed directions or heterogeneous magnitudes, where its subset selection may diverge from Fisher's slope."],"forward_implications":["As all study sample sizes grow with fixed proportions, AW-Fisher's selected weight vector converges to the indicator of nonzero-effect studies, so the gene-by-tissue pattern categories read off the weights are asymptotically correct.","When all true effects share a common nonzero value, AW-Fisher has the same exact Bahadur slope as Fisher's method, so it carries no asymptotic Bahadur efficiency penalty relative to Fisher's classic test.","Weight errors are not symmetric: a true study is dropped with probability decaying like $n \\exp(-n c_k \\lambda_k /2)$, while a null study is kept with probability only $O(1/n)$, so misclassification of real signals disappears much faster than spurious inclusions.","Because $\\hat{w}$ is consistent and the AW-Fisher p-value lies between $L_{obs}$ and a constant multiple of $L_{obs}$, the test's tail probability and the chi-square tail yield the same slope, which is the direct route to asymptotic Bahadur optimality."],"supporting_citations":[{"why":"Introduces the AW-Fisher statistic and binary-weight selector whose properties the paper studies.","marker":"Li et al. (2011)"},{"why":"Defines exact slopes and asymptotic Bahadur optimality, the framework for both theorems.","marker":"Bahadur (1967)"},{"why":"Supplies the tail asymptotic $\\log(1-F_{\\chi_m}(x)) \\sim -x^2/2$ used in Lemma 2 and the consistency proof.","marker":"Bahadur et al. (1960)"},{"why":"Proves Fisher's method is asymptotically Bahadur optimal, the benchmark for Theorem 2.","marker":"Littell and Folks (1971)"},{"why":"Provides Fisher's method, which is the base statistic $T(w;P)$ for each fixed weight vector.","marker":"Fisher (1925)"}],"fun_headline_variants":["Adaptive Fisher weights match classic efficiency","Consistent and Bahadur optimal: AW-Fisher proven","Proof: AW-Fisher hits Fisher's asymptotic efficiency","True study detection: AW-Fisher weights converge","Optimal meta-analysis: adaptive Fisher weights"],"cache_read_input_tokens":13824,"weakest_assumption_plain":"Every study's sample size must grow at the same rate as the average sample size, so $n_k/n$ has a positive limit; otherwise a study with negligible relative sample size cannot be reliably detected by the weight estimator.","fun_headline_variants_meta":{"raw":{"variants":["Adaptive Fisher weights match classic efficiency","Consistent and Bahadur optimal: AW-Fisher proven","Proof: AW-Fisher hits Fisher's asymptotic efficiency","True study detection: AW-Fisher weights converge","Optimal meta-analysis: adaptive Fisher weights"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000669,"raw_usage":{"total_tokens":3055,"prompt_tokens":956,"completion_tokens":2099,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":2027}},"tokens_in":572,"tokens_out":2099,"duration_ms":15413,"temperature":1.0,"reasoning_tokens":2027,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:47:35.288404+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take $K=2$, put a nonzero effect in study 1 only, and let $n_1$ grow like $\\log n_2$ so that $n_1/n_2 \\to 0$; this regime is excluded by the theorem, and if simulation shows $P(\\hat{w}_1 = 1)$ does not approach 1, then the proportional-growth assumption is not merely technical but necessary. Conversely, under equal effects one can estimate $-2\\log P(S > s_{obs})/n$ by simulation and check whether it approaches $\\sum_k \\lambda_k c_k(\\theta)$, as Theorem 2 requires.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines exact slopes and asymptotic Bahadur optimality, the framework for both theorems."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Proves Fisher's method is asymptotically Bahadur optimal, the benchmark for Theorem 2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides Fisher's method, which is the base statistic $T(w;P)$ for each fixed weight vector."}],"review_version":1}