{"id":"d22da3fc-1af2-4ef0-a694-da29c2b3e5ac","arxiv_id":"2507.13615","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A full likelihood method using empirical likelihood for the Copas-like publication bias model lowers finite-sample mean squared error and improves confidence interval coverage, but with no asymptotic efficiency gain.","lead":"This paper proposes a new statistical method for correcting publication bias in meta-analysis, combining a selection model with a nonparametric estimate of the study standard errors. The authors show the method produces confidence intervals with better coverage in simulations, although it offers no asymptotic efficiency gain over existing methods.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The implemented estimator fixes γ1,γ2 at conditional MLEs, so Theorem 1's chi-square/normality results apply to the uncomputed full MLE, not to the method actually used; the flat-plateau argument in Section 5 is heuristic and untested.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern I find: the implemented estimator fixes γ1 and γ2 at conditional MLEs, so Theorem 1's asymptotic results for the full MLE do not rigorously apply to the procedure that is actually used. This is not a minor technicality because the paper's main practical claims—more accurate coverage probabilities for likelihood-ratio intervals—depend on the χ² calibration of the implemented statistic, not on the uncomputed full MLE. The authors' flat-plateau argument in Section 5 is explicitly informal and is not a substitute for a theorem or a distributional check. I also note two secondary issues that do not change the main verdict: Proposition 1 shows no asymptotic efficiency gain, so the abstract's efficiency language should be qualified to finite samples, and the simulations report coverage but not interval widths, leaving open the possibility that the coverage improvement is purchased with impractically wide intervals. None of these issues makes the paper's contribution worthless; the simulations are promising and the full-likelihood formulation is a useful conceptual step. But the theoretical claims should be aligned with the implementation, or the implementation should be changed to approximate the full MLE with a proven guarantee. A concrete simulation-based distributional check of the implemented LR statistic would settle whether the concern actually lands, and the verdict should remain CONDITIONAL pending that check.","tokens_in":15212,"tokens_out":4125,"duration_ms":54426,"concrete_test":"Use a Section 3 configuration (e.g., N0=100, γ12,0=(-0.6,0.8), θ0=0.4, τ0=0.5, ρ0=0.8) and generate B=2000 datasets. For each dataset, compute the likelihood-ratio statistic for H0: θ=θ0 using the Section 2.5 algorithm with γ12 fixed at the conditional MLE, and also with γ12 fixed at the true γ12,0. Compare the empirical distribution of each statistic to χ²_1 via QQ plots and a Kolmogorov-Smirnov test. If the implemented version deviates from χ²_1 while the true-γ12 version does not, the mismatch is caused by the plug-in step and Theorem 1 does not cover the implemented estimator.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim requires that the estimator actually implemented be, at least asymptotically, the full MLE analyzed in Theorem 1. Section 2.5 instead fixes γ12 = (γ1,γ2) at the conditional MLE from Ning et al. (2017)'s EM algorithm because direct maximization is unstable, and then maximizes the remaining profile likelihood. The resulting estimator is not the argmax of ℓ(N,α,γ) over the full parameter vector, so Theorem 1's joint normality and χ² likelihood-ratio limits do not formally apply to the reported procedure. The Discussion explicitly admits that 'the resulting parameter estimates are in essence different from the true maximum likelihood estimates,' and the justification that the full likelihood is a flat plateau is a heuristic, not a proof. A root-n consistent plug-in estimator for a subset of nuisance parameters can change the limiting distribution of a profile likelihood ratio unless a strong orthogonality or flatness condition holds; no such condition is stated or verified. This matters because the headline simulation claims are about coverage probabilities of likelihood-ratio intervals, and Proposition 1 independently shows that full and conditional point estimators are asymptotically equivalent, so the practical value of the method rests on the finite-sample distribution of the implemented LR statistic. If that distribution is not χ², the reported coverage improvements for the implemented method are unsupported. Secondary concerns include the absence of reported interval widths, which makes it unclear whether the improved coverage comes from wider intervals, and the tension between the abstract's efficiency language and Proposition 1's no-asymptotic-gain result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a full likelihood approach for meta-analysis under the Copas-like selection model of Ning et al. (2017). It combines the conditional likelihood for the published effect sizes with an empirical likelihood for the marginal distribution of the study-specific standard errors, and treats the total number of studies N as an unknown binomial parameter. The main theoretical result (Theorem 1) claims joint asymptotic normality of the full likelihood estimators of (N, alpha, gamma) and a central chi-square limit for the likelihood ratio statistic; Theorem 2 and Proposition 1 give the corresponding conditional-likelihood results and establish asymptotic equivalence of the two point estimators. Simulations compare the proposed estimators and likelihood ratio intervals with conditional-likelihood estimators and Wald intervals, and a real data example is analyzed. The discussion acknowledges that the implemented algorithm fixes gamma1 and gamma2 at conditional maximum likelihood estimates rather than maximizing the full likelihood over all parameters.","tokens_in":15445,"tokens_out":6990,"duration_ms":86309,"significance":"If the chi-square calibration of the implemented likelihood ratio statistic could be justified, the paper would provide a practically useful route to confidence intervals for N, theta, tau, and rho that avoids the instability of inverse-probability-weighting and the absurd Wald lower bounds for N. The empirical-likelihood formulation of the marginal distribution of the standard errors is natural, the profile likelihood derivation is coherent, the simulation study is extensive, and the QQ-plot diagnostic is a useful addition. The main limitation is that the asymptotic theory does not cover the two-step estimator actually implemented, and Proposition 1 shows that there is no asymptotic efficiency gain over conditional likelihood, so the contribution rests on finite-sample and interval-calibration properties that are currently supported only by simulation.","major_comments":[{"comment":"The algorithm in Section 2.5 fixes gamma12 at the conditional MLE from Ning et al. (2017) and then maximizes over the remaining parameters, but Theorem 1 is stated for the unconstrained maximizer of ell(N, alpha, gamma). The Discussion in Section 5 concedes that the resulting parameter estimates are 'in essence different from the true maximum likelihood estimates.' Replacing a subset of nuisance parameters by a root-n consistent estimator in a profile likelihood does not in general preserve the chi-square limit of the likelihood ratio statistic unless an orthogonality or flatness condition holds; the flat-plateau statement is a heuristic and no verification is provided. Consequently, the asymptotic justification for the likelihood ratio intervals in Tables 1 and 2 and the QQ-plots in Figure 1 does not formally apply to the estimator actually computed. The authors should either prove the relevant profile-likelihood result for the two-step estimator, verify the negligibility condition empirically by comparing the two-step and full maximizers on simulated data, or explicitly reframe the theoretical claims and provide an alternative calibration for the implemented procedure.","section":"Section 2.5 and Section 5"},{"comment":"The finite-sample coverage comparisons compare likelihood ratio intervals based on the proposed full-likelihood procedure with Wald intervals based on the conditional-likelihood procedure. Since Proposition 1 shows that the two point estimators are asymptotically equivalent, the reported coverage improvements may be due to the use of likelihood ratio calibration rather than to the full-likelihood estimator itself. To support the attribution in the abstract and introduction, the authors should add a comparison with likelihood-ratio-type intervals based on the conditional likelihood or with Wald intervals computed from the full-likelihood estimates; otherwise the conclusion should be softened to claim an improvement of the full-likelihood LR interval over the conditional-likelihood Wald interval as a complete procedure.","section":"Section 3.2 and Proposition 1"},{"comment":"Theorem 1 is stated under Conditions C1 and C2, which are described only as being in the supplementary material, and the proof is deferred as well. The main text should at least state these conditions and comment on whether the simulation designs satisfy them, particularly the positive-definiteness of Omega and the moment conditions needed for the empirical likelihood weights. Without this, the reader cannot assess whether the simulation settings fall within the theoretical regime claimed by Theorem 1.","section":"Theorem 1 and supplementary conditions"}],"minor_comments":[{"comment":"The sentence 'For point estimation of N, the full likelihood and conditional likelihood estimates are 23 and 16, respectively' is reversed relative to Table 3, which reports CL = 23 and FL = 16.","section":"Section 4"},{"comment":"The caption says 'Meta analysis results of the lung cancer data and the premature birth data,' but the text and the rest of Section 4 analyze only the premature birth data.","section":"Table 3 caption"},{"comment":"The quantity s_i is introduced as 'the estimated standard variance'; given its role as the standard error in models (1) and (2), the term 'standard error' or 'standard deviation' would be clearer and would avoid confusion with s_i^2 as a variance.","section":"Section 2.1"},{"comment":"The phrase 'To then end' should be 'To this end,' and the reference 'Contrilled Clinical Trials' should be 'Controlled Clinical Trials.'","section":"Section 5 and references"},{"comment":"The statement that the proof of Theorem 1 implies a chi-square limit for likelihood ratio statistics for any subvector of (N, alpha, gamma) is asserted informally; a formal statement of this profile likelihood result would strengthen the interval-construction claims.","section":"Section 2.4"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of a statistical methodology journal, but the main theorem does not cover the estimator that is actually implemented, and this is a load-bearing issue for the interval coverage claims. I would ask the authors to close this gap or to revise the theoretical claims downward before publication. The concern is correctable within the manuscript's scope, so I am not recommending rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a serious and mostly honest methods paper. The genuinely new piece is combining the Copas-like selection model's conditional likelihood with a semiparametric empirical likelihood for the marginal distribution of standard errors to get a full likelihood. Identifiability for ρ≠0 is stated, and the asymptotic normality and chi-square LR results for the full MLE are nontrivial. The authors also deserve credit for Proposition 1: it says plainly that there is no asymptotic efficiency gain over conditional MLE, which undercuts the abstract's hint of efficiency but is the right thing to report. The simulation evidence does show smaller RMSEs for θ, τ, ρ and N in most settings, and the LR intervals typically have better coverage than Wald intervals.\n\nThe soft spot is exactly where the stress-test lands. The implemented algorithm in Section 2.5 fixes γ1 and γ2 at the conditional MLEs from Ning et al.'s EM, then maximizes the remaining profile likelihood. That estimator is not the argmax of ℓ(N, α, γ), so Theorem 1's joint normality and χ² property do not formally apply to the procedure used in the simulations. The authors admit in the Discussion that the resulting estimates are in essence different from true MLEs, and the flat-plateau argument is heuristic. Without a theorem or a careful numerical check for the plug-in version, the headline coverage claims are not backed by the theory.\n\nThere is a second issue that gets too little attention: the N estimates are wildly volatile in Tables 1 and 2, with RMSEs sometimes ten times the true N0. Coverage for N may be well calibrated, but no interval widths are reported. If the LR interval for N is typically [n, huge], the coverage gain is less impressive. The real-data interval [13,20] is fine, but that is one example.\n\nOverall, the math is coherent under stated assumptions, the data are simulations plus one real example, and the citation pattern is reasonable—their own EL papers are genuinely the starting point. This paper deserves a serious referee, but the authors need to close the gap between Theorem 1 and their algorithm, or re-frame the claims as simulation-based. They should also report interval lengths. For a reading group, it is a useful case study in profile likelihood with plug-in nuisance parameters.","headline":"Serious EL-based full-likelihood method for Copas selection, but the theory covers the uncomputed full MLE while the implementation uses plug-in gammas; the coverage gains need interval widths and a proof for the implemented procedure.","tokens_in":16031,"tokens_out":4029,"would_cite":false,"duration_ms":45774,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62G05"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a full likelihood formed by combining the Copas conditional likelihood with an empirical likelihood for study standard errors gives jointly normal estimators and asymptotically chi-square likelihood ratios, and that…","keywords":["meta-analysis","publication bias","Copas selection model","empirical likelihood","full likelihood","likelihood ratio confidence intervals","between-study heterogeneity","total number of studies"],"falsifier":"Generate simulated meta-analyses from the model with known parameters, but choose settings where the conditional EM estimates of $\\gamma_1$ and $\\gamma_2$ are visibly far from the full likelihood maximizers; compute the implemented two-step estimator and check whether 95% likelihood-ratio intervals for $\\theta$ and $N$ still cover at about 95%. A large drop in coverage when the plugged-in $\\gamma_{12}$ differs materially from the true MLE would refute the claim that the implemented estimator inherits Theorem 1's chi-square limit.","tokens_in":14945,"feed_emoji":"📊","tokens_out":10610,"duration_ms":114989,"temperature":0.7,"pith_summary":"Meta-analysis combines published studies, but publication bias and between-study heterogeneity can distort conclusions. The paper develops a full-likelihood method for a Copas-style selection model by combining the usual conditional likelihood with an empirical likelihood for the marginal distribution of study standard errors. It claims that the resulting maximum likelihood estimators are jointly asymptotically normal and that the full likelihood ratio is asymptotically chi-square, enabling likelihood-ratio confidence intervals for the overall effect size and for the total number of studies. In simulations, these intervals have noticeably more accurate coverage than the conditional-likelihood Wald intervals, and the point estimators usually have smaller mean squared errors. The paper also shows that all underlying parameters are identifiable whenever publication bias is present ($\\rho\\neq 0$).","feed_headline":"Full-likelihood intervals beat Wald for publication bias","feed_subtitle":"Copas selection model plus empirical likelihood yields intervals that hold near-nominal coverage.","key_machinery":"The load-bearing object is the profile semiparametric full log-likelihood, obtained by replacing the unknown marginal distribution of the study standard errors $s_i^*$ with Owen empirical-likelihood weights $p_i=n^{-1}[1+\\lambda\\{\\Phi(\\gamma_1+\\gamma_2/s_i)-\\alpha\\}]^{-1}$ satisfying $\\sum_i p_i=1$ and $\\sum_i p_i\\Phi(\\gamma_1+\\gamma_2/s_i)=\\alpha$. With $v_i(\\gamma)=(\\gamma_1+\\gamma_2/s_i+\\rho s_i(\\theta_i-\\theta)/(\\tau^2+s_i^2))/\\sqrt{1-\\rho^2 s_i^2/(\\tau^2+s_i^2)}$, the profile log-likelihood is $$\\ell(N,\\$\\alpha$,\\gamma)=\\log\\binom{N}{n}+(N-n)\\log(1-\\$\\alpha$)+\\sum_i\\left[\\log\\Phi(v_i(\\gamma))-\\frac12\\log(\\$tau^{2}$+$s_i^{2}$)-\\frac{(\\theta_i-\\$\\theta$)^2}{2(\\$tau^{2}$+$s_i^{2}$)}\\right]-\\sum_i\\log[1+\\$\\lambda$\\{\\Phi(\\gamma_1+\\gamma_2/s_i)-\\$\\alpha$\\}].$$ This turns the conditional Copas likelihood into a full likelihood that also uses information in the observed number of studies $n$ and in the marginal distribution of $s^*$. Theorem 1 establishes joint asymptotic normality of the maximizers and a $\\chi^2_7$ limit for the full likelihood ratio; Proposition 1 shows the asymptotic covariance equals that of the conditional likelihood.","core_discovery":"Under the Copas-like selection model of Ning et al. (2017), the paper proposes a full semiparametric likelihood that joins the conditional likelihood for the observed effect estimates with Owen's empirical likelihood for the unknown marginal distribution of study standard errors. The paper proves that, provided publication bias exists ($\\rho\\neq 0$) and mild regularity conditions hold, the full likelihood maximizers of $(N,\\alpha,\\gamma)$ are jointly asymptotically normal and the full likelihood ratio statistic converges to a central $\\chi^2_7$ distribution. This yields likelihood-ratio confidence intervals for $\\theta$, $\\tau$, $\\rho$, $\\alpha$, and $N$ with asymptotically correct coverage, and consistent estimators of the marginal distributions of standard errors and effect sizes. Proposition 1 shows that the asymptotic variances of the full and conditional estimators of $N$ and $\\gamma$ are identical, so the full likelihood offers no asymptotic point-estimation efficiency gain; its advantage lies in likelihood-ratio inference and in avoiding unstable inverse-probability weights. Simulation results in the paper report smaller mean squared errors and, especially, more accurate coverage probabilities for the full likelihood intervals than for conditional-likelihood Wald intervals, and a real data example illustrates that the two methods can differ on whether publication bias is detected.","pith_inferences":["Editorial inference: the equality of asymptotic variances in Proposition 1 suggests the full likelihood's practical value may generalize across biased-sample problems: point estimation is not improved asymptotically, but likelihood-ratio inference avoids variance estimation and naturally respects parameter constraints such as $N\\ge n$.","Editorial inference: because the flatness in $(\\gamma_1,\\gamma_2)$ is the computational bottleneck, profiling the full likelihood over these two parameters or calibrating the two-step likelihood ratio by bootstrap could yield an implemented estimator for which the stated chi-square theory literally holds, rather than the current fixed-$\\tilde\\gamma_{12}$ version.","Editorial inference: the observed superiority of the chi-square approximation over the normal approximation to the Wald statistic is a finite-sample phenomenon; a natural testable extension is whether it persists under misspecification of the selection equation or of the random-effects distribution."],"forward_implications":["Under the assumed model with $\\rho\\neq 0$, the full likelihood approach gives consistent, jointly normal estimators of $\\theta$, $\\tau$, $\\rho$, $\\alpha$, and $N$, and likelihood-ratio statistics with asymptotic chi-square limits, so confidence intervals can be built without estimating variances.","Because the empirical likelihood weights prevent extreme small probabilities, the estimated marginal distributions of study standard errors and effect sizes avoid the instability of inverse-probability-weighting estimators.","The Proposition 1 result implies that the full likelihood does not improve asymptotic point-estimation efficiency over the conditional likelihood; its benefit is in likelihood-ratio inference rather than in smaller asymptotic variance.","In the premature-birth example, the full likelihood ratio interval for $\\rho$ excludes zero while the conditional Wald interval contains zero, so the method can change the conclusion about whether publication bias is present.","The likelihood-ratio interval for $N$ respects the constraint $N\\ge n$, avoiding Wald intervals whose lower bounds fall below the number of observed studies."],"supporting_citations":[{"why":"Supplies the Heckman-style latent variable selection model from which the Copas model is derived.","marker":"Copas and Li (1997)"},{"why":"Provides the conditional likelihood baseline, the EM algorithm used for conditional estimation, and the conditional MLEs to which the paper fixes $\\gamma_1$ and $\\gamma_2$ in implementation.","marker":"Ning et al. (2017)"},{"why":"Supplies the empirical likelihood machinery used to handle the nonparametric marginal distribution of study standard errors.","marker":"Owen (1990)"},{"why":"Supplies the random-effects model connecting each observed effect estimate to the underlying effect size and between-study heterogeneity.","marker":"DerSimonian and Laird (1986)"},{"why":"Documents the flat likelihood and non-convergence problem that motivates fixing $\\gamma_1$ and $\\gamma_2$ during maximization.","marker":"Copas and Shi (2001)"},{"why":"Provides the observed-information matrix procedure used with the EM algorithm to construct the Wald intervals for the conditional likelihood method.","marker":"Louis (1982)"},{"why":"Introduces the inverse-probability-weighting estimator for the total number of studies whose Wald intervals the paper criticizes.","marker":"Mavridis et al. (2013)"},{"why":"Supplies the premature birth meta-analysis dataset used in the real-data comparison.","marker":"Copas and Jackson (2004)"}],"fun_headline_variants":["Full likelihood sharpens meta-analysis under publication bias","Empirical likelihood fixes coverage in Copas selection models","Better confidence intervals for publication-biased meta-analysis","Full likelihood boosts interval accuracy over conditional methods","Copas-like selection: full likelihood yields correct coverage"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method as actually computed fixes the two parameters that determine how strongly a study is selected for publication to values obtained from a separate fitting procedure, and the paper assumes without proof that this substitution does not change the asymptotic results or interval coverage.","fun_headline_variants_meta":{"raw":{"variants":["Full likelihood sharpens meta-analysis under publication bias","Empirical likelihood fixes coverage in Copas selection models","Better confidence intervals for publication-biased meta-analysis","Full likelihood boosts interval accuracy over conditional methods","Copas-like selection: full likelihood yields correct coverage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00061,"raw_usage":{"total_tokens":2842,"prompt_tokens":951,"completion_tokens":1891,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":567,"completion_tokens_details":{"reasoning_tokens":1833}},"tokens_in":567,"tokens_out":1891,"duration_ms":16301,"temperature":1.0,"reasoning_tokens":1833,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:19:50.649963+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate simulated meta-analyses from the model with known parameters, but choose settings where the conditional EM estimates of $\\gamma_1$ and $\\gamma_2$ are visibly far from the full likelihood maximizers; compute the implemented two-step estimator and check whether 95% likelihood-ratio intervals for $\\theta$ and $N$ still cover at about 95%. A large drop in coverage when the plugged-in $\\gamma_{12}$ differs materially from the true MLE would refute the claim that the implemented estimator inherits Theorem 1's chi-square limit.","supporting_citations":[],"review_version":1}