{"id":"593d1595-40cd-4cb2-b5fe-0939c1e7bf21","arxiv_id":"2508.17716","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Worst-case publication-bias bounds are extended from selection models monotone in standard error to a general selection-model class, computed via nonlinear programming with linear constraints and validated on simulations and two real meta-analyses.","lead":"This statistics paper proposes a new way to compute worst-case bounds for publication bias in meta-analysis, extending a known bound to a wider class of publication-selection models. It uses nonlinear programming with linear constraints to compute approximate worst-case bounds and validates them with simulations and two real meta-analyses.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Approximate worst-case bound may fail to be conservative; abstract gives no guarantee.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: the NLP relaxation might not faithfully represent the true worst-case supremum, and the 'approximate' bound may understate publication bias. Because the paper is abstract-only, this concern cannot be resolved without the full text or code. The reader's UNVERDICTED verdict is appropriate, and my stress test does not change it: the concern is real and central, but it is a demand for verification rather than a demonstrated defect. The proposed concrete test provides a decisive check: by specializing to the C-J class where ground truth is known analytically, one can detect whether the approximation is ever non-conservative. If the test passes, the central claim gains strong support; if it fails, the method is unsafe for sensitivity analysis. Without such evidence, the paper should not be accepted as providing a valid worst-case bound. I therefore keep the verdict unchanged and agree with the reader's framing. No ad hominem, no theatrical language; the critique is about the missing conservatism guarantee, not the authors' competence.","tokens_in":984,"tokens_out":2767,"duration_ms":36499,"concrete_test":"Restrict the proposed general class to the original Copas-Jackson class, where selection probabilities are monotone in study standard error and the exact worst-case bound has a closed analytical form. Generate a large battery of meta-analytic datasets (e.g., 10,000 random combinations of number of studies, sample sizes, true effect size, and heterogeneity), plus the paper's two real-world examples. For each dataset, compute (a) the exact C-J bound using the analytical formula and (b) the proposed NLP approximate bound with the same monotonicity constraint, using the authors' implementation. If any computed approximate bound is less than the corresponding exact C-J bound (beyond small numerical tolerance), the approximation is not conservative, and the claimed generalization fails as a safe sensitivity tool. If it is always ≥, this supports conservatism on the C-J subclass; a second check","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is a worst-case bound for publication bias over a broader class of selection models than Copas-Jackson. The abstract describes the output as 'an approximate worst-case bound via tractable nonlinear programming with linear constraints' and never states whether the approximation is an upper bound (conservative) or a lower bound. In sensitivity analysis for publication bias, a non-conservative approximation is dangerous: if the NLP relaxation undershoots the true supremum of bias over the allowed selection models, the reported interval is too narrow, giving meta-analysts false reassurance. This risk is concrete: many tractable NLP formulations are inner approximations (e.g., finite discretizations of the selection function or convex relaxations that only search a subset), and the abstract provides no theorem or numerical evidence that the approximate bound always lies at or above the true worst-case bound. Without such a guarantee, the method cannot be regarded as a valid sensitivity bound. A secondary but related issue is that the 'general class' of selection models is not defined in the abstract; unless its membership conditions are both broad enough to be interesting and restrictive enough to make the supremum finite and computable, the bound may be vacuous or infinite. The load-bearing premise is that the NLP formulation faithfully and conservatively encodes the supremum over the entire class; this is currently unstated and unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a method for constructing worst-case publication-bias bounds over a broader class of selection models than those allowed by the Copas-Jackson (C-J) bound. The abstract states that the C-J bound covers only selection models monotone in study standard error, while the proposed method weakens this assumption and computes an approximate worst-case bound via tractable nonlinear programming with linear constraints. The effectiveness is claimed to be substantiated by extensive simulations and two real-world meta-analyses. The central contribution, if valid, is a sensitivity-analysis tool that is less assumption-dependent than the existing C-J bound.","tokens_in":1160,"tokens_out":2102,"duration_ms":27286,"significance":"Worst-case bounds for publication bias are valuable because they avoid strong parametric assumptions in sensitivity analysis. The C-J bound is a nonparametric benchmark, but its monotonicity restriction is limiting. If the proposed generalization is rigorous and the NLP approximation is provably conservative, the paper would make a useful methodological contribution. The significance is therefore potentially high. However, because the full text is not available for review and the abstract omits several load-bearing technical details, the significance cannot be confirmed from the submitted material alone.","major_comments":[{"comment":"The phrase 'approximate worst-case bound via tractable nonlinear programming with linear constraints' does not state whether the approximation is an upper bound (conservative) or a lower bound (optimistic). In publication-bias sensitivity analysis, a non-conservative approximation is problematic: if the NLP optimum undershoots the true supremum of bias over the allowed selection models, the resulting interval is too narrow and gives false reassurance. The manuscript must either prove that the NLP formulation yields an upper bound over the entire model class, or specify the sense in which the approximation is valid (e.g., asymptotically conservative under discretization refinement). Without such a statement the central claim is not established.","section":"Abstract, sentence 4"},{"comment":"The 'general class' of selection models is not defined. The contribution rests on the claim that this class strictly generalizes the C-J monotonicity class, but the abstract gives no membership conditions. The paper must formally define the class, state the assumptions imposed on the selection function (e.g., monotonicity in a different variable, bounds on selection probabilities, smoothness), and show that the supremum of bias over this class is finite and representable by the proposed finite-dimensional NLP. Without this, the bound could be vacuous, infinite, or not actually weaker than the C-J assumption.","section":"Abstract, sentence 3"},{"comment":"The validation claim 'extensive simulation studies' and 'two real-world meta-analyses' is asserted without any summary of the results. The paper should report, at minimum, the simulation design (data-generating models, selection mechanisms, sample sizes), the outcome measure (e.g., coverage rate of the true bias or true effect), and a comparison with the C-J bound. Crucially, the simulations must verify conservatism: the approximate worst-case bound should be at or above the true worst-case publication bias in the scenarios studied. The abstract alone gives no evidence for this property, which is the central risk identified above.","section":"Abstract, sentence 5"}],"minor_comments":[{"comment":"The abstract would benefit from a precise statement of the optimization problem and the nature of the approximation, even in qualitative terms. For example, is the NLP an inner or outer approximation of the feasible set, and does the 'linear constraints' formulation follow from discretization or from convexification?","section":"Abstract, general"},{"comment":"The acronym PB is introduced but not used again; consider removing it or using it consistently. Also, the phrase 'provide a more objective evaluation' is vague; the comparison to trim-and-fill should be stated in terms of assumptions or bias, not 'objectivity'.","section":"Abstract, sentence 1"},{"comment":"No software or code availability is mentioned. For a numerical method, sharing code would improve reproducibility; this is a minor point for the revision.","section":"Not applicable"}],"recommendation":"uncertain","confidential_remarks":"This review is based only on the abstract; the full text was not provided. The 'uncertain' recommendation reflects the inability to verify the central derivation, the NLP formulation, and the claimed simulation results, not a specific technical objection. I would urge the editor to obtain the full manuscript before making a decision. If the full paper contains a theorem that the NLP solution is a conservative upper bound over the formally defined model class, and simulations that support this property, the paper could be publishable. As it stands, the abstract alone does not establish the main claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one-sentence take: this is a real extension of the Copas-Jackson worst-case bound, but the abstract leaves open the possibility that the 'approximate worst-case bound' is not an upper bound, and that would understate bias. The paper deserves peer review, but the conservatism question is the first thing I'd put to the authors.\n\nWhat's actually new: weakening the monotone-in-standard-error restriction on selection models and replacing the analytical C-J bound with a nonlinear program with linear constraints. That is a genuine extension if the class is as general as claimed. They also ran simulations and two real meta-analyses, which is more than many theory papers do.\n\nWhere it's soft: the abstract says 'approximate worst-case bound' without stating the direction of the approximation. For a sensitivity bound, you need the output to be an upper bound on the true worst-case bias, at least in expectation or with high probability. If the NLP is an inner approximation, the interval is too narrow and gives false reassurance. The stress-test note is right that no such guarantee is stated. The secondary issue is that the selection-model class is not defined in the abstract, so a reader can't tell whether the linear constraints cover the whole class or just a tractable subset. Both could be resolved in the full text, but they need to be explicit.\n\nI don't share the circularity concern: the method generalizes an external benchmark, and there's no obvious fitting of the target result. The absence of code and data is a minor issue for a methods paper, though it would help.\n\nBottom line: if the full text proves that the NLP optimum is an upper bound on the supremum over the full class—or gives a numerical guarantee that the approximation error is non-negative—this is a solid contribution. If it only gives an approximation without a direction, it's not yet a valid sensitivity bound. That is exactly what peer review should settle.\n\nI'd send it to review, and I'd bring it to reading group to argue about the NLP encoding. I wouldn't cite it yet, but I'd watch for the revision.","headline":"A legitimate generalization of the Copas-Jackson bound, but the abstract never says whether the NLP approximation is conservative, and that is the make-or-break question.","tokens_in":1710,"tokens_out":1837,"would_cite":false,"duration_ms":22111,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that worst-case bounds for publication bias in meta-analysis can be computed over a wider nonparametric class of selection models than the Copas-Jackson bound allows, using tractable nonlinear programming with linear constr","keywords":["publication bias","meta-analysis","selection models","Copas-Jackson bound","worst-case bound","sensitivity analysis","nonlinear programming"],"falsifier":"Construct a concrete non-monotone selection model in the proposed general class, compute the true worst-case bias over that model by exhaustive search or an analytic formula, and then run the proposed nonlinear program. If the program's approximate worst-case bound is smaller than the true worst case, the method understates publication bias and the central claim fails.","tokens_in":743,"feed_emoji":"📈","tokens_out":2934,"duration_ms":35598,"temperature":0.7,"pith_summary":"Publication bias can distort meta-analysis, and sensitivity analysis with selection models is a way to quantify it. The Copas-Jackson bound gives a worst-case bias bound but only for selection models that are monotone in study standard error. This paper tries to establish that a similar worst-case bound can be obtained for a strictly broader class of selection models, dropping that monotonicity assumption. The proposed bound is approximate and computed numerically, and the authors support it with simulations and two real meta-analyses. If the method works, meta-analysts gain a defensible sensitivity bound under weaker assumptions than before.","feed_headline":"Worst-case publication-bias bounds survive weaker assumptions","feed_subtitle":"A generalized Copas-Jackson bound no longer requires publication chance to track study error.","key_machinery":"The key object is the selection model, which describes the probability that a study is published as a function of its outcome and characteristics. The Copas-Jackson bound is an analytical worst-case bias bound over selection models that are monotone in study standard error. The paper's machinery replaces that restrictive class with a broader nonparametric class and reformulates the worst-case bias as an optimization problem. The approximate worst-case bound is then obtained by tractable nonlinear programming with linear constraints, which is the mechanism that makes the generalization computable.","core_discovery":"The paper's central claim is that the Copas-Jackson bound, which holds over selection models where publication probability is monotone in study standard error, can be extended to a general class of selection models. For that general class the paper constructs an approximate worst-case bound by solving a nonlinear program whose constraints are linear. The practical question is: when a meta-analyst cannot credibly assume monotonicity in standard error, can they still put a worst-case number on publication bias? The paper's answer is yes, with a bound that generalizes the Copas-Jackson construction and remains computable in practice.","pith_inferences":["The same computational strategy could be applied to other sensitivity quantities, such as worst-case shifts in effect size, by keeping the selection-model constraints and changing the objective.","If the approximate bound is conservative, it could be inverted to identify which selection models would overturn a meta-analytic conclusion, giving a threshold for criticism.","The method might extend to meta-regression or multivariate outcomes, where monotonicity in standard error is even harder to defend; this extension is not claimed in the paper."],"forward_implications":["Meta-analysts can perform publication-bias sensitivity analysis without requiring publication probability to be monotone in standard error.","The proposed worst-case bound can be computed numerically even though it covers a more flexible class of selection models.","The bound provides a more objective sensitivity check than simple graphical methods like trim-and-fill, while weakening the assumptions of the Copas-Jackson bound.","If the bound is used in practice, conclusions from meta-analyses can be reported with a worst-case bias interval under a broader class of publication mechanisms."],"supporting_citations":[],"fun_headline_variants":["Publication-bias bounds extended beyond monotone selection","Worst-case bias bound without monotonicity assumption","Generalized Copas-Jackson bound for non-monotone selection","Relaxing assumptions on publication-bias bounds","New worst-case bound for publication bias under weaker assumptions"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the nonlinear programming relaxation faithfully represents the full general class of selection models, so that the computed approximate worst-case bound never falls below the true worst-case publication bias.","fun_headline_variants_meta":{"raw":{"variants":["Publication-bias bounds extended beyond monotone selection","Worst-case bias bound without monotonicity assumption","Generalized Copas-Jackson bound for non-monotone selection","Relaxing assumptions on publication-bias bounds","New worst-case bound for publication bias under weaker assumptions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000191,"raw_usage":{"total_tokens":1153,"prompt_tokens":692,"completion_tokens":461,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":436,"completion_tokens_details":{"reasoning_tokens":384}},"tokens_in":436,"tokens_out":461,"duration_ms":4669,"temperature":1.0,"reasoning_tokens":384,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T16:45:22.769056+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a concrete non-monotone selection model in the proposed general class, compute the true worst-case bias over that model by exhaustive search or an analytic formula, and then run the proposed nonlinear program. If the program's approximate worst-case bound is smaller than the true worst case, the method understates publication bias and the central claim fails.","supporting_citations":[],"review_version":1}