{"id":"bf380111-0592-479b-b580-8a4d8b56608c","arxiv_id":"2607.04327","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Reference-anchored meta-analysis uses an external untreated-outcome distribution as both target and instrument to recover target-population effects that published averages overstate under selection and target mismatch.","lead":"A new meta-analysis method uses an external untreated-outcome distribution both to define the patient population of interest and to correct publication bias via control-arm means as instruments. It recovers a known all-trials antidepressant benchmark in a strict holdout and yields smaller target-population effects for insomnia sleep time and diabetes HbA1c.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged exclusion restriction.","rationale":"The reader's strongest_claim accurately restates Theorem 1 and the strict holdout result. The identification argument is clean under the stated assumptions; the external anchor supplies the missing content that published-only funnel/p-value methods lack; the holdout is a genuine out-of-sample recovery of the aggregate estimand; and the method correctly refuses to answer when first-stage F is weak (sleep quality). The only material vulnerability is the exclusion restriction the reader already named. Because the paper probes it with a Copas grid spanning roughly twofold odds shifts per SD of control level and finds no tipping, and because the holdout itself is robust to published-only versus combined anchors, that vulnerability does not overturn the written claim. No further load-bearing concern (e.g., circular use of hidden trials, mechanical artifact from standardization, or unacknowledged non-identification of absolute rates) survives scrutiny of the text and supplementary materials. Verdict therefore remains CONDITIONAL for the same reasons the reader gave: real-world exclusion and external-reference construction still warrant multi-literature caution and a frozen public archive, not because the central derivation fails.","tokens_in":30695,"tokens_out":612,"duration_ms":7739,"concrete_test":"Re-fit the published-only antidepressant holdout (K=71) after adding the design covariates the paper enumerates as blockers of direct C\to R paths (drug class, era, registration status, country if available) into both the outcome model and the selection probit; if Δ* moves outside the reported CI [0.15,0.30] or the all-trials benchmark 0.248 exits the new interval, residual exclusion violation is material.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is identification of the bias-corrected outcome model and relative selection profile (hence Δ*) from published-only (Y,C,U) under exclusion C ⊥ R | Y,U, external anchor of p(C|U), and relevance/completeness (β_C ≠ 0), with the strict Cipriani holdout recovering the all-trials benchmark. That identification (Theorem 1 / Lemmas S1–S2) is standard once the three assumptions are granted; the holdout is correctly described as recovering the aggregate inverse-selection-weighted mean, not individual suppressed trials; first-stage F, anchor-sensitivity, and Copas ρ_C curves are reported a priori. The reader's weakest_assumption (exclusion) is the genuine soft spot, but the paper already bounds residual direct C\to R coupling on an interpretable odds scale and shows no qualitative reversal. No additional load-bearing inconsistency, hidden circularity, or unacknowledged failure mode is required for the written claim to stand.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes reference-anchored meta-analysis: an external untreated-outcome distribution is used both to define a target-population estimand Δ* (bias-corrected effect at a declared control level m*) and, via each trial’s control-arm mean as an externally anchored shadow instrument, to identify publication selection under exclusion C ⊥ R | Y, U, external knowledge of p(C|U), and relevance/completeness (β_C ≠ 0). Theorem 1 (formalized in S1) shows that the bias-corrected outcome model and relative selection profile are identified from published-only (Y,C,U); the absolute publication rate is calibrated from registries. A strict holdout on the Cipriani antidepressant network—removing recovered-unpublished contrasts from both fitting and anchoring—recovers Δ* = 0.22 whose CI contains the all-trials benchmark 0.248. Applications to insomnia TST and placebo-controlled HbA1c yield target-population effects materially smaller than published averages, while a first-stage F diagnostic declines weak-instrument cases (e.g., self-rated sleep quality). Diagnostics include Copas-type exclusion sensitivity, anchor-sensitivity curves, and multi-start/bootstrap calibration.","tokens_in":31017,"tokens_out":1702,"duration_ms":22112,"significance":"If the identification and holdout results hold under the stated assumptions, the contribution is genuine and practically important: it separates publication selection from target mismatch—two problems usually conflated in meta-analysis—and supplies identifying content unavailable to funnel- or p-value-only correctors. The strict Cipriani holdout is a rare out-of-sample positive control in the publication-bias literature and is correctly scoped as recovery of an aggregate inverse-selection-weighted mean, not prediction of individual suppressed trials. Built-in pre-correction diagnostics (first-stage F), auditable anchor construction (Tables 2, M5), registry falsification, Copas ρ_C grids on an interpretable odds scale, and a full replication package are real strengths. The framework connects cleanly to nonresponse-IV and proximal causal inference while making the dual target/instrument role of the external reference explicit. For clinical evidence synthesis this is a substantive advance over published-only selection models and RoBMA-style averaging.","major_comments":[{"comment":"Assumption 1 (exclusion: C ⊥ R | Y, U; M2, Theorem 1) remains the load-bearing soft spot. S7’s Copas ρ_C grid on [-0.5,0.5] (≈ up to ~2× publication odds per SD of control level) shows no qualitative reversal in the three main applications, which is reassuring, but the manuscript should state more explicitly which design covariates U are required to block the most plausible residual pathways (perceived clinical importance, sponsor strategy, drug class, era, country) in each application, and what happens when those covariates are unavailable. A short “when exclusion is least credible” checklist in Box 3 or the Discussion would make the claim auditable rather than asserted.","section":"M2 / S7 / Assumption 1"},{"comment":"The paper correctly distinguishes three control distributions (m0 completed-trial, mA anchor, m* target; M1, Fig. 1) and notes that a publication-selection reading requires m0 ≈ mA. In the insomnia and diabetes applications the anchor is fully external/epidemiological, so the selection component is more naturally “selection of published evidence relative to the declared target.” The main text sometimes still reads as if the correction is pure publication selection. Please make the interpretive default explicit in Results and Discussion (and in the Fig. 3 caption): report the selection-at-published-composition piece and the target-shift piece separately in every application, and state when the former should not be labeled “publication bias.”","section":"Results / Fig. 3 / M1"},{"comment":"Diabetes is presented as an “objective-outcome generalization” at moderate first-stage strength (F = 7.2; Table 1, S12, Table S14). The matched-simulation calibration of F labels is useful, but the primary single-specification Δ* ≈ 0.40 (CI 0.16–0.55) is wide and the registry-constrained/model-averaged value is 0.35. The manuscript should either (i) elevate the registry-constrained estimate as primary with a pre-specified decision rule, or (ii) report a single model-averaged primary with the linear/quadratic and baseline/endpoint instrument variants as sensitivity, so readers are not left choosing among 0.35–0.40. The leave-one-class-out range [0.32, 0.39] is good; fold it into the primary reporting.","section":"Results (Type-2 diabetes) / S12 / Table S13"}],"minor_comments":[{"comment":"Box 2 and M3 describe a probit selection in T = Y/SE with optional quadratic term and a no-selection model-averaging component. State the default primary specification (linear vs quadratic; MA weights on/off) once in the main text so application tables are unambiguous.","section":"Box 2 / M3"},{"comment":"Native-scale conversions (Δ* × reference SD) are conversions of the standardized estimate, not native-scale refits (M5). Flag this more visibly in Fig. 3’s right axis and in the insomnia MCID comparison so readers do not treat −3.5 min as an independent native-scale analysis.","section":"Fig. 3 / M5"},{"comment":"Table 1’s AACT audit ranges vs fitted registry rates (especially antidepressants 0.79 vs 0.59–0.74 lower bound) are carefully footnoted; a one-sentence reminder in the main Results that conclusions are stable across the audited range (already in S6–S7) would help non-specialist readers.","section":"Table 1 / M4"},{"comment":"The education and exercise applications (S13) are useful specificity/stress checks but are buried. A brief pointer in the main Discussion that the method manufactures little correction in near-complete literatures and can move away from zero (blood pressure) would strengthen the “not a blanket shrink-to-zero” claim already made for ACE inhibitors.","section":"Discussion / S13"},{"comment":"Notation: Z_k = (C_k − μ_C^pop)/σ_C is introduced in M1; ensure the main-text first mention of the instrument uses the same standardization as the first-stage F regression so F is reproducible from the text alone.","section":"M1 / Results"},{"comment":"References 20–41 and the full simulation tables (S1, S3–S5, S12) are thorough; a short “what conventional correctors miss and why” paragraph already in Results could cite Table S16 more directly for readers who skip the SI.","section":"Results / Table S16"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is written in a high-impact general-science style (one-sentence summary, boxes, multi-literature applications) while the technical core is solid stat.ME. That is a feature if the journal wants methods that change practice; if the venue is a pure methods journal, the editors may want a slightly tighter main-text focus on identification and the holdout, with clinical applications compressed. I see no circularity or hidden failure mode beyond the exclusion restriction the authors already probe. The replication package and strict holdout are unusually strong for this literature and should weigh heavily in the decision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing to know is that they give meta-analysis an external untreated-outcome distribution that does two jobs at once: it defines the target population (so the estimand is no longer “average over published trials”) and it anchors the control-arm mean as a shadow instrument for selection. That is the real novelty relative to Hedges–Vevea, Copas, PET/PEESE, and RoBMA, all of which stay inside published geometry or assumed selection shapes.\n\nWhat they do well is concrete. Theorem 1 / Lemmas S1–S2 are standard completeness-plus-exclusion logic once you grant the three assumptions; the math is clean. The strict holdout on Cipriani is the right test: the 12 recovered-unpublished contrasts are removed from both the estimation sample and the anchor, and the published-only fit still lands Δ* = 0.22 with a CI that covers the all-trials benchmark 0.248. They are explicit that this recovers the aggregate inverse-selection-weighted mean, not individual suppressed trials. First-stage F is reported before any correction and correctly declines the sleep-quality outcome (F = 2.4). Simulations stay near-unbiased under matched probit, misspecified logit, and a step mechanism when the anchor is valid. Code and data are promised to reproduce everything from one command; that matters.\n\nThe soft spot is exactly the one the reader flagged: exclusion (C ⊥ R | Y, U). If absolute control severity still moves publication through unblocked channels (perceived importance, sponsor strategy, era), the selection profile is not cleanly identified. They run a Copas-type ρ_C grid on an interpretable odds scale and show no qualitative reversal, which is honest but not a proof that residual direct paths are zero in editorial practice. Absolute publication rates need registry calibration and are not identified from published data alone; they treat that correctly. Anchor construction for objective endpoints (TST, HbA1c) is firmer than for psychometric scales, and they gate on that.\n\nThis is for people who do evidence synthesis or clinical epidemiology and care about target transport as much as selection. It deserves a serious referee. I would engage with it and cite the identification strategy and the holdout design.","headline":"Solid identification paper: external untreated-outcome distribution as simultaneous target and shadow instrument for publication selection, with a real holdout that recovers the Cipriani all-trials benchmark.","tokens_in":31616,"tokens_out":547,"would_cite":true,"duration_ms":7115,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10","62F10","62N01"],"pacs":[],"model":"grok-4.5","headline":"An external untreated-outcome distribution lets a meta-analysis estimate the effect in a declared patient population and correct publication selection with the same object.","keywords":["meta-analysis","publication bias","target population","external anchor","shadow instrument","clinical trials","selection models"],"falsifier":"A literature in which registry-linked publication rates, stratified by control severity at fixed effect size and design covariates, show a large residual direct association between control level and publication, or a strict holdout that fails to cover a known all-trials benchmark when first-stage strength is high.","tokens_in":31578,"feed_emoji":"📊","tokens_out":745,"duration_ms":7135,"temperature":0.7,"pith_summary":"A standard meta-analysis averages the trials that were published, not the effect a clinician or policymaker needs for a defined patient group. This paper shows that one external reference—the distribution of the untreated outcome among trial-eligible patients—does two jobs at once: it names the target population, and it turns each trial’s control-arm mean into a shadow instrument that identifies how publication selected the published record. Under an exclusion restriction (control severity affects publication only through the reported effect and design covariates), the bias-corrected effect at the reference, called Δ*, is identified from published data alone; the absolute publication rate is pinned separately by registry calibration. In a strict holdout on an antidepressant literature with recovered unpublished trials, the method recovers the held-out all-trials benchmark without ever seeing the hidden trials. Applied to insomnia total sleep time and placebo-controlled HbA1c trials, the target-population effects are materially smaller than the published averages, while a first-stage strength check refuses to answer when the instrument is too weak. Evidence synthesis becomes a target-explicit estimate whose selection assumptions can be audited.","feed_headline":"External anchors unmask smaller target effects in published trials","feed_subtitle":"One untreated-outcome distribution names the patient population and corrects publication selection","key_machinery":"Reference-anchored meta-analysis: the external untreated-outcome distribution is both the target (fixing the control level at which Δ* is reported) and the anchor that makes each trial’s control-arm mean a shadow instrument for publication selection, identified under the exclusion restriction C ⊥ R | Y, U.","core_discovery":"Under exclusion, an external anchor of the control distribution, and relevance of the control mean for the reported effect, the bias-corrected outcome model and the relative selection profile are identified from published-only data. The target estimand Δ* is the bias-corrected effect evaluated at the external reference control level and averaged over the trial-eligible covariate distribution. A strict holdout that removes recovered unpublished antidepressant contrasts from both fitting and anchoring recovers an all-trials benchmark inside the uncertainty interval, and applications to insomnia and diabetes yield target effects substantially smaller than naïve published averages.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["External anchors unmask target effects smaller than published averages","Reference anchors correct selection bias for true population effects","Untreated-outcome anchors reveal hidden target effects in trials","Anchored meta-analysis yields smaller effects than published trials","Control means as anchors expose overstated effects in published data"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Absolute control-arm severity is assumed to affect whether a trial is published only through the reported effect and observed design covariates, not by a direct path such as perceived clinical importance or sponsor strategy.","fun_headline_variants_meta":{"raw":{"variants":["External anchors unmask target effects smaller than published averages","Reference anchors correct selection bias for true population effects","Untreated-outcome anchors reveal hidden target effects in trials","Anchored meta-analysis yields smaller effects than published trials","Control means as anchors expose overstated effects in published data"]},"model":"grok-4.5","effort":"low","cost_usd":0.008932,"raw_usage":{"total_tokens":2040,"prompt_tokens":731,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":89320000,"prompt_tokens_details":{"text_tokens":731,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1229,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":731,"tokens_out":80,"duration_ms":13007,"temperature":1.0,"reasoning_tokens":1229,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T20:02:11.116554+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A literature in which registry-linked publication rates, stratified by control severity at fixed effect size and design covariates, show a large residual direct association between control level and publication, or a strict holdout that fails to cover a known all-trials benchmark when first-stage strength is high.","supporting_citations":[],"review_version":1}