Pith. sign in

REVIEW 3 major objections 5 minor 1 references

Publication-bias tests miss real bias even with 1,000 studies — estimate it instead

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 12:11 UTC pith:XCEWCWT7

load-bearing objection Useful simulation paper arguing meta-analysts should estimate bias with confidence intervals rather than binary tests, but the headline coverage comparison for z-curve leans on the authors' own prior calibration and an assumption that selection is significance-based only. the 3 major comments →

arxiv 2607.19626 v1 pith:XCEWCWT7 submitted 2026-07-21 stat.AP

Beyond Publication-Bias Detection: Estimating Bias under Uncertainty in Heterogeneous Literatures

classification stat.AP
keywords publication biasmeta-analysisexpected discovery ratez-curveselection modelsconfidence interval coverageheterogeneitystatistical power
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper makes a two-part argument. First, common publication-bias tests have low power in heterogeneous research literatures: in a 5,184-condition simulation with up to 1,000 studies, no test averaged conventional power, and only conditions that exactly matched a model's assumptions came close. Second, the goal should shift to estimating how much bias is compatible with the data. Step-function selection models give confidence intervals that cover the true amount of bias only about half the time, while z-curve's expected-discovery-rate intervals cover it about 97% of the time. Applied to social priming and applied-psychology meta-analyses, the upper bounds of these intervals support worst-case reasoning that significance tests hide.

Core claim

The discovery is a calibration result plus a re-framing. On the same estimand — the publication probability of nonsignificant results, converted from the expected discovery rate (EDR) via the identity w = EDR(1 − ODR)/[ODR(1 − EDR)] — z-curve's bootstrapped confidence interval contains the true bias in about 97% of simulated runs, while the step-function selection model's interval contains it in only about half. The re-framing is that the upper bound of such an interval, not a p-value, is the information meta-analysts need: a wide upper bound shows that a large amount of bias remains compatible with the data even when the test is nonsignificant, and a tight upper bound can actually rule out

What carries the argument

The load-bearing quantity is the expected discovery rate (EDR), the average unconditional probability that a study in a heterogeneous set produces a significant result, estimated by z-curve from the distribution of significant z-values. The paper maps EDR to the selection-model weight via w = EDR(1 − ODR)/(ODR(1 − EDR)), putting z-curve and step-function models on the same estimand: the publication probability of nonsignificant results. Z-curve's bootstrap confidence interval around EDR, converted through this identity, produces the calibrated interval that carries the argument; the step-function model's weight estimate lacks this calibration under heterogeneity and graded selection.

Load-bearing premise

The load-bearing premise is that publication bias primarily operates through statistical significance, so the estimand 'missing nonsignificant results' captures real-world bias; the paper's own Limitations concede that selection among significant results can inflate the EDR estimate by up to 10 percentage points, and the headline coverage number also inherits a 5-point calibration from earlier fitting rather than being independently derived here.

What would settle it

Run the full 5,184-condition simulation on the missing-nonsignificant-results scale and check coverage. If step-function intervals cover the true bias near 95% in the full design, or z-curve intervals fall below 90%, the paper's key comparison fails. A complementary check on real data: in the social-priming reanalysis, if a selection model that allows effect-size-dependent publication produces an upper bound near the step-function model's 49% rather than z-curve's 86%, the claimed worst case depends on the selection assumption.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A nonsignificant publication-bias test should no longer be reported as evidence that bias is absent; with k = 1,000 studies the tests can miss 70% missing nonsignificant results.
  • Meta-analysts should report the upper bound of a bias interval; a wide upper bound means substantial bias remains compatible with the data even when the point estimate is small.
  • Step-function selection-model intervals understate bias uncertainty in heterogeneous literatures, so their use for reassurance can be actively misleading.
  • Sensitivity analyses can replace analyst-chosen selection weights with the lower bound of z-curve's EDR interval, grounding worst-case scenarios in the data.
  • In large, high-powered meta-analyses, a statistically significant bias test can detect only trivial amounts of bias; estimation distinguishes consequential from inconsequential selection.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If this result holds up, reporting guidelines for meta-analysis should ask authors to give a bias-interval upper bound rather than a p-value; the paper's logic extends naturally to any literature where heterogeneity is the norm.
  • The 10-percentage-point EDR inflation from selection among significant results (acknowledged in the Limitations) is the main hazard: if effect-size selection occurs, z-curve's upper bound could understate bias, and a targeted simulation with effect-size-dependent selection would test whether the coverage advantage survives.
  • A testable extension would apply this estimation framework to preregistered multi-lab replication sets, where selection is presumed absent; z-curve intervals should then concentrate near zero bias if the estimand is valid.
  • The same conversion identity could let other discovery-rate estimators be compared on the missing-studies scale, not just the EDR scale, making bias estimates across methods directly comparable.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper argues that publication-bias tests in meta-analysis have low power, especially in heterogeneous conceptual-replication literatures, and that the field should move from null-hypothesis testing to interval estimation of the amount of bias. The author reports a large factorial simulation (5,184 conditions) varying study number, effect size, heterogeneity, sample size, false-discovery rate, selection mechanism, and effect-size distribution, and finds that bias tests remain low-powered even with 1,000 studies unless the data match model assumptions. The paper then compares interval estimators: step-function selection models (Vevea--Woods / weightr) are reported to cover the true bias only about half the time, whereas z-curve's transformed EDR-based intervals are reported to achieve 97% coverage. Two applied illustrations (social priming and applied-psychology meta-analyses) are used to argue that upper confidence bounds are more informative than binary bias tests.

Significance. If the central claims hold, the paper would make a useful methodological contribution: it reframes publication-bias analysis from detection to estimation, provides a large simulation benchmark under realistic heterogeneity, and identifies a practical tool (z-curve) for interval estimation. The low-power result is credible and consistent with prior work (Hedges & Vevea 1996; Renkewitz & Keiner 2019), and the factorial design is extensive with code apparently available on OSF. However, the headline coverage claim for z-curve is not fully self-contained: it depends on a 5-percentage-point calibration made by the same authors in Bartoš & Schimmack (2022), and the paper's own Limitations section concedes a non-conservative failure mode under selection among significant results. These issues affect the strength of the paper's central recommendation and require attention before the claims can be regarded as fully supported.

major comments (3)
  1. [Estimation of Publication Bias / Figure 6] The headline claim that z-curve confidence intervals 'have nominal coverage' relies on a 5-percentage-point adjustment made in Bartoš & Schimmack (2022) to ensure nominal coverage in this same family of simulation scenarios. The present paper's 97% coverage figure is therefore partly a verification of prior calibration rather than an independent demonstration. Because the coverage comparison is the load-bearing evidence for preferring z-curve as an interval estimator, the manuscript should either report what the intervals' coverage would be without the inherited calibration, or clearly frame the result as 'coverage under the calibrated procedure previously validated by the authors' and discuss the limits of this self-reference. As written, the abstract's contrast between 'about half the time' and 'nominal coverage' overstates the independence of the evidence.
  2. [Limitations, p. 43] The paper concedes that selection among significant results—for example, reporting larger effects first—can inflate z-curve's EDR estimate by up to 10 percentage points (Pek et al., 2026; Schimmack & Soto, 2026). This is a non-conservative error for the paper's recommended use of the upper confidence bound: a higher EDR lowers ODR−EDR and lowers the upper bound, so the interval would understate the amount of bias. The coverage simulations reported in the paper do not include this mechanism, and the rebuttal that such selection is unlikely without offsetting QRPs is an empirical conjecture rather than a demonstrated robustness property. The manuscript should either simulate this mechanism and report coverage of the bias interval under it, or substantially soften the claim that z-curve intervals have nominal coverage in realistic heterogeneous literatures.
  3. [Estimation of Publication Bias, coverage-analysis inclusion rule] The coverage comparison retains only simulation runs with true bias between 40% and 80%, excluding 2% of simulations. Since the coverage estimate is the paper's central quantitative result, this exclusion rule needs justification and a sensitivity check. If the excluded runs have more extreme selection bias, the coverage of both methods could differ materially, especially for the step-function model whose intervals are reported to be systematically too narrow. Please report coverage with and without the exclusion, or at least describe the distribution of excluded cases and show that the conclusion is unchanged.
minor comments (5)
  1. [Abstract] The phrase 'I show that confidence intervals from step-function selection models are too narrow' would benefit from specifying the estimand (selection weight for nonsignificant results) and the simulation conditions where this holds; the current wording is easy to overgeneralize.
  2. [Figure 6] The figure caption does not state whether the plotted intervals are one-sided or two-sided, nor does it define 'upper bound' in the selection-weight metric. Please clarify the interval construction in the caption or text.
  3. [Table 1] The table reports rejection rates but no Monte Carlo uncertainty. With 5,184 cells and apparently one replication per cell, reporting standard errors or a statement about the number of replications per condition would help readers calibrate the precision of the power estimates.
  4. [Simulation Design] The caliper width is described as '20% of the criterion value, z = 1.96' but it is not immediately clear whether the interval is [1.96, 1.96+0.392] or symmetric around 1.96. Please define the interval endpoints explicitly.
  5. [Discussion] The paragraph contrasting confidence intervals with sensitivity analyses would be clearer if it acknowledged that a confidence interval is only as assumption-free as the model that produces it; the distinction is rhetorical and could be sharpened.

Circularity Check

2 steps flagged

Headline z-curve coverage claim is partially pre-fitted rather than independently demonstrated; acknowledged effect-size selection among significant results is conceded to inflate EDR non-conservatively by up to 10pp, undermining the headline asymmetry claim.

specific steps
  1. fitted input called prediction [Z-Curve section (p.18) and Estimation of Publication Bias section (p.31); Figure 6]
    "Simulation studies showed that z-curve 2.0 can estimate the EDR with systematic bias in some specific situations. To address this problem the confidence intervals were adjusted by 5 percentage points to ensure nominal coverage in all scenarios (Bartoš & Schimmack, 2022)."

    The paper's headline result is that z-curve's EDR-based confidence intervals have 'nominal coverage' (97% in Figure 6). But the interval procedure was previously calibrated—'adjusted by 5 percentage points to ensure nominal coverage'—on the same family of factorial designs described in the paper ('the present study used a simulation design previously used to validate confidence-interval coverage for z-curve’s EDR estimates'). The new 97% coverage is therefore a re-verification of a procedure tuned to achieve nominal coverage, not an independent out-of-sample demonstration. This is one of several inputs to the central claim that is inherited by self-calibration.

  2. other [Limitations (pp.43-44)]
    "A more serious concern, because it operates in the opposite direction, is selection that leads z-curve to overestimate the EDR and thereby underestimate bias. This can occur when significant results with larger effect sizes are reported in preference to those with smaller effect sizes. One simulation study suggested that this bias may inflate EDR estimates up to 10 percentage points (Pek et al., 2026; Schimmack & Soto, 2026)."

    The load-bearing comparative claim is that z-curve intervals have nominal coverage while step-function intervals cover only ~50%. This coverage is established only under selection mechanisms that depend on statistical significance (step and graded selection dropping nonsignificant results). The paper concedes in its own Limitations that selection among significant results—which the simulation design never varies—can inflate EDR by up to 10 percentage points. Because the bias estimand is ODR − EDR, a 10pp EDR inflation shrinks the estimated bias and its upper bound, so the interval would understate the true amount of bias under that acknowledged mechanism. The paper's rebuttal that such effect-size selection 'is unlikely to occur without questionable research practices' that 'have the oppos

full rationale

The central contribution is the shift from bias detection to bias estimation with calibrated confidence intervals, with the headline result that z-curve's intervals achieve nominal coverage (97%) while step-function selection-model intervals cover only about half the time (Figure 6). The step-function undercoverage finding is self-contained and credible: the full-text simulation design (5,184 conditions, with step and graded selection, non-normal effect distributions) is described in enough detail that the reduction is transparent, and the 50% coverage is presented as a direct simulation output. However, two parts of the argument reduce to inherited inputs rather than independent derivation. First, the paper states explicitly that z-curve's confidence intervals 'were adjusted by 5 percentage points to ensure nominal coverage in all scenarios' in the authors' own prior work (Bartoš & Schimmack, 2022), and that the present simulation design is the same one used for that prior validation. The reported 97% coverage is therefore a check on an interval procedure that was tuned to achieve near-nominal coverage on that same family of scenarios; this is partial circularity in the sense of pattern 2 (fitted input called prediction), because the headline calibration claim inherits the fit rather than being newly established. Second, the paper's own Limitations section concedes that selection among significant results (preferentially reporting larger effects) can inflate EDR estimates by up to 10 percentage points, citing two papers (Pek et al., 2026; Schimmack & Soto, 2026). Under the paper's own estimand, a 10pp EDR inflation would non-conservatively lower the estimated bias and the upper bound. No simulation under that mechanism is reported; the response that effect-size selection is unlikely without offsetting QRPs is an unquantified empirical conjecture. This conceded mechanism is not a circular derivation, but it is a load-bearing gap that is acknowledged in-text and weighs against the 97%-coverage claim being a robust independent result. Self-citations appear throughout (Bartoš & Schimmack, 2022, for the calibration; Brunner & Schimmack, 2020; Schimmack & Soto, 2026, for the effect-size-selection simulation), and the calibration citation is load-bearing, though not fully circular because the paper does run new simulation cells and reports a factorial undercoverage result for step-function models that does not depend on the authors' prior calibration. A fair cir

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

The central claims rest on simulation design choices and on the assumption that selection acts on significance. No new theoretical entities are postulated; EDR and selection weights are re-used from prior literature.

free parameters (3)
  • 5-percentage-point z-curve CI calibration = +/-5 percentage points applied to EDR confidence interval
    The paper states z-curve's confidence intervals were adjusted by 5 percentage points to ensure nominal coverage in all scenarios (Bartoš & Schimmack, 2022). The headline coverage claim inherits this hand-calibrated constant.
  • Caliper test interval width = 20% of z=1.96 (about 0.392 z-units)
    The caliper width was chosen as the widest interval from prior simulation work; this tuning choice directly affects the caliper test's power and Type I error.
  • Coverage-analysis inclusion rule = Retain only simulations with true bias between 40% and 80%
    The coverage comparison excludes 2% of runs outside this bias range, making reported coverage conditional on the target estimand rather than unconditional.
axioms (5)
  • domain assumption Selection operates primarily on statistical significance: significant results are observed, nonsignificant results may be missing.
    Location: Z-Curve section and Limitations (p.43). The entire bias estimand is defined as missing nonsignificant results; effect-size selection or QRPs can violate this.
  • domain assumption The distribution of significant z-values identifies the expected discovery rate (EDR) via a finite mixture model.
    Location: Z-Curve section. The paper relies on this identifiability without proving it; it is inherited from prior z-curve work.
  • domain assumption Latent effect-size distributions for step-function selection models are approximately normal; the simulation's beta and truncated-normal distributions are the relevant misspecifications.
    Location: Selection of Methods, pp.13-14. The step-function model's poor coverage under graded selection is interpreted as assumption violation, which presumes the simulation's non-normal/graded mechanisms are realistic.
  • standard math Monotone transformation of a calibrated EDR confidence interval preserves coverage, and bootstrap resampling of z-curve D-statistics is valid.
    Location: Estimation of Publication Bias. The conversion w = EDR(1-ODR)/(ODR(1-EDR)) is monotone in EDR; coverage transfer follows only if the base interval and bootstrap are valid.
  • domain assumption The applied datasets (Dai et al. social priming; Siegel et al. applied psychology) are accurately extracted and the clustered bootstrap/weightr reanalyses are correctly specified.
    Location: Illustrations 1 and 2. Exclusions such as six effect sizes above d=2.5 and the one-tailed p=.50 step are analyst choices; data extraction errors would change the illustrative conclusions.

pith-pipeline@v1.3.0-alltime-deepseek · 20026 in / 14819 out tokens · 156571 ms · 2026-08-01T12:11:08.141819+00:00 · methodology

0 comments
read the original abstract

Meta-analysts routinely test for publication bias, but a nonsignificant test is inconclusive because it may be a Type 2 error. In a large factorial simulation, I show that publication-bias tests have low power even with 1,000 studies once effect sizes are heterogeneous and the data do not meet the model's assumptions. I argue that the goal should shift from detecting bias to estimating it with confidence intervals. A confidence interval does not only bound the hypothesis that bias is absent; its upper bound also indicates how much bias remains compatible with the data. When the interval is wide, the upper bound does not exclude a large amount of bias, even if the formal test is nonsignificant. I then show that confidence intervals from step-function selection models are too narrow, covering the true amount of bias only about half the time. In contrast, z-curve's confidence interval obtained by transforming the expected discovery rate into a selection weight parameter has nominal coverage. Z-curve is therefore a valuable tool for examining publication bias in meta-analyses, especially when heterogeneity is high. I illustrate the implications of bias estimation with a meta-analysis of social priming and a set of applied-psychology meta-analyses.

Figures

Figures reproduced from arXiv: 2607.19626 by Ulrich Schimmack.

Figure 8
Figure 8. Figure 8: Z-curve plots of two meta-analyses from Siegel et al. (2022). Bias was detected because both meta-analyses were relatively large (k > 100) and most studies had high power, making it possible to detect the absence of a few nonsignificant results. However, the amount of bias is small and had no practical implications for the estimation of the average effect size. These results further underscore the problem … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

1 extracted references

  1. [376]

    discoveries

    https://doi.org/10.1037/cap0000246 PUBLICATION BIAS 49 Schimmack, U., & Bartoš, F. (2023). Estimating the false discovery risk of (randomized) clinical trials in medical journals based on published p-values. PLOS ONE, 18(8), Article e0290084. https://doi.org/10.1371/journal.pone.0290084 Schimmack, U., Heene, M., & Kesavan, K. (2017, February 2). Reconstru...