{"id":"fc9c1f6c-8b74-4e7f-b31c-fee8f6eb8cbe","arxiv_id":"2607.23174","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"Across 358 behavioral-science meta-analyses, outlier-handling treatments changed the significance verdict in 11.5% and the smallest-effect-size verdict in 15.9% of meta-analyses, while leaving the mean effect almost unchanged.","lead":"This study compared four ways of handling outlier results across 358 meta-analyses and found the pooled effect size barely moves, but the statistical verdict flips in a notable share of cases. It gives applied meta-analysts and reviewers a quantitative reference for how much this under-reported methodological choice matters.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Baseline-integrity is the key load-bearing assumption; because its likely effect is conservative, the verdict stands, but a direct test would settle it.","rationale":"The paper is carefully designed, pre-registered, and transparent, with public data and code. The numerical results in Tables 2 and 3 are internally consistent with the reported reversal rates, and the robustness checks (k≥20, k≥5, alternative cutoffs, one-estimate-per-study) bracket the main findings without contradicting them. The baseline-integrity concern is real and matches the reader's weakest-assumption diagnosis: if the BEAR v2 collection already contains meta-analyses whose compilers removed outliers, then the 'do-nothing' baseline is not truly unhandled. The likely direction of this bias is toward understating the sensitivity of meta-analytic conclusions, which would not weaken the paper's central qualitative claim and could strengthen it. The authors explicitly acknowledge the limitation in Section 5.2 and provide some evidence (the retained |d|>10 estimates) that the psychology source was not aggressively cleaned. No internal inconsistency or computational error was found that would threaten the ACCEPT verdict. The proposed concrete test—applying the pipeline to an independently known unprocessed sample—would directly resolve whether baseline contamination materially changes the headline rates. Until such a test is performed, the existing disclosure and conservative-direction argument are sufficient to maintain the reader's verdict.","tokens_in":14790,"tokens_out":11872,"duration_ms":125773,"concrete_test":"Run the identical pre-registered pipeline on an independent collection of meta-analyses whose raw study-level data are known to be unprocessed before any exclusions (for example, a registered-report dataset or a repository archiving all collected estimates prior to outlier decisions). Compare the at-least-one treatment reversal rates for statistical significance and SESOI with the reported cell-level rates (7.7% and 10.3%) and meta-analysis-level rates (11.5% and 15.9%). If the rates in the unprocessed sample are materially higher, baseline contamination was understating sensitivity; if similar or lower, the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Every reported reversal rate is a difference from the 'do-nothing' baseline. Section 5.2 concedes that the BEAR v2 collection 'may contain meta-analyses that have already omitted some outliers,' so the baseline itself may already embed prior outlier-handling decisions. If source compilers removed or corrected extreme estimates, then the treatment-versus-baseline contrasts do not cleanly isolate the effect of outlier handling from the effect of prior cleaning. The authors argue that any such contamination would make their estimates understate true sensitivity, and the retention of 37 estimates with |d|>10 in the psychology source supports that reading. However, the direction is not logically guaranteed: if prior cleaning selectively removed estimates that were propping up significance, the remaining baseline could be more fragile, and the reported 11.5% SS / 15.9% SESOI meta-analysis-level reversal rates could overstate the sensitivity of a truly unhandled synthesis. Because all headline percentages are computed against this baseline, this is the most load-bearing assumption in the paper. It is disclosed and partially mitigated, but it is not directly tested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper quantifies how pre-registered outlier and influence-handling decisions affect meta-analytic conclusions in 358 behavioral-science meta-analyses (k≥10; 16,664 estimates) from the BEAR v2 collection spanning psychology, psychotherapy, and exercise. Four active treatments—dropping the most extreme |d|, deleting estimates with studentized deleted residuals >3, winsorizing at the 5th/95th percentiles, and deleting estimates with |DFBETAS|>2/√k—are compared against a do-nothing baseline under two estimators: random effects (REML with Hartung–Knapp adjustment) and unrestricted weighted least squares (UWLS). Outcomes are absolute change in pooled d, statistical significance (p<.05), and whether |pooled d| reaches a SESOI of ≥0.20. The median absolute change in d is at most 0.047 across treatments. At least one treatment changes statistical significance in 7.7% of 715 estimable estimator-by-meta-analysis cells (11.5% of 358 meta-analyses) and changes SESOI status in 10.3% of cells (15.9% of meta-analyses). Reversals are largely confined to results near the decision boundary; DFBETAS is the most sensitive and winsorizing the least. Robustness checks (k≥20, k≥5, alternative cutoffs, one-estimate-per-study) are consistent. The entire pipeline was pre-registered, the dataset is frozen, and a replication package is provided.","tokens_in":15049,"tokens_out":12581,"duration_ms":112598,"significance":"If correct, this is the first large-scale empirical reference for the sensitivity of meta-analytic significance and substantive-size verdicts to outlier handling, an under-reported researcher degree of freedom. The paper's strengths are its design: a pre-analysis plan with fixed thresholds and edge rules, a frozen dataset with checksum, and a publicly archived replication package that makes the confirmatory results push-button reproducible. The comparison to Langan et al.'s heterogeneity-estimator discordance (7.5%) provides a useful benchmark. The descriptive nature is appropriate; the authors do not claim to identify which treatment is correct. Limitations—possible prior cleaning in the source databases, repeated d/SE rows of unknown provenance in psychology, and the conflation of small-study effects with true outliers—are acknowledged and partially mitigated by the retention of extreme values (|d|>10) in the psychology source. The findings should be of interest to applied meta-analysts, methods researchers, and journal reviewers.","major_comments":[],"minor_comments":[{"comment":"The statement that prior cleaning in the source databases would make the treatment effects \"understate how much these choices matter\" is plausible but not logically guaranteed: if earlier cleaning selectively removed estimates that were propping up significance, the remaining baseline could be more fragile, and the reported reversal rates could overstate sensitivity. The retention of 37 estimates with |d|>10 supports the conservative-direction reading, but a sentence acknowledging the alternative direction and noting that a direct test would require access to source pre-processing would be more precise.","section":"§5.2, Limitations"},{"comment":"The column \"Not estimable\" in Table 4 is not defined in the table notes. The text in §4.4 explains that these arise from metafor REML non-convergence, but the table note should carry this definition for readers who start with the tables.","section":"Table 4, notes"},{"comment":"The SESOI threshold c=0.20 is based on Cohen's small effect for d. Because the psychology studies are converted from Fisher-z correlations, it may be worth adding a sentence clarifying whether the threshold is intended to be directly comparable across the three fields, which have different original metrics.","section":"§3.3, Outcomes"}],"recommendation":"minor_revision","confidential_remarks":"The baseline-integrity concern is the main substantive caveat. I regard it as adequately disclosed and the authors' argument that it likely understates sensitivity is plausible given the retained extreme values. The paper is a strong, reproducible meta-scientific contribution. I would not block acceptance on this caveat; the minor comments are local clarifications."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the paper the meta-analysis methods literature has been missing—a large, pre-registered, push-button-replicable survey of how outlier-handling choices change conclusions. I'd send it to a serious referee; my own verdict is accept with minor revisions.\n\nWhat's new: prior empirical work (Langan et al., Kontopantelis et al.) looked at heterogeneity estimators; nobody has systematically compared outlier treatments across hundreds of meta-analyses. The design is clean: five pre-registered treatments (including do-nothing), two estimators, three outcomes, fixed data with checksum, archived code. The main findings are credible and useful: pooled estimates barely move (median |Δd| ≤ 0.047) but categorical verdicts flip in 7.7% (SS) and 10.3% (SESOI) of cells; at the meta-analysis level, 11.5% and 15.9%. The ranking of treatments (winsorize least, DFBETAS most) is sensible and robust to the registered sensitivity checks.\n\nSoft spots, in order of importance:\n\n1. The baseline-integrity assumption. Every reversal rate is a contrast against the 'do-nothing' baseline, but the BEAR v2 collection may already contain source-level cleaning. The authors disclose this and argue that contamination would understate sensitivity. That argument is plausible but not logically guaranteed; prior cleaning could, in principle, have removed propping-up outliers, making the baseline fragile and inflating reversal rates. The 37 retained |d|>10 estimates show at least the psychology source wasn't aggressively scrubbed, which is reassuring. Still, a direct test—e.g., splitting the sample by evidence of prior cleaning—would settle it. This is a moderate caveat, not a fatal flaw.\n\n2. No confidence intervals on the reversal rates. With 358 meta-analyses, the 11.5% and 15.9% rates have meaningful sampling error; a bootstrap or exact binomial interval would help readers calibrate.\n\n3. Generalizability. The sample is behavioral science, k≥10, dominated by psychology (259/358). The paper is appropriately careful not to overclaim. The k≥5 and k≥20 robustness checks help, but the pattern (more small-k sensitivity) is expected.\n\nThe paper is honest, transparent, and the code and data are public. The authors' practical recommendations (pre-register, report both with and without treatment) follow from the evidence. I'd take the reversal rates as reference benchmarks for behavioral science, not universal constants.\n\nBottom line: deserves a serious referee, not a desk reject. I'd engage with it and cite it.","headline":"Solid, pre-registered empirical benchmark: first large-scale comparison of outlier-handling treatments in meta-analysis; the one soft spot (baseline integrity) is disclosed and likely conservative.","tokens_in":15517,"tokens_out":2605,"would_cite":true,"duration_ms":26373,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Outlier-handling choices flip the verdict in 11.5% of behavioral science meta-analyses.","keywords":["meta-analysis","outlier handling","influence diagnostics","researcher degrees of freedom","pre-registration","sensitivity analysis","smallest effect size of interest","behavioral science"],"falsifier":"One finding that would settle it: re-run the pre-registered treatments on a set of meta-analyses where raw primary-study data are available and the original authors' outlier removals are documented, and build the do-nothing baseline from the unhandled raw estimates. If reversal rates fall to near zero, the reported 11.5% and 15.9% figures are an artifact of baselines that already incorporate outlier handling. Conversely, if reversals were found concentrated among strongly significant or large effects—which the paper claims essentially never change—the central claim would also be contradicted.","tokens_in":14686,"feed_emoji":"🔄","tokens_out":6918,"duration_ms":62774,"temperature":0.7,"pith_summary":"This paper asks whether a meta-analyst's decision about what to do with extreme or highly influential estimates changes the conclusions of research syntheses. Across 358 behavioral science meta-analyses, four pre-registered outlier treatments barely move the pooled effect size—the median absolute change in Cohen's d is at most 0.047—yet at least one treatment reverses the statistical-significance verdict in 11.5% of the meta-analyses and the smallest-effect-of-interest (|d| ≥ 0.20) verdict in 15.9%. The reversals concentrate almost entirely among results already near the decision boundary; strongly significant results essentially never change. The authors read this as evidence that outlier handling is a researcher degree of freedom, and argue that pre-registering the exact handling rule is a practical safeguard.","feed_headline":"11.5% of meta-analyses flip when outliers are handled","feed_subtitle":"The pooled effect barely moves, yet significance and smallest-effect verdicts change—pre-registering outlier rules is key.","key_machinery":"The comparison machinery is a pre-registered grid: each of 358 meta-analyses is estimated under two estimators—random-effects with a small-sample t correction, and unrestricted weighted least squares (an inverse-variance weighted mean with heterogeneity-inflated standard errors)—and each is compared against a 'do nothing' baseline on four treatments: dropping the single most extreme effect, deleting estimates with studentized deleted residuals above 3, winsorizing at the 5th and 95th percentiles, and deleting estimates with |DFBETAS| above 2/√k. Treating one handling choice at a time while holding thresholds and edge rules fixed makes any observed change attributable to the handling rule. Th","core_discovery":"The central claim is that in behavioral science meta-analyses with at least ten estimates, outlier and influence handling decisions—dropping the most extreme estimate, deleting estimates flagged by studentized deleted residuals or DFBETAS, or winsorizing the tails—have little effect on the magnitude of the pooled mean (median |Δd| ≤ 0.047), but a meaningful effect on categorical interpretation. Counting either of two estimators and any of the four treatments, 7.7% of the 715 estimable estimator-by-meta-analysis cells change statistical significance and 10.3% change smallest-effect-of-interest status; at the level of meta-analyses the rates are 11.5% and 15.9%. The authors emphasize that wins","pith_inferences":["The reversal rates are measured against a 'do nothing' baseline taken from compiled meta-analyses that may already have had outliers removed by their original authors; if so, the true sensitivity of meta-analytic conclusions to outlier handling is understated, not overstated.","Because the sample is limited to behavioral science syntheses with at least ten effects, applying the same pre-registered pipeline to medical or other disciplinary meta-analyses, where effect distributions and typical k differ, could meaningfully change the rates; that is a testable extension of the design.","The fact that winsorizing only ever moves results toward statistical significance suggests a caution: a method that shrinks tails without asking why an estimate is extreme could make a synthesis look more conclusive than the data warrant, especially in areas with publication bias.","A direct practical test of this design's value: for a set of meta-analyses with fully documented raw data, compare the conclusions reached by teams that pre-registered versus teams that did not, to see whether pre-specification actually reduces outcome-dependent outlier handling."],"forward_implications":["Meta-analysts should pre-specify their outlier handling rule: a defensible choice of method can change the headline verdict in about one in nine syntheses.","Reviewers should ask for sensitivity analyses that report the central result both with and without outlier treatment; borderline findings are where the risk of reversal lives.","Winsorizing is the least disruptive handling option and DFBETAS-based deletion the most; a meta-analyst wanting a conservative check can use winsorizing, while DFBETAS under unrestricted weighted least squares is the most sensitive diagnostic to report.","Because reversal rates rise when meta-analyses are small (k below 20) and fall when they are large, outlier handling deserves special attention in small syntheses.","These rates are a benchmark: a 7.7% cell-level and 11.5% meta-analysis-level significance-reversal rate is comparable in scale to the discordance previously documented for heterogeneity-estimator choice, so outlier handling is no smaller a degree of freedom."],"fun_headline_variants":["Outlier choices flip 1 in 9 meta-analysis verdicts","Meta-analysis means stable, but verdicts flip 11.5% of the time","How outliers are handled can flip significance in 11.5% of studies","Pre-register outlier rules: verdicts flip often, means barely move","Winsorizing least, DFBETAS most: outlier handling flips 11.5%"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The 'do nothing' baseline is assumed to reflect genuinely unhandled data; if the source meta-analyses had already removed or corrected extreme estimates, the measured reversal rates understate how much outlier handling matters.","fun_headline_variants_meta":{"raw":{"variants":["Outlier choices flip 1 in 9 meta-analysis verdicts","Meta-analysis means stable, but verdicts flip 11.5% of the time","How outliers are handled can flip significance in 11.5% of studies","Pre-register outlier rules: verdicts flip often, means barely move","Winsorizing least, DFBETAS most: outlier handling flips 11.5%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00023,"raw_usage":{"total_tokens":1365,"prompt_tokens":835,"completion_tokens":530,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":424}},"tokens_in":579,"tokens_out":530,"duration_ms":4638,"temperature":1.0,"reasoning_tokens":424,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T03:22:48.224623+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One finding that would settle it: re-run the pre-registered treatments on a set of meta-analyses where raw primary-study data are available and the original authors' outlier removals are documented, and build the do-nothing baseline from the unhandled raw estimates. If reversal rates fall to near zero, the reported 11.5% and 15.9% figures are an artifact of baselines that already incorporate outlier handling. Conversely, if reversals were found concentrated among strongly significant or large effects—which the paper claims essentially never change—the central claim would also be contradicted.","supporting_citations":[],"review_version":1}