REVIEW 3 major objections 5 minor 9 references
This paper argues that multiverse analysis is especially well suited to letting researchers abdicate responsibility for their conclusions and to manufacturing doubt, and proposes judging a multiverse by its worst universe.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Multiverse analysis is argued to encourage abdication of researcher responsibility and to be suited for manufacturing doubt, with a 2025 literature review as tentative support.
T0 review reviewed 2026-08-02 challenge →
load-bearing objection A clear, honest normative critique of multiverse analysis whose empirical evidence is thin but whose proposed conventions are worth debating. the 3 major comments →
Multiverse analysis, abdication of responsibility and manufacturing of doubt
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that the very features that make multiverse analysis attractive—comprehensiveness, size, and the appearance of agnosticism—are the features that make it a ready tool for abdicating scientific responsibility and manufacturing doubt. Based on a hand-coded review of 31 papers with original multiverse analyses published in 2025, the author finds that 45% rely exclusively on probabilistic interpretation, 52% include at least one visualization where results cannot be traced to specific universes, and only 23% attempt any post-hoc evaluation of which universes are better. The author concludes with two proposed conventions: a multiverse should be judged by its worst univ
What carries the argument
The central evaluative device is the worst-universe convention: unless proven otherwise, a critique of any single analysis in a multiverse counts as a critique of the whole, so the multiverse's quality is capped by its weakest member. The paper pairs this with a 'large-is-weak' heuristic: very large multiverses signal either an immature research question or abdicated judgment. The argument also leans on a distinction between probabilistic interpretation (e.g., 'most models show ...'), which requires unjustified assumptions of equal correctness and exchangeability, and possibilistic interpretation (an outcome is possible if any universe produces it), which is weaker and safer.
Load-bearing premise
The argument treats a large multiverse as evidence of abdication, which assumes that most research questions have only a small number of genuinely defensible analytic choices; if many independent, reasonable choices legitimately exist, a large multiverse could be responsible rather than evasive.
What would settle it
Audit all published multiverse analyses with 1,000 or more universes for three markers: pre-registration of inclusion criteria, per-universe data-fit diagnostics, and full traceability of every displayed result to its defining choices. If a substantial share of large multiverses pass all three markers, the claim that large size typically signals abdication would be undermined.
If this is right
- If adopted, the worst-universe norm shifts the burden of proof to multiverse authors, who must defend every included universe rather than let the set as a whole lend false credibility to weak members.
- Large multiverses (1,000 or more universes) would be treated as a red flag, prompting reviewers to ask whether the research question is mature enough to be answered or whether authors have dodged the task of choosing defensible specifications.
- Probabilistic statements such as 'most models show ...' would be discouraged, because they presuppose that all universes are equally likely to be correct—an assumption the paper argues is rarely warranted.
- Reflective multiverses remain useful for contextualizing prior results, but even there the paper argues that untraceable visualizations and probabilistic summaries are undesirable.
- The review's numbers provide a baseline: roughly one in four studies attempts to evaluate which universes are better, and over half include at least one display where readers cannot connect results to specific universes.
Where Pith is reading between the lines
- The worst-universe rule could generalize beyond multiverse analysis to any specification-curve or sensitivity analysis, implying that robustness claims should be judged by the weakest defensible alternative specification rather than by the central estimate.
- The 'large multiverse as red flag' claim rests on an unstated assumption that genuinely defensible analytic choices are few; in fields where many independent, reasonable decisions legitimately exist, the rule would need to be qualified by requiring authors to justify each choice rather than simply counting them.
- A testable extension would compare rates of post-hoc universe evaluation and traceability before and after the proposed conventions gain traction, to see whether social norms actually change analytic practice.
- If probabilistic interpretation is discouraged, researchers will need new tools for aggregating multiverse results that do not assume exchangeability—for example, meta-analytic or decision-theoretic summaries that explicitly weight universes by evidence.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that multiverse analysis, while useful in some applications, is particularly suited to two negative functions: abdication of analytic responsibility and manufacturing of doubt. To support this, the author conducts a small literature review of 31 original multiverse analyses published in 2025, reporting frequent probabilistic interpretation (45.2% exclusive), untraceable visualizations (25.8% no trace), and infrequent post-hoc evaluation of universes (22.6%). The paper then proposes two evaluation conventions: judge a multiverse by its worst universe, and treat large multiverse size as a red flag for immaturity or abdication. It also discourages probabilistic interpretation of multiverse results. The empirical evidence is explicitly described as tentative, and the author repeatedly acknowledges the subjectivity of the coding.
Significance. If the conceptual argument is accepted, the paper is a useful corrective to the uncritical enthusiasm surrounding multiverse analysis. The proposed 'worst universe' convention is a clear, actionable policy that would shift the burden of justification to authors of multiverse studies, and the paper honestly discloses the limitations of its empirical review. The strength of the paper lies in its structured argument and accessible summary of problematic practices in the 2025 literature. The main weakness is that the 'large multiverse as red flag' recommendation depends on an unsecured assumption that genuinely defensible analytic choices are typically few, and the empirical review does not measure defensibility or compare with non-multiverse practice. The paper is a worthwhile contribution to the methods debate if the load-bearing assumptions are made explicit and the recommendations are appropriately bounded.
major comments (3)
- [Section 5] The recommendation that 'large multiverses as a red flag' relies on the premise that a typical research question has only a small number of genuinely defensible analytic choices. This premise is asserted, not argued or tested. A multiverse of e.g. 1000 can arise from a few independent binary choices that are all reasonable, if the question genuinely has many dimensions of preprocessing and modeling uncertainty. Without a measure of defensibility or an empirical distribution of choice spaces, size alone does not signal abdication. The review in §4.2 reports multiverse sizes but never relates size to the fraction of indefensible universes. Please either supply a theoretical argument for small choice spaces, provide data showing that large multiverses contain more indefensible universes, or weaken the claim to a conditional heuristic.
- [Section 4.1, Table 1] The empirical support for 'abdication of responsibility is present' is based on 31 papers with coding criteria that the author states 'evolved' during reading and were not preregistered. The absence of a baseline comparison to non-multiverse papers means that the rates in Table 1 (e.g., 22.6% post-hoc evaluation, 25.8% no trace) do not establish that multiverse analyses are more prone to abdication than ordinary analyses. The paper's hedging is commendable, but the numerical summaries are still presented as evidence. Please either recast the review as a purely illustrative case collection, or strengthen the methods (e.g., prespecified coding, inter-rater reliability, a comparison group).
- [Section 4.3] The manufacturing-doubt claim rests on a single anecdote in which the author's inference about the usefulness of the multiverse for 'sowing doubt' is accompanied by biographical details about Bendavid's public positions. The manuscript correctly says intentions cannot be inferred, but the biographical information is irrelevant to the methodological critique and risks making the argument appear ad hominem. If kept, the point about the 'three quarters of universes biased to find null' should stand alone; otherwise, the anecdote should be cut or stripped of personal details. As written, this section is the least rigorous part of the paper and invites valid criticism that distracts from the structural argument in Section 3.
minor comments (5)
- [Title/Abstract] The variant 'analyzes' appears where 'analyses' is intended (both in the title and abstract). Please correct.
- [Section 3] Typo: 'they may ran a large multiverse' should be 'they may run'. Also 'alwasy' in Section 4.3 should be 'always'.
- [Section 1, author affiliation] The affiliation line contains 'V ´Uvalu' with a stray accent; should be 'V Úvalu' or proper Czech formatting.
- [Section 4.1] The phrase 'for all results where I could obtain the full text' is not quantified in the results; consider reporting how many of the 31 had full-text access.
- [Figure 1] The x-axis labels use non-standard bin notation (e.g., '1 − 10' rather than '1–10'). Consider unifying and ensuring bins do not overlap at boundaries (e.g., 10 appears in both first and second bin).
Circularity Check
No significant circularity: the paper's claims are normative/empirical arguments with disclosed measurement limitations, not derivations that reduce to their inputs.
full rationale
The paper's derivation chain is not a formal one. The central claims—that multiverse analysis can enable abdication of responsibility and manufacturing of doubt, and that the community should adopt worst-universe evaluation and treat large multiverses as a red flag—are presented as normative recommendations supported by a small, explicitly tentative literature review. The empirical section codes 31 published multiverse analyses using stated proxies (e.g., probabilistic interpretation, traceability of visualizations, post-hoc evaluation of universes). The coding is subjective and not preregistered, and the author acknowledges this ('The search and the specific criteria were not preregistered'; 'the low number of papers and subjectivity of the classifications is also a reason to not overly generalize the conclusions'). That is a validity limitation, not circularity: the coded variables are not defined in terms of the paper's conclusions, and the conclusions are not fitted values drawn from the coded data. The recommendation about the 'worst universe' is a proposed convention, not a result derived from a definition. The large-multiverse red flag rests on an unproven assumption about how many defensible analytic choices a typical question has, but that is an unsupported empirical/normative premise, not a circular reduction. The only self-reference is the author's statement that he has used multiverse analysis in his own published work; it is anecdotal and not load-bearing. No equation, fitted parameter, or unique solution is relabeled as a prediction, and no self-citation chain is used to forbid alternatives. The paper is self-contained in the sense relevant to circularity analysis, and the honest finding is no significant circularity.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Probabilistic interpretation of a multiverse requires that included universes are exchangeable and equally likely to be correct.
- ad hoc to paper If a single universe in a multiverse is invalid, the whole multiverse should be judged no higher than that universe unless authors justify it.
- ad hoc to paper A large multiverse indicates either an immature research question or abdication of responsibility.
- domain assumption The 2025 OpenAlex sample and the author's manual coding are adequate to demonstrate that the problems exist in the literature.
- standard math Possibilistic logic: if any universe yields an outcome, the outcome cannot be ruled out.
Cite this review
Pith. "Pith review of Multiverse analysis, abdication of responsibility and manufacturing of doubt." pith.science (2026). https://pith.science/paper/UND2D3T5
@misc{pith2026260714623,
author = {Pith},
title = {Pith review of: Multiverse analysis, abdication of responsibility and manufacturing of doubt},
year = {2026},
howpublished = {\url{https://pith.science/paper/UND2D3T5}},
note = {Machine review of arXiv:2607.14623}
}
read the original abstract
I argue that multiverse analysis is highly suited to two undesirable uses: abdication of researcher's responsibility for their conclusion and manufacturing of doubt. A review of multiverse analyses published in 2025 provides tentative empirical support that abdication of responsibility is present in the literature and I mention anecdotal evidence that multiverse has been used for manufacturing of doubt about Covid-19 precautions. To mitigate negative effects if multiverse analysis becomes widely used I suggest the community adopts two conventions for evaluating multiverse analyzes: evaluating multiverses by the single worst universe they contain and considering large size of a multiverse as a sign of weakness rather than a praiseworthy achievement.
Reference graph
Works this paper leans on
-
[1]
What's a multiverse good for anyway?
Rohrer, Julia M and Hullman, Jessica and Gelman, Andrew. What's a multiverse good for anyway?. PsyArXiv
-
[2]
Fuck nuance
Healy, Kieran. Fuck nuance. Sociol. Theory
-
[3]
A survey of tasks and visualizations in multiverse analysis reports
Hall, Brian D and Liu, Yang and Jansen, Yvonne and Dragicevic, Pierre and Chevalier, Fanny and Kay, Matthew. A survey of tasks and visualizations in multiverse analysis reports. Comput. Graph. Forum
-
[4]
Epidemic outcomes following government responses to COVID-19 : Insights from nearly 100,000 models
Bendavid, Eran and Patel, Chirag J. Epidemic outcomes following government responses to COVID-19 : Insights from nearly 100,000 models. Sci. Adv
-
[5]
Is scientific reform an unwinnable arms race?
Munaf \`o , Marcus R and Davey Smith, George. Is scientific reform an unwinnable arms race?. PLoS Biol
-
[6]
2024 , eprint=
Analysis of Potential Biases and Validity of Studies Using Multiverse Approaches to Assess the Impacts of Government Responses to Epidemics , author=. 2024 , eprint=
2024
-
[7]
and Tuerlinckx, F
Steegen, S. and Tuerlinckx, F. and Gelman, A. and Vanpaemel, W. , year = 2016, title =. Perspectives on Psychological Science , volume = 11, pages =
2016
-
[8]
Robustness is better assessed with a few thoughtful models than with billions of regressions , journal =. 2025 , doi =. doi:10.1073/pnas.2521917122 , author =
-
[9]
A traveler's guide to the multiverse: Promises, pitfalls, and a framework for the evaluation of analytic decisions
Del Giudice, Marco and Gangestad, Steven W. A traveler's guide to the multiverse: Promises, pitfalls, and a framework for the evaluation of analytic decisions. Adv. Methods Pract. Psychol. Sci
This paper was first reviewed by deepseek-v4-flash on August 2, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.