REVIEW 3 major objections 5 minor 18 references
Enough?
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper argues that RCTs are statistically privileged for one estimand but are not 'enough' for generalizable science; their true value is ontological, creating novel states of the world.
desk verdict A fair, clear-headed commentary that lands the 'enough for what?' question but overstates its case by stipulating an extreme standard of sufficiency. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a four-part decomposition of the causal estimand into sample $S$, site $P$, realization of treatment $R$, and time $T$, each of which has a selection process that functions like a propensity score. The paper combines this decomposition with the notion of 'shoe leather'—design-based knowledge of selection probabilities—to show what an RCT does and does not solve: it gives agnostic inference on the realized cell $(s,p,r,t)$ but leaves the other cells open, and for $T$ shoe leather is impossible because future designs cannot receive selection probabilities before they exist. A second piece is the performative reading of experiments, in which running an RCT imposes a simplified ontology on the world and thereby makes the propensity score known by construction.
What would settle it
Demonstrate that treatment effects estimated in a single-site, single-time RCT reliably predict effects at other sites and times without any selection models or ignorability assumptions; repeated success in such out-of-sample prediction would falsify the claim that the scientific endeavor requires controlling $S$, $P$, $R$, $T$.
Extended reading notes
Core claim
The paper's central claim is that the argument it responds to—that nonparametric identification is not enough while randomized controlled trials are—is correct as far as estimation goes, yet incomplete as a statement about science. RCTs provide agnostic inference on the estimand $\mathbb{E}[\tau(S,P,R,T)\mid S=s,P=p,R=r,T=t]$ but say almost nothing about the other selections—which sample, which site, which realization of treatment, which time—that would be needed for generalizable knowledge. The authors therefore distinguish 'enough' as a statistical property from 'enough' as a property of a scientific enterprise, and they argue that on the latter standard no design is enough. The specialness of RCTs lies not in their epistemology but in their ontology: they perform a model of the world by creating the very state they study.
Load-bearing premise
The argument depends on the premise that scientific sufficiency requires agnostic control over generalization across sample, site, treatment realization, and time; if 'enough' is allowed to mean a well-estimated causal effect in one setting, the conclusion that RCTs are not enough does not follow.
Editorial extensions
If this is right
- A successful RCT in one site and time cannot, by itself, support a generalizable scientific conclusion; extending its result requires either design-based control over site and time selection or explicit assumptions that those selections are ignorable.
- Calls to make RCTs the gold standard of causal inference are better read as sociological claims about the status of methods, not as consequences of statistical theorems.
- The practical importance of the earlier paper's critique depends on empirical facts about the complexity of naturally occurring propensity score functions—facts that do not yet exist and should be collected by metascience.
- Recommender systems' known propensity scores are a case of 'artificial' rather than 'natural' experiments: the scores are knowable only because the virtual world was engineered to be simple, and access to them is limited by corporate power.
- The slogan 'no causation without control' replaces 'no causation without manipulation' and points statisticians toward the political and social dimensions of experimental designs.
Reading between the lines
- A testable extension of the paper's logic would be to run multi-site, multi-time randomized trials with random site selection, measuring how much treatment effects vary across the $S$, $P$, $R$, $T$ dimensions.
- Implicit in the paper but not stated: if RCTs create novel states of the world, then choosing which possible worlds to create is a value-laden decision, adding an ethical dimension to experimental design that statistics alone cannot resolve.
- One operationalization of the call for empirical metascience would be a public registry of propensity score functions estimated in observational studies, with measures of complexity, to replace conflicting intuitions with data.
- The account also redirects gold-standard debates toward institutional questions: who gets to run experiments, who owns the resulting propensity scores, and whose questions get randomized.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a critical commentary on Aronow et al. (2025) (ARSSS), who argue that RCTs are 'enough' while nonparametric identification in observational studies is not. Dimmery and Munger agree with ARSSS on the statistical point about experimental versus observational design, but argue that 'enough' cannot be settled from within statistics because it is a sociological claim about the goals of science. They propose that scientific sufficiency would require agnostic, design-based control over at least four dimensions—units (S), sites (P), theory realizations (R), and time (T)—and argue that no empirical method can satisfy this for T, so nothing can be 'enough.' They call for empirical metascience to replace statisticians' intuitions about the complexity of naturally occurring propensity score functions, illustrate the limits of 'known' propensity scores with recommender systems, and conclude that RCTs are special for ontological rather than epistemological reasons: they create novel states of the world. The final answer to the title question is 'No' for the broad scientific sense of 'enough,' but 'Absolutely' for a special place for RCTs in the ongoing process of societal learning.
Significance. The paper's value is as a commentary that broadens the causal-inference debate: it distinguishes statistical sufficiency from scientific sufficiency, highlights the social and institutional context in which 'enough' is decided, and draws attention to temporal validity and to the lack of metascientific evidence about propensity-score complexity. The recommender-systems discussion and the call for empirical work on the distribution of propensity-score complexity are constructive, and the paper is transparent about the absence of data for its metascientific proposal. However, the central negative conclusion rests on a stipulated epistemic standard and is accompanied by an overstrong ontological claim about RCTs; as a result, the paper's contribution is conditional on accepting a particular view of scientific sufficiency rather than on a demonstrated argument against ARSSS's position.
major comments (3)
- [Section 3, estimand E[τ(S,P,R,T)|...]] The negative answer to the title question is load-bearing on the claim that the 'larger scientific endeavor ... requires control on (at least) these other source of errors' (Section 3). This standard—agnostic, design-based control over S, P, R, and T—is asserted rather than argued. The paper does not engage with the possibility that scientific knowledge could be sufficient through explicit model-based transportability assumptions or by changing the target estimand, in which case an RCT plus stated assumptions can be 'enough' for a specified question. Because the conclusion 'No' follows only under the stipulated standard, the authors should either defend that standard or explicitly frame the conclusion as conditional on it.
- [Section 3, temporal validity] The claim that 'shoe-leather cannot be a solution' for temporal validity because an experiment's selection probability is 0% until conceived is too strong. Design-based inference over T is logically possible whenever a sampling frame of times is specified and selection probabilities are known (for example, a predetermined random start date within a defined set of possible times); the difficulty is practical and institutional rather than logical. If the authors intend the stronger skeptical point about induction, that point applies equally to all empirical methods and makes 'nothing can be enough' trivially true; the paper should clarify why this does not undermine the comparative relevance of the RCT-specific critique.
- [Section 6, final paragraph] The positive account that RCTs 'create novel states of the world' and that 'since that state of the world is created by the experimenter, it is known to align with reality' overstates the experimenter's control. Randomization fixes the assignment mechanism, but noncompliance, implementation failure, missing outcomes, and downstream effects mean the realized state is not fully created or fully known. The claim should be qualified to 'successfully implemented' experimental assignments, or the ontological claim should be formulated more carefully; otherwise the paper risks replacing one exaggerated pedestal with another.
minor comments (5)
- [Abstract and full text, line 1] There is a typo in 'randomize d controlled trials' in the abstract; the same line appears broken in the full-text rendering.
- [Section 2] 'Lipshitz' should be 'Lipschitz'.
- [Section 3] The notation E[τ(S,P,R,T)|S=s, P=p, R=r, T=t] is informal; please define the probability space and clarify how the conditioning relates to the usual definition of an estimand.
- [Section 5] The example 'XUserThreshold03: 20%, XUserThreshold10: 80%' is not explained; a brief description of what these thresholds represent would help readers who are not familiar with recommender-system log data.
- [Section 6] The abstract says the authors agree with ARSSS 'with respect to experimental versus observational research,' while the conclusion answers 'No' to whether RCTs are enough; please reconcile these statements by specifying the exact scope of the agreement.
Circularity Check
No significant circularity: the paper is an explicitly normative and sociological commentary, not a derivation, and its self-citation is not load-bearing.
full rationale
This paper makes no statistical predictions, fits no parameters, and presents no derivation chain whose output could reduce to its inputs. Its central claim—that RCTs are not 'enough' for the broader scientific enterprise—is avowedly a sociological and philosophical argument about the meaning of 'enough,' not a result derived from the statistical argument of Aronow et al. (2025). The Section 3 framing that the 'larger scientific endeavor ... requires control on (at least) these other source of errors' is a stipulated standard, not a fitted or predicted quantity; whether that standard is appropriate is a correctness or robustness concern, not a circularity. The only self-citation is Munger (2023), which is used to characterize the 'agnostic impulse' and temporal validity, but the paper independently argues the temporal-validity point (e.g., the impossibility of pre-committing to experiments in 2200–2210), so the citation is not load-bearing in a circular way. The Section 6 claim that RCTs 'create novel states of the world' and that such states are 'known to align with reality' is a philosophical characterization, not an empirical prediction obtained from an input; it does not rename or repackage a fitted result. No step in the paper equates an output to an input by construction, and no uniqueness theorem or prior result by the authors is invoked to force the conclusion.
Assumptions & free parameters
assumptions (3)
- domain assumption The scientific goal of theory-testing requires generalizable knowledge across units, sites, treatment realizations, and times (the S,P,R,T dimensions).
- domain assumption Design-based inference over the timing of a study (T) is impossible because one cannot assign a probability to conducting an experiment at a future time.
- domain assumption The social world's complexity prevents us from "speaking its language" without controlling it, so observational data cannot be agnostic about ontology.
Cite this review
Pith. "Pith review of Enough?." pith.science (2026). https://pith.science/paper/CBUDXDZE
@misc{pith2026250112161,
author = {Pith},
title = {Pith review of: Enough?},
year = {2026},
howpublished = {\url{https://pith.science/paper/CBUDXDZE}},
note = {Machine review of arXiv:2501.12161}
}
read the original abstract
We respond to Aronow et al. (2025)'s paper arguing that randomized controlled trials (RCTs) are "enough," while nonparametric identification in observational studies is not. We agree with their position with respect to experimental versus observational research, but question what it would mean to extend this logic to the scientific enterprise more broadly. We first investigate what is meant by "enough," arguing that this is fundamentally a sociological claim about the relationship between statistical work and larger social and institutional processes, rather than something that can be decided from within the logic of statistics. For a more complete conception of "enough," we outline all that would need to be known -- not just knowledge of propensity scores, but knowledge of many other spatial and temporal characteristics of the social world. Even granting the logic of the critique in Aronow et al. (2025), its practical importance is a question of the contexts under study. We argue that we should not be satisfied by appeals to intuition about the complexity of "naturally occurring" propensity score functions. Instead, we call for more empirical metascience to begin to characterize this complexity. We apply this logic to the example of recommender systems developed by Aronow et al. (2025) as a demonstration of the weakness of allowing statisticians' intuitions to serve in place of metascientific data. Rather than implicitly deciding what is "enough" based on statistical applications the social world has determined to be most profitable, we argue that practicing statisticians should explicitly engage with questions like "for what?" and "for whom?" in order to adequately answer the question of "enough?"
Reference graph
Works this paper leans on
-
[1]
Site selection bias in program evaluation
Hunt Allcott. Site selection bias in program evaluation. The Quarterly Journal of Economics, 130 0 (3): 0 1117--1165, 2015
work page 2015
-
[2]
Nonparametric identification is not enough, but randomized controlled trials are
PM Aronow, James M Robins, Theo Saarinen, Fredrik S \"a vje, and Jasjeet Sekhon. Nonparametric identification is not enough, but randomized controlled trials are. Observational Studies, 2025
work page 2025
-
[3]
John Langshaw Austin. How to do things with words. Harvard university press, 1975
work page 1975
- [4]
-
[5]
The social scientist as methodological servant of the experimenting society
Donald T Campbell. The social scientist as methodological servant of the experimenting society. Policy Studies Journal, 2 0 (1): 0 72, 1973
work page 1973
-
[6]
Study designs for extending causal inferences from a randomized trial to a target population
Issa J Dahabreh, Sebastien JP A Haneuse, James M Robins, Sarah E Robertson, Ashley L Buchanan, Elizabeth A Stuart, and Miguel A Hern \'a n. Study designs for extending causal inferences from a randomized trial to a target population. American journal of epidemiology, 190 0 (8): 0 1632--1642, 2021
work page 2021
-
[7]
A constitution of democratic experimentalism
Michael C Dorf and Charles F Sabel. A constitution of democratic experimentalism. Colum. L. Rev., 98: 0 267, 1998
work page 1998
- [8]
Show all 18 references
-
[9]
fishing expedition
Andrew Gelman and Eric Loken. The garden of forking paths: Why multiple comparisons can be a problem, even when there is no “fishing expedition” or “p-hacking” and the research hypothesis was posited ahead of time. Department of Statistics, Columbia University, 348 0 (1-17): 0 3, 2013
2013
-
[10]
From sate to patt: combining experimental with observational studies to estimate population treatment effects
Erin Hartman, Richard Grieve, Roland Ramsahai, and Jasjeet S Sekhon. From sate to patt: combining experimental with observational studies to estimate population treatment effects. JR Stat. Soc. Ser. A Stat. Soc.(forthcoming). doi, 10: 0 1111, 2015
2015
-
[11]
Metaphysics and measurement
Alexandre Koyr \'e . Metaphysics and measurement. 1968
1968
-
[12]
Temporal validity as meta-science
Kevin Munger. Temporal validity as meta-science. Research & Politics, 10 0 (3): 0 20531680231187271, 2023
2023
-
[13]
Experimentalist governance
Charles F Sabel and Jonathan Zeitlin. Experimentalist governance. 2012
2012
-
[14]
The sciences of the artificial
Herbert Alexander Simon. The sciences of the artificial. 1969
1969
-
[15]
Stuart, Stephen R
Elizabeth A. Stuart, Stephen R. Cole, Catherine P. Bradshaw, and Philip J. Leaf. The use of propensity scores to assess the generalizability of results from randomized trials. Journal of the Royal Statistical Society: Series A (Statistics in Society), 174 0 (2): 0 369--386, 20...
2011 arXiv
-
[16]
How much can we generalize from impact evaluations? Journal of the European Economic Association, 18 0 (6): 0 3045--3089, 2020
Eva Vivalt. How much can we generalize from impact evaluations? Journal of the European Economic Association, 18 0 (6): 0 3045--3089, 2020
2020
-
[17]
All of nonparametric statistics
Larry Wasserman. All of nonparametric statistics. Springer Science & Business Media, 2006
2006
-
[18]
The generalizability crisis
Tal Yarkoni. The generalizability crisis. Behavioral and Brain Sciences, 45, 2022
2022
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.