Pith. sign in

REVIEW 3 major objections 5 minor 18 references

Enough?

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper argues that RCTs are statistically privileged for one estimand but are not 'enough' for generalizable science; their true value is ontological, creating novel states of the world.

desk verdict A fair, clear-headed commentary that lands the 'enough for what?' question but overstates its case by stipulating an extreme standard of sufficiency. read the letter →

arxiv 2501.12161 v1 pith:CBUDXDZE submitted 2025-01-21 stat.ME

classification stat.ME
keywords randomizedcontrolledtrialscausalinferencepropensityscoresgeneralizabilityempiricalmetasciencetemporalvalidityshoeleatherdesign-based
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Randomized controlled trials are statistically privileged for estimating a treatment effect in one sample, at one site, with one version of treatment, at one time. The authors of this reply accept that privilege but argue that it does not make RCTs 'enough' for the broader scientific goal of generalizable knowledge. The meaning of 'enough' cannot be settled from inside statistics, because it depends on what science is for and who will use the results. The paper's affirmative claim is that RCTs are special because they create novel states of the world—the experimenter performs a simplified model of the world rather than assuming one—and that this ontological power, not their estimation properties, is what merits them a special place.

What carries the argument

The central machinery is a four-part decomposition of the causal estimand into sample $S$, site $P$, realization of treatment $R$, and time $T$, each of which has a selection process that functions like a propensity score. The paper combines this decomposition with the notion of 'shoe leather'—design-based knowledge of selection probabilities—to show what an RCT does and does not solve: it gives agnostic inference on the realized cell $(s,p,r,t)$ but leaves the other cells open, and for $T$ shoe leather is impossible because future designs cannot receive selection probabilities before they exist. A second piece is the performative reading of experiments, in which running an RCT imposes a simplified ontology on the world and thereby makes the propensity score known by construction.

What would settle it

Demonstrate that treatment effects estimated in a single-site, single-time RCT reliably predict effects at other sites and times without any selection models or ignorability assumptions; repeated success in such out-of-sample prediction would falsify the claim that the scientific endeavor requires controlling $S$, $P$, $R$, $T$.

Watch

Extended reading notes

Core claim

The paper's central claim is that the argument it responds to—that nonparametric identification is not enough while randomized controlled trials are—is correct as far as estimation goes, yet incomplete as a statement about science. RCTs provide agnostic inference on the estimand $\mathbb{E}[\tau(S,P,R,T)\mid S=s,P=p,R=r,T=t]$ but say almost nothing about the other selections—which sample, which site, which realization of treatment, which time—that would be needed for generalizable knowledge. The authors therefore distinguish 'enough' as a statistical property from 'enough' as a property of a scientific enterprise, and they argue that on the latter standard no design is enough. The specialness of RCTs lies not in their epistemology but in their ontology: they perform a model of the world by creating the very state they study.

Load-bearing premise

The argument depends on the premise that scientific sufficiency requires agnostic control over generalization across sample, site, treatment realization, and time; if 'enough' is allowed to mean a well-estimated causal effect in one setting, the conclusion that RCTs are not enough does not follow.

Editorial extensions

If this is right

  • A successful RCT in one site and time cannot, by itself, support a generalizable scientific conclusion; extending its result requires either design-based control over site and time selection or explicit assumptions that those selections are ignorable.
  • Calls to make RCTs the gold standard of causal inference are better read as sociological claims about the status of methods, not as consequences of statistical theorems.
  • The practical importance of the earlier paper's critique depends on empirical facts about the complexity of naturally occurring propensity score functions—facts that do not yet exist and should be collected by metascience.
  • Recommender systems' known propensity scores are a case of 'artificial' rather than 'natural' experiments: the scores are knowable only because the virtual world was engineered to be simple, and access to them is limited by corporate power.
  • The slogan 'no causation without control' replaces 'no causation without manipulation' and points statisticians toward the political and social dimensions of experimental designs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension of the paper's logic would be to run multi-site, multi-time randomized trials with random site selection, measuring how much treatment effects vary across the $S$, $P$, $R$, $T$ dimensions.
  • Implicit in the paper but not stated: if RCTs create novel states of the world, then choosing which possible worlds to create is a value-laden decision, adding an ethical dimension to experimental design that statistics alone cannot resolve.
  • One operationalization of the call for empirical metascience would be a public registry of propensity score functions estimated in observational studies, with measures of complexity, to replace conflicting intuitions with data.
  • The account also redirects gold-standard debates toward institutional questions: who gets to run experiments, who owns the resulting propensity scores, and whose questions get randomized.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper is a critical commentary on Aronow et al. (2025) (ARSSS), who argue that RCTs are 'enough' while nonparametric identification in observational studies is not. Dimmery and Munger agree with ARSSS on the statistical point about experimental versus observational design, but argue that 'enough' cannot be settled from within statistics because it is a sociological claim about the goals of science. They propose that scientific sufficiency would require agnostic, design-based control over at least four dimensions—units (S), sites (P), theory realizations (R), and time (T)—and argue that no empirical method can satisfy this for T, so nothing can be 'enough.' They call for empirical metascience to replace statisticians' intuitions about the complexity of naturally occurring propensity score functions, illustrate the limits of 'known' propensity scores with recommender systems, and conclude that RCTs are special for ontological rather than epistemological reasons: they create novel states of the world. The final answer to the title question is 'No' for the broad scientific sense of 'enough,' but 'Absolutely' for a special place for RCTs in the ongoing process of societal learning.

Significance. The paper's value is as a commentary that broadens the causal-inference debate: it distinguishes statistical sufficiency from scientific sufficiency, highlights the social and institutional context in which 'enough' is decided, and draws attention to temporal validity and to the lack of metascientific evidence about propensity-score complexity. The recommender-systems discussion and the call for empirical work on the distribution of propensity-score complexity are constructive, and the paper is transparent about the absence of data for its metascientific proposal. However, the central negative conclusion rests on a stipulated epistemic standard and is accompanied by an overstrong ontological claim about RCTs; as a result, the paper's contribution is conditional on accepting a particular view of scientific sufficiency rather than on a demonstrated argument against ARSSS's position.

major comments (3)
  1. [Section 3, estimand E[τ(S,P,R,T)|...]] The negative answer to the title question is load-bearing on the claim that the 'larger scientific endeavor ... requires control on (at least) these other source of errors' (Section 3). This standard—agnostic, design-based control over S, P, R, and T—is asserted rather than argued. The paper does not engage with the possibility that scientific knowledge could be sufficient through explicit model-based transportability assumptions or by changing the target estimand, in which case an RCT plus stated assumptions can be 'enough' for a specified question. Because the conclusion 'No' follows only under the stipulated standard, the authors should either defend that standard or explicitly frame the conclusion as conditional on it.
  2. [Section 3, temporal validity] The claim that 'shoe-leather cannot be a solution' for temporal validity because an experiment's selection probability is 0% until conceived is too strong. Design-based inference over T is logically possible whenever a sampling frame of times is specified and selection probabilities are known (for example, a predetermined random start date within a defined set of possible times); the difficulty is practical and institutional rather than logical. If the authors intend the stronger skeptical point about induction, that point applies equally to all empirical methods and makes 'nothing can be enough' trivially true; the paper should clarify why this does not undermine the comparative relevance of the RCT-specific critique.
  3. [Section 6, final paragraph] The positive account that RCTs 'create novel states of the world' and that 'since that state of the world is created by the experimenter, it is known to align with reality' overstates the experimenter's control. Randomization fixes the assignment mechanism, but noncompliance, implementation failure, missing outcomes, and downstream effects mean the realized state is not fully created or fully known. The claim should be qualified to 'successfully implemented' experimental assignments, or the ontological claim should be formulated more carefully; otherwise the paper risks replacing one exaggerated pedestal with another.
minor comments (5)
  1. [Abstract and full text, line 1] There is a typo in 'randomize d controlled trials' in the abstract; the same line appears broken in the full-text rendering.
  2. [Section 2] 'Lipshitz' should be 'Lipschitz'.
  3. [Section 3] The notation E[τ(S,P,R,T)|S=s, P=p, R=r, T=t] is informal; please define the probability space and clarify how the conditioning relates to the usual definition of an estimand.
  4. [Section 5] The example 'XUserThreshold03: 20%, XUserThreshold10: 80%' is not explained; a brief description of what these thresholds represent would help readers who are not familiar with recommender-system log data.
  5. [Section 6] The abstract says the authors agree with ARSSS 'with respect to experimental versus observational research,' while the conclusion answers 'No' to whether RCTs are enough; please reconcile these statements by specifying the exact scope of the agreement.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an explicitly normative and sociological commentary, not a derivation, and its self-citation is not load-bearing.

full rationale

This paper makes no statistical predictions, fits no parameters, and presents no derivation chain whose output could reduce to its inputs. Its central claim—that RCTs are not 'enough' for the broader scientific enterprise—is avowedly a sociological and philosophical argument about the meaning of 'enough,' not a result derived from the statistical argument of Aronow et al. (2025). The Section 3 framing that the 'larger scientific endeavor ... requires control on (at least) these other source of errors' is a stipulated standard, not a fitted or predicted quantity; whether that standard is appropriate is a correctness or robustness concern, not a circularity. The only self-citation is Munger (2023), which is used to characterize the 'agnostic impulse' and temporal validity, but the paper independently argues the temporal-validity point (e.g., the impossibility of pre-committing to experiments in 2200–2210), so the citation is not load-bearing in a circular way. The Section 6 claim that RCTs 'create novel states of the world' and that such states are 'known to align with reality' is a philosophical characterization, not an empirical prediction obtained from an input; it does not rename or repackage a fitted result. No step in the paper equates an output to an input by construction, and no uniqueness theorem or prior result by the authors is invoked to force the conclusion.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper relies on normative assumptions about what scientific knowledge requires and on the impossibility of temporal randomization. It introduces no free parameters or invented entities; its empirical claims are explicitly unsupported calls for further research.

assumptions (3)
  • domain assumption The scientific goal of theory-testing requires generalizable knowledge across units, sites, treatment realizations, and times (the S,P,R,T dimensions).
    Invoked in Section 3 when the authors assert that the RCT estimand has "little bearing on the larger scientific endeavor" unless these dimensions are controlled; this is a normative view of what science requires.
  • domain assumption Design-based inference over the timing of a study (T) is impossible because one cannot assign a probability to conducting an experiment at a future time.
    Section 3 argues that "shoe-leather cannot be a solution" for temporal validity, which assumes the impossibility or implausibility of temporal randomization.
  • domain assumption The social world's complexity prevents us from "speaking its language" without controlling it, so observational data cannot be agnostic about ontology.
    Section 2 argues that observational work assumes an ontology while RCTs perform one; this is a philosophical premise about language and measurement.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Enough?." pith.science (2026). https://pith.science/paper/CBUDXDZE

@misc{pith2026250112161,
  author       = {Pith},
  title        = {Pith review of: Enough?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CBUDXDZE}},
  note         = {Machine review of arXiv:2501.12161}
}
read the original abstract

We respond to Aronow et al. (2025)'s paper arguing that randomized controlled trials (RCTs) are "enough," while nonparametric identification in observational studies is not. We agree with their position with respect to experimental versus observational research, but question what it would mean to extend this logic to the scientific enterprise more broadly. We first investigate what is meant by "enough," arguing that this is fundamentally a sociological claim about the relationship between statistical work and larger social and institutional processes, rather than something that can be decided from within the logic of statistics. For a more complete conception of "enough," we outline all that would need to be known -- not just knowledge of propensity scores, but knowledge of many other spatial and temporal characteristics of the social world. Even granting the logic of the critique in Aronow et al. (2025), its practical importance is a question of the contexts under study. We argue that we should not be satisfied by appeals to intuition about the complexity of "naturally occurring" propensity score functions. Instead, we call for more empirical metascience to begin to characterize this complexity. We apply this logic to the example of recommender systems developed by Aronow et al. (2025) as a demonstration of the weakness of allowing statisticians' intuitions to serve in place of metascientific data. Rather than implicitly deciding what is "enough" based on statistical applications the social world has determined to be most profitable, we argue that practicing statisticians should explicitly engage with questions like "for what?" and "for whom?" in order to adequately answer the question of "enough?"

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

18 extracted references · 16 canonical work pages

  1. [1]

    Site selection bias in program evaluation

    Hunt Allcott. Site selection bias in program evaluation. The Quarterly Journal of Economics, 130 0 (3): 0 1117--1165, 2015

  2. [2]

    Nonparametric identification is not enough, but randomized controlled trials are

    PM Aronow, James M Robins, Theo Saarinen, Fredrik S \"a vje, and Jasjeet Sekhon. Nonparametric identification is not enough, but randomized controlled trials are. Observational Studies, 2025

  3. [3]

    How to do things with words

    John Langshaw Austin. How to do things with words. Harvard university press, 1975

  4. [4]

    Designing freedom

    Stafford Beer. Designing freedom. House of Anansi, 1993

  5. [5]

    The social scientist as methodological servant of the experimenting society

    Donald T Campbell. The social scientist as methodological servant of the experimenting society. Policy Studies Journal, 2 0 (1): 0 72, 1973

  6. [6]

    Study designs for extending causal inferences from a randomized trial to a target population

    Issa J Dahabreh, Sebastien JP A Haneuse, James M Robins, Sarah E Robertson, Ashley L Buchanan, Elizabeth A Stuart, and Miguel A Hern \'a n. Study designs for extending causal inferences from a randomized trial to a target population. American journal of epidemiology, 190 0 (8): 0 1632--1642, 2021

  7. [7]

    A constitution of democratic experimentalism

    Michael C Dorf and Charles F Sabel. A constitution of democratic experimentalism. Colum. L. Rev., 98: 0 267, 1998

  8. [8]

    Freedman

    David A. Freedman. Statistical models and shoe leather. Sociological Methodology, 21: 0 291--313, 1991. ISSN 00811750, 14679531. URL http://www.jstor.org/stable/270939

Show all 18 references
  1. [9]

    fishing expedition

    Andrew Gelman and Eric Loken. The garden of forking paths: Why multiple comparisons can be a problem, even when there is no “fishing expedition” or “p-hacking” and the research hypothesis was posited ahead of time. Department of Statistics, Columbia University, 348 0 (1-17): 0 3, 2013

  2. [10]

    From sate to patt: combining experimental with observational studies to estimate population treatment effects

    Erin Hartman, Richard Grieve, Roland Ramsahai, and Jasjeet S Sekhon. From sate to patt: combining experimental with observational studies to estimate population treatment effects. JR Stat. Soc. Ser. A Stat. Soc.(forthcoming). doi, 10: 0 1111, 2015

  3. [11]

    Metaphysics and measurement

    Alexandre Koyr \'e . Metaphysics and measurement. 1968

  4. [12]

    Temporal validity as meta-science

    Kevin Munger. Temporal validity as meta-science. Research & Politics, 10 0 (3): 0 20531680231187271, 2023

  5. [13]

    Experimentalist governance

    Charles F Sabel and Jonathan Zeitlin. Experimentalist governance. 2012

  6. [14]

    The sciences of the artificial

    Herbert Alexander Simon. The sciences of the artificial. 1969

  7. [15]

    Stuart, Stephen R

    Elizabeth A. Stuart, Stephen R. Cole, Catherine P. Bradshaw, and Philip J. Leaf. The use of propensity scores to assess the generalizability of results from randomized trials. Journal of the Royal Statistical Society: Series A (Statistics in Society), 174 0 (2): 0 369--386, 20...

  8. [16]

    How much can we generalize from impact evaluations? Journal of the European Economic Association, 18 0 (6): 0 3045--3089, 2020

    Eva Vivalt. How much can we generalize from impact evaluations? Journal of the European Economic Association, 18 0 (6): 0 3045--3089, 2020

  9. [17]

    All of nonparametric statistics

    Larry Wasserman. All of nonparametric statistics. Springer Science & Business Media, 2006

  10. [18]

    The generalizability crisis

    Tal Yarkoni. The generalizability crisis. Behavioral and Brain Sciences, 45, 2022

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.