{"id":"f7b6223f-cf8f-4ef2-bcdc-19eec6c10cfe","arxiv_id":"2501.12161","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"RCTs are statistically special because they set the treatment probability, but they are not scientifically sufficient because generalization across people, places, treatments, and times requires knowledge we rarely have.","lead":"This paper argues that deciding whether randomized controlled trials are \"enough\" is not just a statistical calculation but a social question about what science should know and for whom. It calls for empirical metascience to measure how complex real-world treatment assignment really is, rather than trusting statisticians' intuition.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's 'No' to 'RCTs are enough?' rests on a stipulated standard—agnostic inference over S, P, R, T—that is asserted rather than defended; if a narrower, substantive standard of sufficiency is adopted, the conclusion does not follow.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the paper's 'No' to the title question depends on requiring agnostic inference over all four dimensions, and that requirement is stipulated rather than established. My reading agrees, and I add that the Section 6 ontological claim is a second pillar that is also unsupported: the experimenter creates the assignment mechanism but not the complete realized world-state, so the claim that an RCT-created state is 'known to align with reality' overreaches. Because the reader already flagged both the stipulated standard and the unsupported ontological claim, the conditional verdict stands. No new objection requires moving the verdict; the paper remains a useful conditional commentary but should be explicit that its negative conclusion is relative to a particular epistemic standard. The concrete test proposed would settle whether the sufficiency standard can be weakened without collapsing the argument, and the same test would clarify whether the paper's central claim is about all scientific knowledge or only about a specific agnostic ideal.","tokens_in":7918,"tokens_out":4590,"duration_ms":50964,"concrete_test":"Formalize the sufficiency standard in Section 3 as a premise P1: a method is 'enough' only if it yields agnostic inference over all four dimensions (S, P, R, T). Then check whether P1 is entailed by any of the paper's own examples (site sampling, temporal validity, Yarkoni). If P1 can be weakened to 'enough for a well-specified decision problem under stated assumptions' without contradicting those examples, the central negative conclusion fails as a general statement. Concretely, construct a decision problem with a known target population and a transportability assumption (for instance, constant effect over sites), and show that an RCT plus this assumption yields a valid decision; if so, 'enough' does not require agnosticism over all dimensions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central negative claim—that RCTs are not 'enough'—is conditional on an unusually strong definition of scientific sufficiency introduced in Section 3: knowledge is 'enough' only if it provides agnostic, design-based control over (at least) the unit sample S, site P, realization of theory R, and time T. The paper never argues that any scientific claim must satisfy this standard; it simply states that the 'larger scientific endeavor ... requires control on (at least) these other source of errors.' This is not derived from ARSSS's statistical argument, and it makes the negative result nearly trivial: because the problem of induction prevents agnostic control over future T (as the paper admits for the period 2200–2210), no empirical method, including RCTs, can be 'enough.' The conclusion is then a stipulation about epistemic values rather than an argument against the sufficiency of RCTs. If a reader adopts a narrower standard—for example, reliable causal effect estimates for a specified target population under explicit transportability assumptions—an RCT plus those assumptions can be 'enough,' and Section 3 fails to engage that possibility. The paper's positive account in Section 6 ('created by the experimenter, it is known to align with reality') is also unsupported: randomization fixes the assignment mechanism, not the full state of the world; noncompliance, implementation failure, and unknown downstream effects mean the realized state is not fully created or known by the experimenter.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a critical commentary on Aronow et al. (2025) (ARSSS), who argue that RCTs are 'enough' while nonparametric identification in observational studies is not. Dimmery and Munger agree with ARSSS on the statistical point about experimental versus observational design, but argue that 'enough' cannot be settled from within statistics because it is a sociological claim about the goals of science. They propose that scientific sufficiency would require agnostic, design-based control over at least four dimensions—units (S), sites (P), theory realizations (R), and time (T)—and argue that no empirical method can satisfy this for T, so nothing can be 'enough.' They call for empirical metascience to replace statisticians' intuitions about the complexity of naturally occurring propensity score functions, illustrate the limits of 'known' propensity scores with recommender systems, and conclude that RCTs are special for ontological rather than epistemological reasons: they create novel states of the world. The final answer to the title question is 'No' for the broad scientific sense of 'enough,' but 'Absolutely' for a special place for RCTs in the ongoing process of societal learning.","tokens_in":8206,"tokens_out":6211,"duration_ms":65019,"significance":"The paper's value is as a commentary that broadens the causal-inference debate: it distinguishes statistical sufficiency from scientific sufficiency, highlights the social and institutional context in which 'enough' is decided, and draws attention to temporal validity and to the lack of metascientific evidence about propensity-score complexity. The recommender-systems discussion and the call for empirical work on the distribution of propensity-score complexity are constructive, and the paper is transparent about the absence of data for its metascientific proposal. However, the central negative conclusion rests on a stipulated epistemic standard and is accompanied by an overstrong ontological claim about RCTs; as a result, the paper's contribution is conditional on accepting a particular view of scientific sufficiency rather than on a demonstrated argument against ARSSS's position.","major_comments":[{"comment":"The negative answer to the title question is load-bearing on the claim that the 'larger scientific endeavor ... requires control on (at least) these other source of errors' (Section 3). This standard—agnostic, design-based control over S, P, R, and T—is asserted rather than argued. The paper does not engage with the possibility that scientific knowledge could be sufficient through explicit model-based transportability assumptions or by changing the target estimand, in which case an RCT plus stated assumptions can be 'enough' for a specified question. Because the conclusion 'No' follows only under the stipulated standard, the authors should either defend that standard or explicitly frame the conclusion as conditional on it.","section":"Section 3, estimand E[τ(S,P,R,T)|...]"},{"comment":"The claim that 'shoe-leather cannot be a solution' for temporal validity because an experiment's selection probability is 0% until conceived is too strong. Design-based inference over T is logically possible whenever a sampling frame of times is specified and selection probabilities are known (for example, a predetermined random start date within a defined set of possible times); the difficulty is practical and institutional rather than logical. If the authors intend the stronger skeptical point about induction, that point applies equally to all empirical methods and makes 'nothing can be enough' trivially true; the paper should clarify why this does not undermine the comparative relevance of the RCT-specific critique.","section":"Section 3, temporal validity"},{"comment":"The positive account that RCTs 'create novel states of the world' and that 'since that state of the world is created by the experimenter, it is known to align with reality' overstates the experimenter's control. Randomization fixes the assignment mechanism, but noncompliance, implementation failure, missing outcomes, and downstream effects mean the realized state is not fully created or fully known. The claim should be qualified to 'successfully implemented' experimental assignments, or the ontological claim should be formulated more carefully; otherwise the paper risks replacing one exaggerated pedestal with another.","section":"Section 6, final paragraph"}],"minor_comments":[{"comment":"There is a typo in 'randomize d controlled trials' in the abstract; the same line appears broken in the full-text rendering.","section":"Abstract and full text, line 1"},{"comment":"'Lipshitz' should be 'Lipschitz'.","section":"Section 2"},{"comment":"The notation E[τ(S,P,R,T)|S=s, P=p, R=r, T=t] is informal; please define the probability space and clarify how the conditioning relates to the usual definition of an estimand.","section":"Section 3"},{"comment":"The example 'XUserThreshold03: 20%, XUserThreshold10: 80%' is not explained; a brief description of what these thresholds represent would help readers who are not familiar with recommender-system log data.","section":"Section 5"},{"comment":"The abstract says the authors agree with ARSSS 'with respect to experimental versus observational research,' while the conclusion answers 'No' to whether RCTs are enough; please reconcile these statements by specifying the exact scope of the agreement.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"This is a commentary paper with no new data or formal proofs; its contribution is conceptual and metascientific. It may be a good fit for a journal that publishes discussion or response pieces, but the editor should consider whether the paper's scope matches the journal's methodological profile. The central claim is conditional on an unargued standard of scientific sufficiency, which the authors can address by reframing the conclusion as conditional or by adding a defense of the standard; the positive ontological claim also needs qualification. With those revisions, the paper could be a useful contribution to the debate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is a fair, well-written commentary that lands one solid punch—ARSSS's 'enough' is a practical and sociological claim, not a statistical one—but its decisive 'No' rests on a stipulated standard of sufficiency that the paper never defends.\n\nWhat is actually new: the authors extend Yarkoni and Munger's earlier generalizability arguments into a multi-dimensional (S, P, R, T) design-based frame, and they make the temporal validity point sharply: shoe leather cannot solve generalization over future time. The call for empirical metascience over statisticians' intuitions about propensity score complexity is well aimed and is the most useful contribution. The recommender systems section is also good—it shows that 'known' propensity scores are not enough unless the ontology and the access are there, and it raises the power question without melodrama.\n\nWhat the paper does well: it grants ARSSS's statistical argument fully and then disputes the scope. That is the right way to engage. The writing is clear, the citations are appropriate, and the examples are well chosen.\n\nSoft spots. Section 3's standard—that science requires agnostic control over S, P, R, T—is asserted, not argued. The authors never explain why a substantive standard, such as reliable causal estimates for a specified target population under explicit transportability assumptions, would not count as 'enough.' On their own standard they admit nothing can be enough (future T is impossible by induction), which makes the 'No' close to trivial. Section 6's ontological claim—'the true value of RCTs is ontological: they create novel states of the world'—is stated more strongly than supported. Randomization fixes the assignment mechanism, not the full realized state; noncompliance, implementation failure, and downstream effects remain. So that sentence overclaims.\n\nNone of this sinks the paper. The core observation—that 'enough' is a question about goals and context, not just estimation—holds, and the call for metascience is sound. It is a commentary, not a new method or result, and should be judged as such.\n\nWho this is for: methodologists and metascience researchers following the RCT-vs-observational debate. It deserves a serious referee if the venue treats commentary as scholarship. My own recommendation: engage with it, but ask the authors to temper the 'No' to 'not enough for generalizable knowledge under the standard we propose,' and to hedge or support the ontological claim. I wouldn't cite it in my own next paper unless I write on this debate, but I'd be glad to see it in print.","headline":"A fair, clear-headed commentary that lands the 'enough for what?' question but overstates its case by stipulating an extreme standard of sufficiency.","tokens_in":8697,"tokens_out":2787,"would_cite":false,"duration_ms":27167,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that RCTs are statistically privileged for one estimand but are not 'enough' for generalizable science; their true value is ontological, creating novel states of the world.","keywords":["randomized controlled trials","causal inference","propensity scores","generalizability","empirical metascience","temporal validity","shoe leather","design-based inference"],"falsifier":"Demonstrate that treatment effects estimated in a single-site, single-time RCT reliably predict effects at other sites and times without any selection models or ignorability assumptions; repeated success in such out-of-sample prediction would falsify the claim that the scientific endeavor requires controlling $S$, $P$, $R$, $T$.","tokens_in":7735,"feed_emoji":"🎲","tokens_out":9146,"duration_ms":86459,"temperature":0.7,"pith_summary":"Randomized controlled trials are statistically privileged for estimating a treatment effect in one sample, at one site, with one version of treatment, at one time. The authors of this reply accept that privilege but argue that it does not make RCTs 'enough' for the broader scientific goal of generalizable knowledge. The meaning of 'enough' cannot be settled from inside statistics, because it depends on what science is for and who will use the results. The paper's affirmative claim is that RCTs are special because they create novel states of the world—the experimenter performs a simplified model of the world rather than assuming one—and that this ontological power, not their estimation properties, is what merits them a special place.","feed_headline":"RCTs are not enough for science","feed_subtitle":"Randomization fixes the estimand, not the sample, site, treatment, or time, so 'enough' is a sociological question.","key_machinery":"The central machinery is a four-part decomposition of the causal estimand into sample $S$, site $P$, realization of treatment $R$, and time $T$, each of which has a selection process that functions like a propensity score. The paper combines this decomposition with the notion of 'shoe leather'—design-based knowledge of selection probabilities—to show what an RCT does and does not solve: it gives agnostic inference on the realized cell $(s,p,r,t)$ but leaves the other cells open, and for $T$ shoe leather is impossible because future designs cannot receive selection probabilities before they exist. A second piece is the performative reading of experiments, in which running an RCT imposes a simplified ontology on the world and thereby makes the propensity score known by construction.","core_discovery":"The paper's central claim is that the argument it responds to—that nonparametric identification is not enough while randomized controlled trials are—is correct as far as estimation goes, yet incomplete as a statement about science. RCTs provide agnostic inference on the estimand $\\mathbb{E}[\\tau(S,P,R,T)\\mid S=s,P=p,R=r,T=t]$ but say almost nothing about the other selections—which sample, which site, which realization of treatment, which time—that would be needed for generalizable knowledge. The authors therefore distinguish 'enough' as a statistical property from 'enough' as a property of a scientific enterprise, and they argue that on the latter standard no design is enough. The specialness of RCTs lies not in their epistemology but in their ontology: they perform a model of the world by creating the very state they study.","pith_inferences":["A testable extension of the paper's logic would be to run multi-site, multi-time randomized trials with random site selection, measuring how much treatment effects vary across the $S$, $P$, $R$, $T$ dimensions.","Implicit in the paper but not stated: if RCTs create novel states of the world, then choosing which possible worlds to create is a value-laden decision, adding an ethical dimension to experimental design that statistics alone cannot resolve.","One operationalization of the call for empirical metascience would be a public registry of propensity score functions estimated in observational studies, with measures of complexity, to replace conflicting intuitions with data.","The account also redirects gold-standard debates toward institutional questions: who gets to run experiments, who owns the resulting propensity scores, and whose questions get randomized."],"forward_implications":["A successful RCT in one site and time cannot, by itself, support a generalizable scientific conclusion; extending its result requires either design-based control over site and time selection or explicit assumptions that those selections are ignorable.","Calls to make RCTs the gold standard of causal inference are better read as sociological claims about the status of methods, not as consequences of statistical theorems.","The practical importance of the earlier paper's critique depends on empirical facts about the complexity of naturally occurring propensity score functions—facts that do not yet exist and should be collected by metascience.","Recommender systems' known propensity scores are a case of 'artificial' rather than 'natural' experiments: the scores are knowable only because the virtual world was engineered to be simple, and access to them is limited by corporate power.","The slogan 'no causation without control' replaces 'no causation without manipulation' and points statisticians toward the political and social dimensions of experimental designs."],"supporting_citations":[{"why":"Supplies the target claim that nonparametric identification is not enough while RCTs are, and the recommender-system example examined here.","marker":"Aronow et al. (2025)"},{"why":"Introduces 'shoe leather' as design-based knowledge of selection mechanisms, the standard by which the paper measures what an RCT provides and what generalization would need.","marker":"Freedman (1991)"},{"why":"Contributes the concept of temporal validity and the argument that choosing when to run a study is itself a selection process.","marker":"Munger (2023)"},{"why":"Provides the generalizability-crisis argument that every design element should be treated as a random draw, which the paper extends to less-controlled real-world RCTs.","marker":"Yarkoni (2022)"},{"why":"Shows how propensity-score-like weights are used to generalize from randomized trials, illustrating the paper's S, P, R, T generalization machinery.","marker":"Stuart et al. (2011)"},{"why":"Documents site selection bias in program evaluation, motivating the need to randomize or model site selection.","marker":"Allcott (2015)"},{"why":"Provides evidence on how much impact evaluations generalize across contexts, supporting the paper's concern about the site dimension.","marker":"Vivalt (2020)"},{"why":"Supplies the distinction between natural and artificial worlds used to argue that recommender-system propensity scores are knowable only because the environment was engineered.","marker":"Simon (1969)"}],"fun_headline_variants":["RCTs fix the estimand, not the science","Enough is sociological, not just statistical","Even randomization can't make 'enough' general","What is enough? Ask for whom, not just how"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on the premise that scientific sufficiency requires agnostic control over generalization across sample, site, treatment realization, and time; if 'enough' is allowed to mean a well-estimated causal effect in one setting, the conclusion that RCTs are not enough does not follow.","fun_headline_variants_meta":{"raw":{"variants":["RCTs fix the estimand, not the science","Enough is sociological, not just statistical","Even randomization can't make 'enough' general","What is enough? Ask for whom, not just how"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1334,"prompt_tokens":999,"completion_tokens":335,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":285}},"tokens_in":615,"tokens_out":335,"duration_ms":3802,"temperature":1.0,"reasoning_tokens":285,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:26:25.419141+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Demonstrate that treatment effects estimated in a single-site, single-time RCT reliably predict effects at other sites and times without any selection models or ignorability assumptions; repeated success in such out-of-sample prediction would falsify the claim that the scientific endeavor requires controlling $S$, $P$, $R$, $T$.","supporting_citations":[{"cited_title":"Nonparametric identification is not enough, but randomized controlled trials are","cited_arxiv_id":null,"evidence_quote":"Supplies the target claim that nonparametric identification is not enough while RCTs are, and the recommender-system example examined here."},{"cited_title":"Temporal validity as meta-science","cited_arxiv_id":null,"evidence_quote":"Contributes the concept of temporal validity and the argument that choosing when to run a study is itself a selection process."},{"cited_title":"The generalizability crisis","cited_arxiv_id":null,"evidence_quote":"Provides the generalizability-crisis argument that every design element should be treated as a random draw, which the paper extends to less-controlled real-world RCTs."},{"cited_title":"Site selection bias in program evaluation","cited_arxiv_id":null,"evidence_quote":"Documents site selection bias in program evaluation, motivating the need to randomize or model site selection."},{"cited_title":"How much can we generalize from impact evaluations? Journal of the European Economic Association, 18 0 (6): 0 3045--3089, 2020","cited_arxiv_id":null,"evidence_quote":"Provides evidence on how much impact evaluations generalize across contexts, supporting the paper's concern about the site dimension."},{"cited_title":"The sciences of the artificial","cited_arxiv_id":null,"evidence_quote":"Supplies the distinction between natural and artificial worlds used to argue that recommender-system propensity scores are knowable only because the environment was engineered."}],"review_version":1}