REVIEW 3 major objections 4 minor 1 cited by
A Bureaucratic Theory of Statistics
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper argues that randomized controlled trials and statistical tests are best understood as 'ex ante policy': rules fixed before data collection that govern future actions, making statistics' primary purpose regulation rather than…
desk verdict A clearly written opinion piece that usefully names the regulatory role of statistical rules, but the stronger claim that regulation is the purpose of statistics is not established by the argument. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing construct is 'ex ante policy,' defined as a rule specified by a regulatory body before data collection that dictates which future actions are acceptable. The paper's worked example rewrites the two-sample proportions test as a correlation threshold, approving a treatment when $R(X,Y) \ge t/\sqrt{n}$, with $t=1.96$ yielding the conventional z-test at level $\alpha=0.05$ and an equivalent chi-squared test. This identity shows that a p-value threshold is an arbitrary evidentiary bar, chosen by convention, not a theorem. Confidence intervals are then defined as randomized algorithms with an ex ante success probability, and Bayesian decision procedures count as ex ante policy whenever prior, likelihood, and utility are fixed in advance.
What would settle it
A broad survey of statistical use across fields such as astronomy, particle physics, genomics, and economic forecasting would falsify the claim if it found a substantial class of applications in which no pre-specified action rule exists and results are used only to update scientific understanding.
Extended reading notes
Core claim
The paper's central claim is that the purpose of a statistical test is regulation. It grounds this in the observation that randomized trials and significance tests, whatever their epistemic ambitions, serve as pre-specified rules that decide whether a drug comes to market, whether a software feature ships, and whether a paper is published. The author introduces 'ex ante policy' as the umbrella term for such rules, and argues that both frequentist procedures and Bayesian decision theory fit under it, because both require the decision-relevant ingredients to be fixed before data collection. The ex ante/ex post distinction is then used to explain persistent confusion: tests carry verifiable, prescriptive guarantees, whereas inference about whether a theory is true is complex, subjective, and not reducible to a numerical algorithm. Consequently, debates over thresholds and rituals are better understood as debates over regulatory conventions than over the foundations of knowledge.
Load-bearing premise
The argument depends on accepting that the purpose of a system is what it does and on treating regulatory applications like drug trials, feature rollouts, and journal review as representative of statistics as a whole; if measurement, forecasting, or exploratory modeling are equally central, the conclusion that statistics' purpose is regulation does not follow.
Editorial extensions
If this is right
- Debates about the correct p-value threshold become deliberations about a regulatory convention, which can be resolved by stakeholder agreement rather than by epistemology.
- Confidence intervals should be communicated as algorithms whose coverage probability is over prospective repetitions, ending the temptation to read a realized interval as an epistemic statement.
- Bayesian and frequentist methods that fix their decision rules in advance enter the same framework, so methodology choices can be compared on transparency and fairness rather than on philosophical loyalty.
- A methodology research agenda opens for designing statistical rules, including the question of when observational data are acceptable as a basis for policy.
Reading between the lines
- An implicit corollary is that the replication crisis can be read as a regulatory failure—rules that no longer encode a community's values—rather than as a failure of inference; this reading suggests fixes like renegotiating evidentiary bars rather than searching for better estimators.
- The ex ante/ex post split predicts that the perceived authority of a statistical result depends less on its epistemic justification than on the transparency and perceived fairness of the rule that produced it, which could be tested in surveys of stakeholders.
- A natural extension is to classify published statistical applications as ex ante policy or ex post inference and compare their evidentiary standards, which would provide a quantitative check on whether regulation really dominates practice.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This commentary argues that statistical methods, especially randomized controlled trials and hypothesis tests, are best understood not as instruments of epistemic inference but as 'ex ante policy': rules specified before data collection to govern future actions. The author introduces this term, illustrates it with FDA drug approval standards, A/B testing, and journal peer review, and reformulates the z-test as a correlation threshold (Eq. 1). He then contrasts ex ante policy with ex post inference, argues that the ex ante frame resolves debates about p-values and confidence intervals, and concludes that 'the purpose of statistical tests is regulation' (Section 4). The paper is written as an essay intended to reframe how statisticians and methodologists think about their field.
Significance. The paper identifies a genuine and underappreciated function of statistical methods: pre-specified decision rules provide transparency, fairness, and accountability in governance. The mathematical reformulation in Eq. (1) is standard and clearly explained, and the historical examples (Bradford Hill, FDA, Kefauver-Harris) are relevant and cited. If the reframing were accepted, it could productively reorient pedagogy, methodology, and discussions of p-values and confidence intervals. However, the central claim as stated is not established: the paper offers an interpretive lens rather than a testable theory, and its conclusion that regulation is the purpose of statistics rests on an unargued selection of examples. The paper is transparent about its own framing and contains no fitted parameters, which is appropriate for a commentary; its strength is in the clarity of the proposed distinction rather than in rigorous proof.
major comments (3)
- [Section 4 and opening] The load-bearing claim, 'The purpose of statistical tests is regulation,' is not entailed by the evidence presented. The argument uses Stafford Beer's maxim 'the purpose of a system is what it does' (opening), but a system does many things, and the paper does not provide a criterion for why the regulatory uses (FDA approvals, A/B tests, journal review) fix the purpose of statistics, rather than ex post inference uses such as estimating treatment effects in epidemiology, measuring physical constants, or forecasting. Section 2 explicitly distinguishes ex ante policy from ex post inference, and the same methods are routinely used in both modes. Without a principled basis for weighting these uses, the conclusion that regulation is the defining purpose does not follow; at most, the paper shows that regulation is an important use. The paper also appears to contradict itself by saying 'Rulemaking is certainly not the singular valuable application of statistics' yet later stating a singular purpose. Please reconcile these statements and either soften the conclusion to 'a' purpose or justify the priority claim.
- [Section 1] The definition of ex ante policy as 'the specification of rules by some regulatory body that governs acceptable future actions' is broad enough to cover almost any pre-specified statistical procedure, and the paper then concludes that 'most common Frequentist methods' are examples. This creates a circularity risk: if every pre-specified rule is defined as policy, then finding that statistics is policy is true by definition. The paper needs a limiting criterion—for example, a distinction between rules adopted by an actual governance body and the merely procedural features of a test—or an independent benchmark for what counts as policy. As written, the conceptual claim is too elastic to be informative.
- [Section 2] The paper asserts that 'science doesn't progress via Popperian means' and that 'there are no computable Bayesian updates that crunch empirical data into crisp posteriors.' These are strong empirical and historical claims used to downplay the importance of ex post inference. They are unsupported by evidence or argument, and the central thesis does not depend on the strong versions of these claims. Please either provide supporting citations or hedge these statements so that the argument does not rely on contested historiography.
minor comments (4)
- [Section 4] In the list of designs, 'differences and differences' should be 'difference-in-differences.'
- [Section 4] 'AB tests' should be written as 'A/B tests.'
- [References] The Imbens and Rubin reference contains formatting errors: '1 edition edition' and the month '4 2015' should be cleaned up.
- [Footnote 1] The Greek letter in 'ˆβ' is garbled in the text; please ensure it renders correctly.
Circularity Check
No circularity: an interpretive commentary with no fitted parameters, derived predictions, or self-citation chain.
full rationale
This paper is an opinion/commentary rather than a derivation: it contains no fitted parameters, no equations whose outputs are recycled as predictions, and no self-citations. The central move is to define 'ex ante policy' as pre-specified rules designed before data collection to govern future actions, then to argue from examples (FDA trials, A/B tests, journal review) and Stafford Beer's maxim that statistical tests are best understood as regulatory instruments. That definition is introduced transparently ('I introduce my own umbrella term for such statistical rulemaking: ex ante policy'), and the conclusion is presented as an interpretive recommendation, not as a theorem derived from the definition. The possible objection that the definition is broad enough to cover most testing procedures is a substantive critique of the framework's scope, not a circularity in the argument chain: no specific claim reduces by construction to its own input.
Assumptions & free parameters
assumptions (4)
- domain assumption The purpose of a system is what it does (Stafford Beer).
- domain assumption Selected applications (drug approval, A/B testing, journal review) are representative of what statistical tests do.
- domain assumption Ex ante policy comes with verifiable theorems, ex post inference does not.
- domain assumption There are no computable Bayesian updates that crunch empirical data into crisp posteriors.
invented entities (1)
-
ex ante policy
Cite this review
Pith. "Pith review of A Bureaucratic Theory of Statistics." pith.science (2026). https://pith.science/paper/A64NNK7U
@misc{pith2026250103457,
author = {Pith},
title = {Pith review of: A Bureaucratic Theory of Statistics},
year = {2026},
howpublished = {\url{https://pith.science/paper/A64NNK7U}},
note = {Machine review of arXiv:2501.03457}
}
read the original abstract
This commentary proposes a framework for understanding the role of statistics in policy-making, regulation, and bureaucratic systems. I introduce the concept of "ex ante policy," describing statistical rules and procedures designed before data collection to govern future actions. Through examining examples, particularly clinical trials, I explore how ex ante policy serves as a calculus of bureaucracy, providing numerical foundations for governance through clear, transparent rules. The ex ante frame obviates heated debates about inferential interpretations of probability and statistical tests, p-values, and rituals. I conclude by calling for a deeper appreciation of statistics' bureaucratic function and suggesting new directions for research in policy-oriented statistical methodology.
Forward citations
Cited by 1 Pith paper
-
Aggregated Individual Reporting for Post-Deployment Evaluation
The authors formalize a mechanism for collecting and aggregating public reports about deployed AI systems, aiming to surface unknown harms and enable accountability.
Reference graph
Works this paper leans on
-
[1]
Banerjee, Sylvain Chassang, Sergio Montero, and Erik Snowberg
Abhijit V. Banerjee, Sylvain Chassang, Sergio Montero, and Erik Snowberg. A theory of experimenters: Robustness, randomization, and balance. American Economic Review, 110 0 (4): 0 1206--30, 2020
work page 2020
-
[2]
Stafford Beer. The Heart of Enterprise. John Wiley & Sons, 1979
work page 1979
-
[3]
The ASA Statement on p-Values: Context, Process, and Purpose
Yoav Benjamini. It's not the p-values' fault. The American Statistician, 2016. URL https://doi.org/10.1080/00031305.2016.1154108. Fourth article in the online supplement to "The ASA Statement on p-Values: Context, Process, and Purpose."
arXiv 2016
-
[4]
Statistical Power Analysis for the Behavioral Sciences
Jacob Cohen. Statistical Power Analysis for the Behavioral Sciences. Academic Press, 1969
work page 1969
-
[5]
Gerd Gigerenzer. Mindless statistics. The Journal of Socio-Economics, 33 0 (5): 0 587--606, 2004. doi:https://doi.org/10.1016/j.socec.2004.09.033
-
[6]
Rink Hoekstra, Richard D. Morey, Jeffrey N. Rouder, and Eric-Jan Wagenmakers. Robust misinterpretation of confidence intervals. Psychonomic Bulletin & Review, 21 0 (5): 0 1157--1164, October 2014. ISSN 1069-9384, 1531-5320. doi:10.3758/s13423-013-0572-3. URL https://link.springer.com/10.3758/s13423-013-0572-3
-
[7]
Guido W. Imbens and Donald B. Rubin. Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press, New York, 1 edition edition, 4 2015. ISBN 978-0-521-88588-1. 00005
work page 2015
-
[8]
When the Alpha is the Omega : P - Values , `` Substantial Evidence ,'' and the 0.05 Standard at FDA
Lee Kennedy-Shaffer. When the Alpha is the Omega : P - Values , `` Substantial Evidence ,'' and the 0.05 Standard at FDA . Food and drug law journal, 72 0 (4): 0 595, 2017. URL https://pmc.ncbi.nlm.nih.gov/articles/PMC6169785/
work page 2017
Show all 11 references
-
[9]
Deborah G. Mayo. The statistics wars and intellectual conflicts of interest. Conservation Biology, 36 0 (1): 0 e13861, 2021
2021
-
[10]
Interview with D on R ubin
Don Rubin. Interview with D on R ubin. Observational Studies, 8 0 (2), 2022
2022
-
[11]
Personal communication
Philip Stark. Personal communication. And he has a sign on his office door, October 2024
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.