{"id":"b4578b50-7939-4670-901d-7ed8a3abe35e","arxiv_id":"2501.03457","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Statistical testing is reframed as bureaucratic rule-making rather than inference about truth.","lead":"An essay argues that the main role of statistical tests is to give organizations clear rules for approving drugs, software, and papers, not to prove scientific truths. It proposes the term 'ex ante policy' for rules fixed before data collection and says this view could settle old fights about p-values.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim overstates what Beer's maxim can establish: statistical methods also function as ex post inference, so selecting regulatory uses as 'what statistics does' is an unargued premise.","rationale":"The reader identified the same weakest assumption: the argument depends on adopting Beer's maxim and then treating regulatory applications as representative of what statistics does. I agree, and I would sharpen the point: Beer's maxim is underdetermined because any system has multiple effects. The paper never justifies why the regulatory effect should be privileged over the inferential effect. This is not a technical error or an internal inconsistency, but it is a gap in the argument's central inference. Because the paper is a commentary rather than an empirical study, the verdict UNVERDICTED remains appropriate: the conceptual claim is coherent but not established. My concrete test would provide evidence on whether the premise about actual statistical practice holds, but even if it fails, the paper could be read more modestly as highlighting an underappreciated role rather than asserting the singular purpose of statistics. Thus no change to the reader's verdict is needed.","tokens_in":6274,"tokens_out":2558,"duration_ms":29357,"concrete_test":"Construct a representative corpus of applied statistical practice—for example, all Applied Research articles in JASA over the past decade—and code each study's stated purpose as either (a) choosing a policy or rule to govern future actions or (b) estimating or inferring an effect, parameter, or theory. If category (b) constitutes a substantial fraction (say more than 20%), then 'what statistics does' is not predominantly regulation, directly undermining the Beer-maxim premise. Additionally, apply the same coding to the paper's own examples, especially the Salk vaccine trial, which was framed at the time both as evidence about vaccine efficacy and as the basis for a policy decision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, stated in the abstract and Section 4 as 'the purpose of statistical tests is regulation,' rests on Stafford Beer's maxim that 'the purpose of a system is what it does.' Applied literally, this maxim is ambiguous: a system does many things, in many contexts. The paper chooses regulatory examples—FDA drug approval, A/B testing, journal review—and then asserts that these are what statistics does. But the same methods are routinely used for ex post inference: estimating treatment effects in epidemiology, measuring physical constants, forecasting, and exploratory data analysis. The paper even acknowledges this by contrasting 'ex ante policy' with 'ex post inference' in Section 2, but it never gives a criterion for why the regulatory uses are the ones that fix statistics' purpose. Without such a criterion, the conclusion that statistics is 'the mathematics of bureaucracy' does not follow from the evidence presented. The claim might be true as a partial account, but the paper's own examples do not establish that regulation is the primary or defining function. This is a load-bearing gap because the entire reframing depends on identifying a single dominant purpose.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This commentary argues that statistical methods, especially randomized controlled trials and hypothesis tests, are best understood not as instruments of epistemic inference but as 'ex ante policy': rules specified before data collection to govern future actions. The author introduces this term, illustrates it with FDA drug approval standards, A/B testing, and journal peer review, and reformulates the z-test as a correlation threshold (Eq. 1). He then contrasts ex ante policy with ex post inference, argues that the ex ante frame resolves debates about p-values and confidence intervals, and concludes that 'the purpose of statistical tests is regulation' (Section 4). The paper is written as an essay intended to reframe how statisticians and methodologists think about their field.","tokens_in":6408,"tokens_out":2874,"duration_ms":28948,"significance":"The paper identifies a genuine and underappreciated function of statistical methods: pre-specified decision rules provide transparency, fairness, and accountability in governance. The mathematical reformulation in Eq. (1) is standard and clearly explained, and the historical examples (Bradford Hill, FDA, Kefauver-Harris) are relevant and cited. If the reframing were accepted, it could productively reorient pedagogy, methodology, and discussions of p-values and confidence intervals. However, the central claim as stated is not established: the paper offers an interpretive lens rather than a testable theory, and its conclusion that regulation is the purpose of statistics rests on an unargued selection of examples. The paper is transparent about its own framing and contains no fitted parameters, which is appropriate for a commentary; its strength is in the clarity of the proposed distinction rather than in rigorous proof.","major_comments":[{"comment":"The load-bearing claim, 'The purpose of statistical tests is regulation,' is not entailed by the evidence presented. The argument uses Stafford Beer's maxim 'the purpose of a system is what it does' (opening), but a system does many things, and the paper does not provide a criterion for why the regulatory uses (FDA approvals, A/B tests, journal review) fix the purpose of statistics, rather than ex post inference uses such as estimating treatment effects in epidemiology, measuring physical constants, or forecasting. Section 2 explicitly distinguishes ex ante policy from ex post inference, and the same methods are routinely used in both modes. Without a principled basis for weighting these uses, the conclusion that regulation is the defining purpose does not follow; at most, the paper shows that regulation is an important use. The paper also appears to contradict itself by saying 'Rulemaking is certainly not the singular valuable application of statistics' yet later stating a singular purpose. Please reconcile these statements and either soften the conclusion to 'a' purpose or justify the priority claim.","section":"Section 4 and opening"},{"comment":"The definition of ex ante policy as 'the specification of rules by some regulatory body that governs acceptable future actions' is broad enough to cover almost any pre-specified statistical procedure, and the paper then concludes that 'most common Frequentist methods' are examples. This creates a circularity risk: if every pre-specified rule is defined as policy, then finding that statistics is policy is true by definition. The paper needs a limiting criterion—for example, a distinction between rules adopted by an actual governance body and the merely procedural features of a test—or an independent benchmark for what counts as policy. As written, the conceptual claim is too elastic to be informative.","section":"Section 1"},{"comment":"The paper asserts that 'science doesn't progress via Popperian means' and that 'there are no computable Bayesian updates that crunch empirical data into crisp posteriors.' These are strong empirical and historical claims used to downplay the importance of ex post inference. They are unsupported by evidence or argument, and the central thesis does not depend on the strong versions of these claims. Please either provide supporting citations or hedge these statements so that the argument does not rely on contested historiography.","section":"Section 2"}],"minor_comments":[{"comment":"In the list of designs, 'differences and differences' should be 'difference-in-differences.'","section":"Section 4"},{"comment":"'AB tests' should be written as 'A/B tests.'","section":"Section 4"},{"comment":"The Imbens and Rubin reference contains formatting errors: '1 edition edition' and the month '4 2015' should be cleaned up.","section":"References"},{"comment":"The Greek letter in 'ˆβ' is garbled in the text; please ensure it renders correctly.","section":"Footnote 1"}],"recommendation":"major_revision","confidential_remarks":"This is a viewpoint/commentary piece, so the standard of empirical proof does not fully apply. However, the central thesis—that statistics' purpose is regulation—needs either to be substantially qualified ('a purpose' rather than 'the purpose') or supported by a criterion for why regulatory uses are definitive. The report's major comments focus on this selection issue, which is the main obstacle to acceptance. The paper is well written and engaging, and the revised version could be a valuable contribution if the overreach is addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a good, readable essay that gives a name to a real phenomenon — the use of pre-specified statistical rules as regulatory instruments. The 'ex ante policy' framing is genuinely handy for talking about FDA trials, A/B tests, and journal gatekeeping without getting tangled in p-value metaphysics. But the paper's stronger claim, that the purpose of statistics is regulation, is not supported by the argument it offers.\n\nWhat's good: the paper is clear and honest. Equation (1) correctly recasts the z-test as a correlation threshold, which is a nice pedagogical point. The distinction between ex ante guarantees and ex post inference is useful and does dissolve some pseudo-debates. The paper is also candid that it is a commentary, not a research claim, and it engages seriously with the literature it cites.\n\nWhere it goes soft: the central inference from 'statistical methods are often used for regulation' to 'the purpose of statistics is regulation' relies on Stafford Beer's maxim, but the maxim doesn't do the work assigned to it. A system does many things; picking out regulatory uses as the ones that fix purpose needs a criterion, and none is given. The same methods are used for measurement, forecasting, and exploratory data analysis. The paper acknowledges this contrast in Section 2 but doesn't explain why the regulatory examples are the defining ones. Also, the claim that 'science doesn't progress via Popperian means' is a strong empirical statement with no supporting evidence; it's not needed for the ex ante policy point and it weakens the essay.\n\nNone of this makes the paper worthless. As a provocation and a vocabulary, it's useful. Statisticians who teach or think about the role of their field will get something from it, and it would be a good reading-group piece.\n\nRecommendation: I'd send it to peer review as a commentary — a serious editor could reasonably accept it for discussion — even though I'd push back on the overstated thesis. It's not a technical result, so a methods journal might say no, but an opinions/commentary venue is right.","headline":"A clearly written opinion piece that usefully names the regulatory role of statistical rules, but the stronger claim that regulation is the purpose of statistics is not established by the argument.","tokens_in":6956,"tokens_out":2606,"would_cite":true,"duration_ms":25015,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62-02","62A01","62F03"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that randomized controlled trials and statistical tests are best understood as 'ex ante policy': rules fixed before data collection that govern future actions, making statistics' primary purpose regulation rather than…","keywords":["ex ante policy","statistical inference","statistical testing","bureaucracy","randomized controlled trials","p-values","regulation","confidence intervals"],"falsifier":"A broad survey of statistical use across fields such as astronomy, particle physics, genomics, and economic forecasting would falsify the claim if it found a substantial class of applications in which no pre-specified action rule exists and results are used only to update scientific understanding.","tokens_in":6013,"feed_emoji":"📋","tokens_out":8912,"duration_ms":80237,"temperature":0.7,"pith_summary":"This paper argues that what statisticians call causal inference and hypothesis testing is, in practice, rulemaking. Randomized controlled trials and statistical tests are best understood as 'ex ante policy': procedures and evidentiary bars fixed before data are collected, intended to govern subsequent actions such as approving a drug or shipping a software feature. On this reading, the RCT's main role is regulation, and statistics is the mathematics of bureaucracy. The benefit of the framing is that it sidesteps long-running disputes about the meaning of p-values or Bayesian versus frequentist inference, replacing them with questions about which transparent, fair rules a community wants.","feed_headline":"The purpose of statistical tests is regulation","feed_subtitle":"Reframing randomized trials and p-values as ex ante policy, not inference about truth.","key_machinery":"The load-bearing construct is 'ex ante policy,' defined as a rule specified by a regulatory body before data collection that dictates which future actions are acceptable. The paper's worked example rewrites the two-sample proportions test as a correlation threshold, approving a treatment when $R(X,Y) \\ge t/\\sqrt{n}$, with $t=1.96$ yielding the conventional z-test at level $\\alpha=0.05$ and an equivalent chi-squared test. This identity shows that a p-value threshold is an arbitrary evidentiary bar, chosen by convention, not a theorem. Confidence intervals are then defined as randomized algorithms with an ex ante success probability, and Bayesian decision procedures count as ex ante policy whenever prior, likelihood, and utility are fixed in advance.","core_discovery":"The paper's central claim is that the purpose of a statistical test is regulation. It grounds this in the observation that randomized trials and significance tests, whatever their epistemic ambitions, serve as pre-specified rules that decide whether a drug comes to market, whether a software feature ships, and whether a paper is published. The author introduces 'ex ante policy' as the umbrella term for such rules, and argues that both frequentist procedures and Bayesian decision theory fit under it, because both require the decision-relevant ingredients to be fixed before data collection. The ex ante/ex post distinction is then used to explain persistent confusion: tests carry verifiable, prescriptive guarantees, whereas inference about whether a theory is true is complex, subjective, and not reducible to a numerical algorithm. Consequently, debates over thresholds and rituals are better understood as debates over regulatory conventions than over the foundations of knowledge.","pith_inferences":["An implicit corollary is that the replication crisis can be read as a regulatory failure—rules that no longer encode a community's values—rather than as a failure of inference; this reading suggests fixes like renegotiating evidentiary bars rather than searching for better estimators.","The ex ante/ex post split predicts that the perceived authority of a statistical result depends less on its epistemic justification than on the transparency and perceived fairness of the rule that produced it, which could be tested in surveys of stakeholders.","A natural extension is to classify published statistical applications as ex ante policy or ex post inference and compare their evidentiary standards, which would provide a quantitative check on whether regulation really dominates practice."],"forward_implications":["Debates about the correct p-value threshold become deliberations about a regulatory convention, which can be resolved by stakeholder agreement rather than by epistemology.","Confidence intervals should be communicated as algorithms whose coverage probability is over prospective repetitions, ending the temptation to read a realized interval as an epistemic statement.","Bayesian and frequentist methods that fix their decision rules in advance enter the same framework, so methodology choices can be compared on transparency and fairness rather than on philosophical loyalty.","A methodology research agenda opens for designing statistical rules, including the question of when observational data are acceptable as a basis for policy."],"supporting_citations":[{"why":"Supplies the cybernetic maxim 'the purpose of a system is what it does,' used to fix statistics' purpose from observed practice.","marker":"Beer [1979]"},{"why":"Documents the drug regulator's de facto standard of two trials at the 0.05 level, the paper's main example of ex ante policy.","marker":"[Kennedy-Shaffer, 2017]"},{"why":"Provides the correlation-based reformulation of tests used in the paper's threshold example.","marker":"Cohen [1969]"},{"why":"Names the null ritual; the paper recasts this ritual as regulation rather than epistemology.","marker":"Gigerenzer [2004]"},{"why":"Shows that causal inference's stated purpose is policy, supporting the bureaucratic reading.","marker":"Imbens and Rubin [2015]"},{"why":"Points to the type of methods research that a policy-oriented agenda would pursue.","marker":"Banerjee et al. [2020]"}],"fun_headline_variants":["Statistics as bureaucratic rule","Why p-values are policy, not proof","Ex ante policy: the real job of statistical tests","Reframing trials as regulation, not inference","Statistical tests: pre-specified rules for decisions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on accepting that the purpose of a system is what it does and on treating regulatory applications like drug trials, feature rollouts, and journal review as representative of statistics as a whole; if measurement, forecasting, or exploratory modeling are equally central, the conclusion that statistics' purpose is regulation does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Statistics as bureaucratic rule","Why p-values are policy, not proof","Ex ante policy: the real job of statistical tests","Reframing trials as regulation, not inference","Statistical tests: pre-specified rules for decisions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1236,"prompt_tokens":800,"completion_tokens":436,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":416,"completion_tokens_details":{"reasoning_tokens":370}},"tokens_in":416,"tokens_out":436,"duration_ms":4350,"temperature":1.0,"reasoning_tokens":370,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:52:26.589898+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A broad survey of statistical use across fields such as astronomy, particle physics, genomics, and economic forecasting would falsify the claim if it found a substantial class of applications in which no pre-specified action rule exists and results are used only to update scientific understanding.","supporting_citations":[{"cited_title":"The Heart of Enterprise","cited_arxiv_id":null,"evidence_quote":"Supplies the cybernetic maxim 'the purpose of a system is what it does,' used to fix statistics' purpose from observed practice."},{"cited_title":"When the Alpha is the Omega : P - Values , `` Substantial Evidence ,'' and the 0.05 Standard at FDA","cited_arxiv_id":null,"evidence_quote":"Documents the drug regulator's de facto standard of two trials at the 0.05 level, the paper's main example of ex ante policy."},{"cited_title":"Statistical Power Analysis for the Behavioral Sciences","cited_arxiv_id":null,"evidence_quote":"Provides the correlation-based reformulation of tests used in the paper's threshold example."}],"review_version":1}