Pith. sign in

REVIEW 4 major objections 4 minor 18 references

The p<0.05 ritual became science's gatekeeper because its mechanical, context-free procedure scaled with postwar massification—not because it solved inference.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 18:06 UTC pith:KSFWOR7K

load-bearing objection A plausible and genuinely useful institutional re-framing of the NHST puzzle, but the central causal claim—massification selected procedural self-sufficiency—remains a functionalist conjecture, not a demonstrated result. the 4 major comments →

arxiv 2603.14757 v3 pith:KSFWOR7K submitted 2026-03-16 stat.OT

The Rise of Null Hypothesis Significance Testing (NHST): Institutional Massification and the Emergence of a Procedural Epistemology

classification stat.OT
keywords NHSTp-valuestatistical significanceinstitutional massificationprocedural self-sufficiencystatistical reformhistory of statisticsscience education
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Null Hypothesis Significance Testing—the p≤0.05 ritual—has survived sixty years of criticism because it never primarily served statistical inference. This paper argues NHST became dominant as a social technology for institutional massification: after World War II, American universities had to train and vet enormous numbers of researchers quickly, and NHST's mechanical, context-stripping procedure allowed that. The paper treats NHST's weaknesses—ignoring modeling assumptions, replacing judgment with a fixed cutoff—as functional features under massification, not bugs. If true, statistical reform cannot be achieved by better teaching or by swapping in a better test; the institutional conditions that select for procedural epistemology must change too.

Core claim

The paper's central claim is that NHST is the offspring of a forced marriage between two incompatible testing frameworks, and that its rise cannot be explained by its technical merits. What NHST solved was not the technical problem of statistical inference but the social problem of rapidly bringing a massive number of people into scientific practice. Its procedural self-sufficiency—purging research context, specifying decisions in advance, replacing expert judgment with a routine—made it independently redeployable across heterogeneous settings with uneven expertise. NHST therefore became the obligatory passage point in many postwar scientific fields, and its technical slippages became the ve

What carries the argument

The key object is NHST itself, defined here as a hybrid decision routine: state a null hypothesis (typically 'no effect'), compute a p-value, reject if p≤0.05, otherwise 'accept' the null—with no specified alternative hypothesis and no check of modeling assumptions. The load-bearing concept is 'procedural self-sufficiency': the procedure itself performs adjudication and persuasion, independent of local context, so it can be taught by non-specialists, used by novices, and audited by administrators. The paper argues this property is what makes NHST scalable and why it became institutionalized.

Load-bearing premise

The paper's load-bearing premise is that postwar institutional massification was the main selective force favoring NHST, an assumption asserted through historical narrative rather than tested against rival explanations such as textbook economics, disciplinary gatekeeping, or simple inertia.

What would settle it

A comparative historical study tracking a rapidly expanding postwar scientific field that did not adopt NHST—or showing that NHST became entrenched before the enrollment surge—would challenge the claim. More directly, if adoption timing across disciplines does not track growth rates in students, doctorates, and journal output, the massification explanation loses its core support.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Statistical reform through improved education alone is unlikely to dislodge NHST, since the source of its dominance is institutional, not cognitive.
  • Any replacement method—including Bayesian approaches—will face the same massification pressure and may be simplified into a threshold ritual, as the paper notes is already happening with Bayes factors.
  • Institutions that adopted NHST to coordinate mass science now depend on it as a resource-allocation and credentialing device; reform must redesign those incentives rather than only teach p-values correctly.
  • NHST's persistence is thus not evidence that researchers are ignorant or lazy; it is the expected outcome of a system built to scale.
  • The historical pattern suggests that technically demanding practices, when massified, tend to drift into standardized procedures that abandon their epistemic origins.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the massification explanation is right, the adoption of NHST should track growth rates in enrollment, doctorate production, and journal output across disciplines and countries; this is checkable and would give the thesis sharper teeth.
  • The argument suggests a caution for current Bayesian reform: any procedure that becomes a standard for evaluating claims will tend to be used as a black box, so reform should focus on institutional evaluation practices—what journals, funders, and promotion committees reward—not just on introducing a different test.
  • The paper's historical focus on postwar United States leaves open whether contemporary massification in other countries is re-creating the same procedural epistemology; that would be a natural extension.
  • A comparative test could look at fields that expanded rapidly but did not adopt NHST, or at periods when NHST was entrenched before the enrollment surge; such cases would help separate massification from other drivers.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper examines the historical rise and persistence of Null Hypothesis Significance Testing (NHST) in the postwar American academy. It argues that NHST, despite its well-documented technical deficiencies, became dominant not because it solved statistical inference but because it served as a 'social technology of institutional massification': its context-stripping, mechanically procedural character allowed scientific training and research to scale rapidly under conditions of huge enrollment growth, a shortage of trained statisticians, and federal investment in science education. The author proposes that NHST achieved 'procedural self-sufficiency,' making it an 'obligatory passage point' in many fields. The argument is developed through a historical narrative from Fisher and Neyman-Pearson, through wartime and Cold War expansion, to the stabilization of a fused, simplified testing procedure in textbooks and teaching.

Significance. If the institutional-massification thesis is correct, it would provide a genuinely novel explanation for NHST's persistence against six decades of statistical criticism and for the limited impact of reforms aimed at statistical education. The paper draws on credible primary and secondary sources (NRC 1947, Wilks 1951, Lehmann 2008, Hotelling 1948, Gigerenzer et al. 1989, Box 1978) and usefully foregrounds the institutional and demographic context generally neglected in purely cognitive or pedagogical accounts. The conceptual proposal of 'procedural self-sufficiency' is a potentially generative lens, and the paper makes a falsifiable general prediction: wherever institutional massification occurs, standardized, context-stripping procedural methods will tend to dominate. The presentation is clear and historically informed.

major comments (4)
  1. [NHST as a Technology of Institutional Massification (pp. 20-22)] The central causal claim—that NHST's procedural features were 'selected' by postwar institutional massification—is asserted rather than demonstrated. The evidence presented establishes covariation: rapid expansion, a statistician shortage, and eventual NHST dominance. But the argument moves from 'these features are well suited to massification' to 'NHST rose to the challenge' without testing against alternatives such as textbook commercial incentives (e.g., Snedecor's cookbook), journal editorial policies and gatekeeping, path dependence in curricula, disciplinary jurisdictional battles (which the paper itself documents at Berkeley), or the low cognitive overhead of a dichotomous rule. To support the primary-causal claim, the paper needs either a comparative design (e.g., disciplines with different rates of massification, countries without the same postwar surge) or explicit consideratio
  2. [Procedural self-sufficiency (pp. 6, 20-22)] The concept of 'procedural self-sufficiency' is defined largely by the persistence it is asked to explain: NHST is said to spread because it is procedurally self-sufficient, while procedural self-sufficiency is inferred from NHST's spread. This risks circularity. The paper should specify operational indicators of procedural self-sufficiency that are independent of eventual dominance—for example, the degree to which a method's prescribed steps can be executed with minimal domain expertise, the extent to which outputs are auditable, or the variation across methods in the need for case-by-case judgment. The paper would be stronger if it identified a testable prediction, such as a positive correlation between massification pressure and the adoption of the most mechanical features of NHST across disciplines or over time.
  3. [Disciplinary scope and counterexamples (pp. 18-22, the industrial quality control case)] The paper's own historical material shows that Neyman-Pearson decision rules stabilized in industrial quality control without university massification (p. 11, 22), implying that proceduralism can take root in other niches and raising the question of which specific feature, if any, was uniquely selected by massification. The thesis would be considerably strengthened by addressing such counterexamples explicitly and by comparing contexts where NHST did not dominate or where dominance preceded massification. As it stands, the scope of the claim—'many postwar scientific fields'—is narrower than the US-focused evidence, and the generalization to other countries and disciplines is not examined.
  4. [Alternative mechanisms acknowledged but not weighed (pp. 12-17, 20-21)] The paper mentions, in passing, institutional rivalries, the role of Neyman's PhD pipeline, and the lobbying of statistics departments, but these are each plausible causal forces in their own right. For example, Neyman's success at Berkeley and the reproduction of academic statisticians could explain the spread of a decision-theoretic vocabulary independent of massification. The manuscript should explicitly separate the explanation of NHST's initial stabilization from the explanation of its continued reproduction, and should state what evidence would distinguish a massification-driven account from a pure supply-side account centered on the interests of a growing statistics profession. This is a load-bearing distinction that the current narrative blurs.
minor comments (4)
  1. [Throughout] There are numerous typographical and reference-consistency errors: 'procedural self-efficiency' should be 'procedural self-sufficiency' (Section 3); 'Salvage, 1976' should be 'Savage, 1976' in the reference list; 'Humes' is spelled 'Hummes' in the references; 'Biometirka' should be 'Biometrika'; 'V ol.' has a stray space; and the figure axis labels show broken year ranges like '194 1' (Figures 4–6). These should be corrected.
  2. [References (e.g., Agresti 2023, Sismondo 2010)] Some references are incomplete or incorrectly formatted: Agresti (2023) is listed with a volume/issue but the page range is missing; Sismondo (2010) lists 'Wiley-Backwell' instead of 'Wiley-Blackwell'; the Huberty citation includes an unexpected URL fragment; and the Smith/Simmons reference appears to be to a conference presentation but lacks full publication details. Please review all references for completeness.
  3. [Figure captions] Figures 4–6 lack full source descriptions in the captions. The text mentions sources (NCES, NSF NCSES, OpenAlex) but the captions themselves are insufficiently informative for a standalone read. Add brief source notes to each figure.
  4. [p. 6 (shadow analogy)] The 'shadows and the object' analogy is engaging but could be tightened; as written it risks implying that statistical inference is as automatic as vision, which undercuts the later point that context-stripping is problematic. A sentence clarifying the limits of the analogy would help.

Circularity Check

0 steps flagged

No significant circularity: the NHST-massification argument is a self-contained historical interpretation, and the only self-citation is non-load-bearing.

full rationale

This is a historical-sociological essay, not a formal derivation, and I find no step in which a claimed result is equivalent to its inputs by construction. The central concept, 'procedural self-sufficiency,' is given independent content: it refers to standardized, context-stripping procedures whose outputs can circulate without repeated epistemic negotiation (Section 'NHST as a Technology of Institutional Massification'). That content is documented from historical sources (Snedecor's cookbook, Hotelling's complaints about non-specialist teachers, the Fisher/Neyman-Pearson textbook fusion) rather than inferred from NHST's persistence alone. The causal step from those features to postwar dominance is an interpretation, not a reduction: the paper follows a mechanism (mechanical procedure → low coordination costs → scalable instruction/adjudication) rather than defining 'self-sufficiency' as 'that which spreads.' The sentence 'NHST scales well because it is procedurally self-sufficient' is a summary of that mechanism, not a tautology. The single self-citation (Ting & Greenland 2024) supports a peripheral point about deterministic framing and is not load-bearing; the acknowledgment that Sander Greenland suggested the term 'programmable' is an intellectual credit, not an imported theorem. The paper's functionalist selection claim is open to empirical challenge—it lacks a systematic comparative test against alternative mechanisms—but that is a question of evidence and causal identification, not circularity under the strict standard applied here. Score 2 reflects the presence of one minor non-load-bearing self-citation and the interpretive character of the causal argument, without any definitional, fitted-prediction, or self-citation-chain circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 1 invented entities

The paper is a qualitative historical-sociological argument. It introduces no numerical free parameters, but its causal thesis rests on several domain assumptions about postwar institutional history and on a functionalist construct that is not independently measured.

axioms (4)
  • domain assumption Actor-Network Theory's translation and black-box concepts can be transferred from laboratory science to statistical methods.
    The paper explicitly takes ANT as its point of entry (Introduction, p. 3) and uses Latour's concepts to describe NHST's spread; if this transfer is invalid, the explanatory frame fails.
  • domain assumption Postwar American institutional massification was the primary selective pressure that stabilized NHST.
    The central causal claim presupposes that expansion of enrollments, doctorates, and journal output (Figures 4-6) drove methodological dominance. Alternative drivers are mentioned but not assessed.
  • ad hoc to paper NHST's context-stripping and procedural features were functionally adaptive for massification, not incidental.
    This is the core explanatory step: NHST's flaws are called 'functional features rather than flaws' because they aided scaling (Abstract, p. 21). This selectionist reading is inferred from NHST's persistence rather than independently evidenced.
  • domain assumption The historical documents and retrospective accounts cited are reliable evidence for the state of statistical training.
    The argument relies on quotations from Hotelling, Lehmann, Wilks, NRC/RSS reports, and textbook reviews; no independent verification or systematic sourcing protocol is provided.
invented entities (1)
  • Procedural self-sufficiency (as a property of inferential technologies) no independent evidence
    purpose: Explains why NHST could travel across heterogeneous postwar settings without renegotiating epistemic assumptions.
    Introduced as an analytic construct (pp. 21-22); it is inferred from NHST's observed persistence and used as the mechanism of that persistence, so it lacks an independent falsifiable handle.

pith-pipeline@v1.3.0-alltime-deepseek · 16689 in / 8834 out tokens · 92598 ms · 2026-08-02T18:06:31.911887+00:00 · methodology

0 comments
read the original abstract

It has long been a puzzle why, despite sustained reform efforts, many applied scientific fields remain dominated by Null Hypothesis Significance Testing (NHST), a framework that dichotomizes study results and privileges "statistically significant" findings. This paper examines that puzzle by situating the development and rise of NHST within its historical and institutional context. Taking Actor-Network Theory as a point of entry, the analysis identifies the conditions under which particular inferential technologies stabilize and endure. The analysis shows that, although NHST does not resolve the technical problem of statistical inference, it came to dominate as a social technology that addressed the most pressing institutional challenge of the postwar period: the mass expansion of scientific networks. Under conditions of rapid institutional growth, NHST's technical slippages--purging research context and replacing epistemic judgment with mechanical procedures--became functional features rather than flaws. These features enabled procedural self-sufficiency across settings marked by heterogeneous goals and uneven expertise, thereby sealing NHST's position as the obligatory passage point in many postwar scientific fields.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

18 extracted references · 2 linked inside Pith

  1. [1]

    statistically significant

    1 The Rise of Null Hypothesis Significance Testing (NHST): Institutional Massification and the Emergence of a Procedural Epistemology Carol Ting* Department of Communication, University of Macau Version 2 – March 2026 Note. This paper examines the rise of Null Hypothesis Significance Testing (NHST) from a historical and sociological perspective. While eng...

  2. [2]

    significant

    Reject the null if 𝑝≤0.05; otherwise accept the null This mechanical routine concludes with labeling results as either significant or insignificant. Unaware of the peculiar origin of the term “significant” (Shafer, 2020), many NHST users treat the label as self-explanatory and perceive little need for interpretation. The dichotomous classification becomes...

  3. [13]

    significant

    As the technically demanding original texts passed through successive rounds of translation by textbook authors and instructors, confusion and misconception were not only likely but structurally reinforced. Huberty’s review (1993) of popular statistics textbooks suggests that a confusing mixture of terms had already appeared in textbooks before the war. T...

  4. [16]

    programmable

    Bush’s skillful persuasion paved the way for the establishment of National Science Foundation and the sustained injection of federal funds into higher education and scientific research. Under this national zeitgeist, higher education expanded rapidly, and applied science departments, programs, and research centers proliferated across university campuses (...

  5. [39]

    Replication Crisis

    https://doi.org/10.2307/2342435 Fisher, & A., R. (1956). Statistical methods and scientific inference. Oliver and Boyd. Gigerenzer, G., Porter, T., & Daston, L. (1989). The Empire of Chance: How Probability Changed Science and Everyday Life. Cambridge University Press. Gigerenzer, G. (2004). Mindless statistics. The Journal of Socio-Economics, 33(5), 587–...

  6. [106]

    forced marriage

    describes as a chimerical ancestry—the offspring of a “forced marriage” between two theoretically incompatible frameworks. The two sides fought long and hard, but as their ideas disseminated through the rapid postwar expansion of scientific networks, neither remained intact. I argue that massive post war institutional expansion (massification) of American...

  7. [227]

    p < 0.05

    https://www.science.org/doi/10.1126/science.abd7628 UN Economic and Social Council. (1949). An international programme for education and training in statistics. https://unstats.un.org/unsd/statcom/doc49/1949-56-EducationStats.pdf Wasserstein, R. L., & Lazar, N. A. (2016). The ASA’s Statement on p-Values: Context, Process, and Purpose. The American Statist...

  8. [690]

    14 The success and visibility of statisticians’ participation in the war effort spurred post-war demand for trained statisticians, as both public and private sectors sought to improve operation through statistical analysis (National Research Council [NRC], 1947; UN Statistical Commission, 1949; E. S. Pearson, 1959). Academia was another major driver of de...

  9. [1900]

    On the Mathematical Foundations of Theoretical Statistics

    and William Gosset’s seminal development of the small sample distribution later used in t-test (Student,1 1908). However, prior to the work of Ronald A. Fisher (1890-1962), there was no unified logical framework for reasoning from sample to population: the concepts of population and sample were not yet clearly distinguished, and statistical tools were oft...

  10. [1911]

    As an emerging discipline, statistics departments typically began with graduate programs only, oriented toward specialist training rather than broad instructional demand

    was the only statistics department before the war. As an emerging discipline, statistics departments typically began with graduate programs only, oriented toward specialist training rather than broad instructional demand. In brief, the pre-war capacity for statistical training was quite limited, and the institutional pipeline required for mass undergradua...

  11. [1925]

    at once polished and awkward

    as a practical manual for researchers. In it, Fisher guides users through common inferential tasks with worked examples showing how his method could be applied to data. To facilitate significance testing, the book also provides tables of critical values for test statistics, allowing researchers to determine whether their results met chosen cutoff threshol...

  12. [1938]

    Over the next sixteen years, Neyman worked tirelessly to pursue this goal and finally succeeded in

    As his student Lehmann recalled, the department chair who made the appointment likely got more 13 than he had hoped for, as Neyman’s ambition was to establish an independent statistics department (Lehmann, 2008). Over the next sixteen years, Neyman worked tirelessly to pursue this goal and finally succeeded in

  13. [1944]

    Personnel and Training Problems Created by the Recent Growth of Applied Statistics in the United States

    brought more than two million veterans to college and university campuses (Olson, 1973). According to Pulitzer prize winner Edward Humes (2006), the GI Bill supported the education of approximately 450,000 engineers and 91,000 scientists, and the resulting influx of students placed severe strain on existing training capacity. Government funded research fu...

  14. [1945]

    Magna Carta of America Science

    Seizing the moment when American trust in science was at its historical peak, this document—often hailed as the “Magna Carta of America Science” (Thorp, 2020)—argued powerfully that basic scientific research was a national strategic resource essential to public health, national security, and economic welfare. Bush emphasized the importance of maintaining ...

  15. [1951]

    forced marriage

    p.3) Under these conditions, reliance on routinized procedures and standardized rules increasingly substituted for deep statistical expertise and contextual judgment. Of course, as many statistics departments were established in the two decades after World War II, it would be naïve to ignore the issue of resource allocation in such debates. Growing cluste...

  16. [1955]

    the upper echelons of North American Statistics and beyond, for a decade or more

    Lehmann describes Neyman’s struggle as follows: The struggle to convert a one-man appointment as professor of mathematics into a substantial separate department of statistics did not, of course, take place in a vacuum. It required the acquisition of a faculty, the associated office and laboratory space, a corresponding expansion of the course pro- gram, a...

  17. [1986]

    a technology of institutional massification

    as a point of entry. Concepts of black box and the translation model of fact-building are useful for analyzing the emergence and spread of NHST. At the same time, by anchoring the NHST network in the postwar period of rapid expansion of science, this study foregrounds the role of institutional incentives and demographic change, situating these processes w...

  18. [2017]

    helping to satisfy immediate instructional needs

    p. 40). Commenting in his Presidential Address at the 1950 Annual Meeting of the American Statistical Association, Sam Wilks acknowledged the contribution of self-taught instructors in “helping to satisfy immediate instructional needs”, while also emphasizing the resulting quality problems in statistical education: As a whole and as it now stands, our sta...