Pith. sign in

REVIEW 4 major objections 3 minor 21 references

The ultimate issue error in scientific inference: mistaking parameters for hypotheses

T0 review · 4 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The probability of a parameter is not the probability of a hypothesis, so p<0.05 alone cannot support a claim like 'the treatment works'.

desk verdict A clear, honest restatement of the parameter-vs-hypothesis distinction, but the vitamin D 'zero evidence' conclusion is an artifact of the author's simplifying choices, not a consequence of the WoE algebra. read the letter →

arxiv 2411.15398 v2 pith:5J7WM665 submitted 2024-11-23 stat.ME q-bio.QMstat.AP

classification stat.MEq-bio.QMstat.AP
keywords weightofevidenceultimateissueerrorhypothesistestingp-valuesBayesfactorsstudydesignbiasvitaminDCOVID-19likelihoodratio
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that standard statistical tests answer a different question from the one scientists usually pose. A p-value, Bayes factor, or posterior probability describes a parameter in a model; it does not give the probability that a scientific hypothesis such as 'this drug is effective' is true. The paper proposes a weight-of-evidence (WoE) approach, a logarithmic likelihood ratio that folds in study design, bias, power, and prior plausibility, and applies it to a widely cited vitamin D and COVID-19 study. If the argument is right, statements like 'the treatment is effective (p<0.05)' are non sequiturs, and many published negative results carry little or no evidence for the null hypothesis.

What carries the argument

The central object is the augmented likelihood ratio, expressed as weight of evidence in decibel units: $$\text{WoE}(H_1:E^*,I) = 10\log_{10}\left[\frac{P(E^*\mid H_1,I)}{P(E^*\mid H_0,I)}\right] + 10\log_{10}\left[\frac{P(H_1,I)}{P(H_0,I)}\right].$$ It carries the argument by converting a study result into support for a hypothesis, with $P(E^*\mid H_0,I)$ the false-positive probability adjusted for bias and $P(E^*\mid H_1,I)$ the power adjusted for design flaws. The logarithmic scale makes independent pieces of evidence additive, and the prior term lets background plausibility enter explicitly. In the worked examples, a significant result with 60% power and a 15% false-positive rate gives WoE $\approx 6$ (0.8 probability for the hypothesis), while a negative result gives WoE $\approx -3$ (0.67 probability for the null), showing that both outcomes are nearly uninformative.

What would settle it

Re-run the vitamin D analysis using a likelihood ratio built from the continuous test statistic or the full logistic-regression likelihood instead of coding the result as $p>0.05$. If the resulting weight of evidence is not close to the paper's $-0.75$, then the binary $\alpha$/power collapse is what produces the 'zero evidence' conclusion.

Watch

Extended reading notes

Core claim

The paper's central claim is that the probability of a parameter and the probability of a hypothesis are distinct quantities, and conflating them—the 'ultimate issue error'—makes statements such as 'the treatment is effective (p<0.05)' non sequiturs. A parameter is a quantitative feature of a statistical model; a hypothesis is a testable proposition whose truth depends on background information, construct validity, competing explanations, and prior plausibility. The paper argues that no direct quantitative relationship links p-values or Bayes factors to hypothesis probabilities, and that moving from parameter evidence to hypothesis support requires the augmented likelihood ratio $$\text{WoE}(H_1:E^*,I) = 10\log_{10}\left[\frac{P(E^*\mid H_1,I)}{P(E^*\mid H_0,I)}\right] + 10\log_{10}\left[\frac{P(H_1,I)}{P(H_0,I)}\right].$$ In the vitamin D example, a large negative study that concluded 'no association' yields WoE near $-0.75$ when power is corrected for proxy measurement and outcome misclassification, moving the probability of no association from 0.5 to only 0.54—essentially zero evidence for the null.

Load-bearing premise

The worked examples assume that all a study's evidence can be compressed into a single binary outcome—significant or not—with false-positive probability $\alpha$ and power, so the numerical WoE values depend on that simplification rather than on the actual continuous data.

Editorial extensions

If this is right

  • If the argument is correct, 'statistically significant' results do not by themselves support a hypothesis; support also depends on false-positive risk, power, and prior plausibility.
  • A study can have power and false-positive values such that it can never provide evidence for the alternative hypothesis: if the false-positive probability exceeds the power, the WoE cannot favor $H_1$.
  • For designing informative experiments, lowering the false-positive rate improves the WoE more than increasing power: $\alpha=0.01$ gives WoE $\approx 19$, while 95% power at $\alpha=0.05$ gives WoE $\approx 12.8$.
  • For evidence about the absence of an effect, minimizing false negatives matters more than the nominal alpha.
  • The vitamin D/COVID-19 study, despite its very large sample, provides essentially no evidence against an association once low power from proxy vitamin D measurements and outcome misclassification are accounted for.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct implication the author leaves implicit is that the same logic applies to Bayes factors: a large Bayes factor computed from a parameter does not by itself give the probability that the hypothesis is true; it must also be adjusted for study design, bias, and prior plausibility.
  • The WoE framework suggests a practical audit tool for published 'negative' studies: report power and a bias-adjusted false-positive rate, and treat studies with WoE within about $\pm3$ decibels as uninformative for meta-analyses.
  • A testable extension would be to calibrate WoE-derived probabilities against replication outcomes in large multi-study data sets: if hypotheses with WoE greater than 12 replicate at rates far below 0.95, the chosen power and alpha adjustments would need rethinking.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The paper argues that in most biological and social-science settings, statistical inference concerns parameters, whereas scientific claims concern hypotheses; therefore p-values, Bayes factors, and posterior probabilities of parameters are not probabilities of hypotheses. It introduces the label 'ultimate issue error' and proposes a Weight of Evidence (WoE) framework based on the likelihood ratio of binary 'positive/negative test' events, extended with a prior term, to combine numerical results with study-design information. Equations (1)-(5) present the algebra. The paper illustrates the idea with a hypothetical antidepressant trial and applies it to a UK Biobank vitamin D / COVID-19 study, concluding that the study provides 'zero evidence' against an association and should be excluded from meta-analyses.

Significance. The conceptual distinction between parameter probabilities and hypothesis probabilities is correct and pedagogically valuable, and the WoE framework is a transparent application of Bayes' theorem. The paper's separation of parametric inference from hypothesis evaluation, its use of the diagnostic-test analogy, and its explicit acknowledgment that subjectivity and potential for misuse are inherent to the approach are strengths. However, the applied vitamin D demonstration is not yet convincing: the quantitative conclusions depend on a binary simplification of the data and on hand-selected values for power and false-positive probability. If revised to use the graded likelihood and a sensitivity analysis, the example could support the central claim; as written, the numeric illustration overreaches.

major comments (4)
  1. [Real world example (Eq. 5); 'For simplicity' paragraph] The vitamin D WoE calculation collapses the study's evidence into a binary event E = 'p > 0.05'. Equation (5) is then applied with P(E|H1)=1-power and P(E|H0)=1-alpha. This discards the graded likelihood: Hastie et al. report an adjusted OR=1.00 with 95% CI 0.998 to 1.01, an estimate that is precise and very close to the null. Evaluated at the observed estimate, a likelihood-based measure would give substantially more evidence against a clinically meaningful OR=0.80 than 10 log10(0.8/0.9) = -0.75 dB. Thus the computed WoE and the 'zero evidence' conclusion are consequences of the binary simplification, not of the actual data alone.
  2. [Real world example: power calculation] The numerator P(E|H1) is evaluated at a single point alternative OR=0.80, but H1 = 'an association exists' is a composite hypothesis. P(negative result | H1) should be an average over the prior distribution of possible effect sizes under H1, not the power at one chosen OR. A true OR of 0.95 would produce a lower power and hence a larger numerator, while a true OR of 0.70 would produce the opposite. Without specifying a prior over effect sizes under H1, the quantity P(E|H1) is not well-defined and the resulting WoE is not identifiable.
  3. [Real world example: 'let's estimate the power to be 20%'] The reduction of power from 65% to 20% is asserted rather than derived, as the text states 'Simulations could be conducted if an accurate power estimate is important, but let's estimate the power to be 20%.' No sensitivity analysis is reported for the chosen alpha (0.1) and power (0.2), even though the WoE changes from -4.1 to -0.75 across the two power values considered. The conclusion that the evidence is approximately zero is therefore highly sensitive to unvalidated input values; the paper should provide a sensitivity table over plausible ranges of alpha, power, and prior odds, or justify the 20% value with a simulation.
  4. [Discussion: 'zero evidence' and 'can therefore be ignored'] The conclusion that the study provides 'zero evidence' and 'can therefore be ignored, and certainly not included in meta-analyses' is not entailed by the calculation. A WoE of -0.75 dB is negative but small; it is not zero evidence. Moreover, the claim that the study can be ignored in meta-analyses does not follow from the binary WoE, because that simplification discards the precision of the adjusted estimate. A precise null estimate like OR=1.00 with CI 0.998-1.01 is informative for meta-analysis and should not be excluded on the basis of the simplified WoE value alone.
minor comments (3)
  1. [Designing informative experiments] The sentence 'whereas decreasing false positives to alpha=0.01 increases the WoE to 10 log10(0.95/0.05) ≈ 19' appears to be a copy-paste error: the second expression should presumably read 10 log10(0.8/0.01) ≈ 19, since the power is unchanged at 0.8.
  2. [Equation (3)] The notation P(H1,I)/P(H0,I) is nonstandard and could be misread; the intended quantity is the prior odds P(H1|I)/P(H0|I) conditional on background information I.
  3. [P(Parameter) ≠ P(Hypothesis)] The informal notation P(Parameter) and P(Hypothesis) is used without defining the events underlying either probability; defining these events explicitly would prevent confusion with the later likelihood notation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the WoE framework is a direct Bayes-theorem calculation with stated inputs; the vitamin D example's conclusion follows from explicitly chosen power and false-positive values, not from circular reasoning.

full rationale

The paper's central derivation is the WoE identity in Equations 1-3, which is a direct application of Bayes' theorem: WoE = 10 log10(P(E|H1)/P(E|H0)) + prior odds. No step defines H1 or H0 in terms of the WoE output, and the prior term is set to zero in the applications by explicitly assuming equal prior probabilities. The vitamin D calculation uses Equation 5 with clearly stated inputs: specificity 1 - alpha = 0.9 and sensitivity power initially set to 65% and later hand-estimated at 20%, giving WoE values of -4.1 and -0.75. The conclusion that the study provides 'zero evidence' is a verbal interpretation of the small computed magnitude, not a separate empirical claim derived from that conclusion. There is no fitted parameter relabeled as a prediction, no load-bearing self-citation, and no uniqueness theorem imported from prior work by the author. The main limitation is that the applied result is sensitive to the hand-set power and false-positive values, so the example's conclusion reflects those assumptions more than the raw data; however, those values are stated assumptions, not outputs of the same calculation. The formal WoE framework is self-contained and non-circular, and the applied arithmetic is transparently conditional on its inputs.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

No free parameters are fitted to data in the traditional sense; instead, every WoE calculation imports user-chosen probabilities for false positives, power, and priors. The vitamin D example's conclusion is especially sensitive to the asserted power reduction to 20%, which is not derived from data. The paper introduces no new entities, forces, or conserved quantities.

free parameters (5)
  • False-positive probability (alpha) for vitamin D example = 0.05 nominal, adjusted to 0.10
    Used as P(E|H0) in Equations 4 and 5; the author raises it from 0.05 to 0.1 because the positive and negative groups differed on baseline variables and multiple analyses were run.
  • Initial power for vitamin D study = 65%
    Calculated from an assumed odds ratio of 0.80, 449 cases, and N=348,598; the R code is mentioned but not present, and the effect size is chosen by the author.
  • Adjusted power for vitamin D study = 20%
    Asserted after discussing 10-14 year measurement lag and likely misclassification of untested people as negative; no simulation, formula, or external estimate is provided.
  • Prior odds of association vs no association = 1 (log prior ratio = 0)
    Set to equal priors so the prior term in Equation 3 drops out; this is a stated convention for drawing conclusions independent of previous experiments.
  • Clinically meaningful effect size for power calculation = OR = 0.80
    Chosen by the author to define the alternative hypothesis in the power calculation; a smaller or larger assumed effect would change the power and WoE.
assumptions (5)
  • domain assumption Scientific evidence can be dichotomized into a positive or negative test result based on p<0.05.
    Equations 4 and 5 treat E as a binary significant or non-significant result; the vitamin D example codes negative result as p>0.05, discarding graded evidence.
  • domain assumption P(E|H0) equals the false-positive rate (adjustable by judgment) and P(E|H1) equals statistical power.
    Invoked in the section How can we quantify these terms; assumes the likelihood ratio for hypotheses is captured by alpha and power alone.
  • domain assumption Background information can be expressed as subjective probabilities P(E*|H1,I) and P(E*|H0,I), plus a prior odds ratio.
    Equation 3 requires these probabilities; the paper acknowledges subjectivity but gives no elicitation or calibration method.
  • domain assumption An experiment has operating characteristics like a diagnostic test, with sensitivity equal to power and false positive rate equal to alpha.
    The analogy in Equation 4 and surrounding text; assumes stable error rates that are independent of the observed data.
  • domain assumption Equal prior odds are appropriate when drawing conclusions independent of previous experiments.
    Used in both worked examples to set the prior term to zero; this is a convention, not a theorem.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The ultimate issue error in scientific inference: mistaking parameters for hypotheses." pith.science (2026). https://pith.science/paper/5J7WM665

@misc{pith2026241115398,
  author       = {Pith},
  title        = {Pith review of: The ultimate issue error in scientific inference: mistaking parameters for hypotheses},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5J7WM665}},
  note         = {Machine review of arXiv:2411.15398}
}
read the original abstract

Statistical inference often conflates the probability of a parameter with the probability of a hypothesis, a critical misunderstanding termed the ultimate issue error. This error is pervasive across the social, biological, and medical sciences, where null hypothesis significance testing (NHST) is mistakenly understood to be testing hypotheses rather than evaluating parameter estimates. Here, we advocate for using the Weight of Evidence (WoE) approach, which integrates quantitative data with qualitative background information for more accurate and transparent inference. Through a detailed example involving the relationship between vitamin D (25-hydroxy vitamin D) levels and COVID-19 risk, we demonstrate how WoE quantifies support for hypotheses while accounting for study design biases, power, and confounding factors. These findings emphasise the necessity of combining statistical metrics with contextual evaluation. This offers a structured framework to enhance reproducibility, reduce false interpretations, and foster robust scientific conclusions across disciplines.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

21 extracted references · 18 canonical work pages

  1. [1]

    Aitken, Colin, Alex Biedermann, Silvia Bozza, and Franco Taroni. 2021. Statistics and the Evaluation of Evidence for Forensic Scientists. 3rd ed. Wiley & Sons, Limited, John

  2. [2]

    Allen, Ronald J, and Michael S Pardo. 2019. ``Relative Plausibility and Its Critics.'' The International Journal of Evidence & Proof 23 (1--2): 5--59. https://doi.org/10.1177/1365712718813781

  3. [3]

    Button, Katherine S, John P A Ioannidis, Claire Mokrysz, Brian A Nosek, Jonathan Flint, Emma S J Robinson, and Marcus R Munafo. 2013. ``Power Failure: Why Small Sample Size Undermines the Reliability of Neuroscience.'' Nat Rev Neurosci 14 (5): 365--76. https://doi.org/10.1038/nrn3475

  4. [4]

    Colman, Andrew M., ed. 2009. A Dictionary of Psychology. 3. ed. Oxford Reference Online. Oxford: Oxford Univ. Press

  5. [5]

    Mazess, and Linda L

    Davies, Gareth, Richard B. Mazess, and Linda L. Benskin. 2021. ``Letter to the Editor in Response to the Article: `Vitamin D Concentrations and COVID-19 Infection in UK Biobank' (Hastie Et Al.).'' Diabetes & Metabolic Syndrome: Clinical Research & Reviews 15 (2): 643--44. https://doi.org/10.1016/j.dsx.2021.02.016

  6. [6]

    Earman, John, and Clark Glymour. 1980. ``Relativity and Eclipses: The British Eclipse Expeditions of 1919 and Their Predecessors.'' Historical Studies in the Physical Sciences 11 (1): 49--85. https://doi.org/10.2307/27757471

  7. [7]

    Edwards, A W F. 1992. Likelihood. 2nd ed. Baltimore, MD: Johns Hopkins University Press

  8. [8]

    Fairfield, Tasha, and Andrew E. Charman. 2022. Social Inquiry and Bayesian Inference: Rethinking Qualitative Research. University of Cambridge Press

Show all 21 references
  1. [9]

    Good, I J. 1950. Probability and the Weighing of Evidence. London: Charles Griffin & Company

  2. [10]

    Mackay, Frederick Ho, Carlos A

    Hastie, Claire E., Daniel F. Mackay, Frederick Ho, Carlos A. Celis-Morales, Srinivasa Vittal Katikireddi, Claire L. Niedzwiedz, Bhautesh D. Jani, et al. 2020. ``Vitamin D Concentrations and COVID-19 Infection in UK Biobank.'' Diabetes & Metabolic Syndrome: Clinical Research &;...

  3. [11]

    Higgins, J. P. T., D. G. Altman, P. C. Gotzsche, P. Juni, D. Moher, A. D. Oxman, J. Savovic, K. F. Schulz, L. Weeks, and J. A. C. Sterne. 2011. `` The Cochrane Collaboration's tool for assessing risk of bias in randomised trials .'' BMJ 343 (oct18 2): d5928--28. https://doi.or...

  4. [12]

    M., and D

    Hoenig, J. M., and D. M. Heisey. 2001. ``The Abuse of Power: The Pervasive Fallacy of Power Calculations for Data Analysis.'' The American Statistician 55 (1): 19--24

  5. [13]

    Jaynes, E. T. 2003. Probability Theory: The Logic of Science. Cambridge, UK: Cambridge University Press

  6. [14]

    Sneve, M

    Jorde, R., M. Sneve, M. Hutchinson, N. Emaus, Y. Figenschau, and G. Grimnes. 2010. ``Tracking of Serum 25-Hydroxyvitamin D Levels During 14 Years in a Population-Based Study and During 12 Months in an Intervention Study.'' American Journal of Epidemiology 171 (8): 903--8. http...

  7. [15]

    Juslin, Peter, Hakan Nilsson, Anders Winman, and Marcus Lindskog. 2011. ``Reducing Cognitive Biases in Probabilistic Reasoning by the Use of Logarithm Formats.'' Cognition 120 (2): 248--67. https://doi.org/10.1016/j.cognition.2011.05.004

  8. [16]

    Levine, M., and M. H. Ensom. 2001. ``Post Hoc Power Analysis: An Idea Whose Time Has Passed? https://www.ncbi.nlm.nih.gov/pubmed/11310512'' Pharmacotherapy 21 (4): 405--9

  9. [17]

    Hovey, Jean Wactawski-Wende, Christopher A

    Meng, Jennifer E., Kathleen M. Hovey, Jean Wactawski-Wende, Christopher A. Andrews, Michael J. LaMonte, Ronald L. Horst, Robert J. Genco, and Amy E. Millen. 2012. ``Intraindividual Variation in Plasma 25-Hydroxyvitamin D Measures 5 Years Apart Among Postmenopausal Women.'' Can...

  10. [18]

    Pardo, Michael S., and Ronald J. Allen. 2007. ``Juridical Proof and the Best Explanation.'' Law and Philosophy 27 (3): 223--68. https://doi.org/10.1007/s10982-007-9016-4

  11. [19]

    Peirce, Charles Sanders. 2014. Illustrations of the Logic of Science. Edited by Cornelis De Waal. New York: Open Court

  12. [20]

    Polya, George. 1954. Mathematics and Plausible Reasoning. Vol. I and II. Mansfield Centre, CT: Martino Fine Books

  13. [21]

    Senn, Stephen J. 2002. ``Power Is Indeed Irrelevant in Interpreting Completed Studies. https://www.ncbi.nlm.nih.gov/pubmed/12458264'' BMJ 325 (7375): 1304. CSLReferences document

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.