REVIEW 4 major objections 3 minor 21 references
The ultimate issue error in scientific inference: mistaking parameters for hypotheses
T0 review · 4 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The probability of a parameter is not the probability of a hypothesis, so p<0.05 alone cannot support a claim like 'the treatment works'.
desk verdict A clear, honest restatement of the parameter-vs-hypothesis distinction, but the vitamin D 'zero evidence' conclusion is an artifact of the author's simplifying choices, not a consequence of the WoE algebra. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the augmented likelihood ratio, expressed as weight of evidence in decibel units: $$\text{WoE}(H_1:E^*,I) = 10\log_{10}\left[\frac{P(E^*\mid H_1,I)}{P(E^*\mid H_0,I)}\right] + 10\log_{10}\left[\frac{P(H_1,I)}{P(H_0,I)}\right].$$ It carries the argument by converting a study result into support for a hypothesis, with $P(E^*\mid H_0,I)$ the false-positive probability adjusted for bias and $P(E^*\mid H_1,I)$ the power adjusted for design flaws. The logarithmic scale makes independent pieces of evidence additive, and the prior term lets background plausibility enter explicitly. In the worked examples, a significant result with 60% power and a 15% false-positive rate gives WoE $\approx 6$ (0.8 probability for the hypothesis), while a negative result gives WoE $\approx -3$ (0.67 probability for the null), showing that both outcomes are nearly uninformative.
What would settle it
Re-run the vitamin D analysis using a likelihood ratio built from the continuous test statistic or the full logistic-regression likelihood instead of coding the result as $p>0.05$. If the resulting weight of evidence is not close to the paper's $-0.75$, then the binary $\alpha$/power collapse is what produces the 'zero evidence' conclusion.
Extended reading notes
Core claim
The paper's central claim is that the probability of a parameter and the probability of a hypothesis are distinct quantities, and conflating them—the 'ultimate issue error'—makes statements such as 'the treatment is effective (p<0.05)' non sequiturs. A parameter is a quantitative feature of a statistical model; a hypothesis is a testable proposition whose truth depends on background information, construct validity, competing explanations, and prior plausibility. The paper argues that no direct quantitative relationship links p-values or Bayes factors to hypothesis probabilities, and that moving from parameter evidence to hypothesis support requires the augmented likelihood ratio $$\text{WoE}(H_1:E^*,I) = 10\log_{10}\left[\frac{P(E^*\mid H_1,I)}{P(E^*\mid H_0,I)}\right] + 10\log_{10}\left[\frac{P(H_1,I)}{P(H_0,I)}\right].$$ In the vitamin D example, a large negative study that concluded 'no association' yields WoE near $-0.75$ when power is corrected for proxy measurement and outcome misclassification, moving the probability of no association from 0.5 to only 0.54—essentially zero evidence for the null.
Load-bearing premise
The worked examples assume that all a study's evidence can be compressed into a single binary outcome—significant or not—with false-positive probability $\alpha$ and power, so the numerical WoE values depend on that simplification rather than on the actual continuous data.
Editorial extensions
If this is right
- If the argument is correct, 'statistically significant' results do not by themselves support a hypothesis; support also depends on false-positive risk, power, and prior plausibility.
- A study can have power and false-positive values such that it can never provide evidence for the alternative hypothesis: if the false-positive probability exceeds the power, the WoE cannot favor $H_1$.
- For designing informative experiments, lowering the false-positive rate improves the WoE more than increasing power: $\alpha=0.01$ gives WoE $\approx 19$, while 95% power at $\alpha=0.05$ gives WoE $\approx 12.8$.
- For evidence about the absence of an effect, minimizing false negatives matters more than the nominal alpha.
- The vitamin D/COVID-19 study, despite its very large sample, provides essentially no evidence against an association once low power from proxy vitamin D measurements and outcome misclassification are accounted for.
Reading between the lines
- A direct implication the author leaves implicit is that the same logic applies to Bayes factors: a large Bayes factor computed from a parameter does not by itself give the probability that the hypothesis is true; it must also be adjusted for study design, bias, and prior plausibility.
- The WoE framework suggests a practical audit tool for published 'negative' studies: report power and a bias-adjusted false-positive rate, and treat studies with WoE within about $\pm3$ decibels as uninformative for meta-analyses.
- A testable extension would be to calibrate WoE-derived probabilities against replication outcomes in large multi-study data sets: if hypotheses with WoE greater than 12 replicate at rates far below 0.95, the chosen power and alpha adjustments would need rethinking.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that in most biological and social-science settings, statistical inference concerns parameters, whereas scientific claims concern hypotheses; therefore p-values, Bayes factors, and posterior probabilities of parameters are not probabilities of hypotheses. It introduces the label 'ultimate issue error' and proposes a Weight of Evidence (WoE) framework based on the likelihood ratio of binary 'positive/negative test' events, extended with a prior term, to combine numerical results with study-design information. Equations (1)-(5) present the algebra. The paper illustrates the idea with a hypothetical antidepressant trial and applies it to a UK Biobank vitamin D / COVID-19 study, concluding that the study provides 'zero evidence' against an association and should be excluded from meta-analyses.
Significance. The conceptual distinction between parameter probabilities and hypothesis probabilities is correct and pedagogically valuable, and the WoE framework is a transparent application of Bayes' theorem. The paper's separation of parametric inference from hypothesis evaluation, its use of the diagnostic-test analogy, and its explicit acknowledgment that subjectivity and potential for misuse are inherent to the approach are strengths. However, the applied vitamin D demonstration is not yet convincing: the quantitative conclusions depend on a binary simplification of the data and on hand-selected values for power and false-positive probability. If revised to use the graded likelihood and a sensitivity analysis, the example could support the central claim; as written, the numeric illustration overreaches.
major comments (4)
- [Real world example (Eq. 5); 'For simplicity' paragraph] The vitamin D WoE calculation collapses the study's evidence into a binary event E = 'p > 0.05'. Equation (5) is then applied with P(E|H1)=1-power and P(E|H0)=1-alpha. This discards the graded likelihood: Hastie et al. report an adjusted OR=1.00 with 95% CI 0.998 to 1.01, an estimate that is precise and very close to the null. Evaluated at the observed estimate, a likelihood-based measure would give substantially more evidence against a clinically meaningful OR=0.80 than 10 log10(0.8/0.9) = -0.75 dB. Thus the computed WoE and the 'zero evidence' conclusion are consequences of the binary simplification, not of the actual data alone.
- [Real world example: power calculation] The numerator P(E|H1) is evaluated at a single point alternative OR=0.80, but H1 = 'an association exists' is a composite hypothesis. P(negative result | H1) should be an average over the prior distribution of possible effect sizes under H1, not the power at one chosen OR. A true OR of 0.95 would produce a lower power and hence a larger numerator, while a true OR of 0.70 would produce the opposite. Without specifying a prior over effect sizes under H1, the quantity P(E|H1) is not well-defined and the resulting WoE is not identifiable.
- [Real world example: 'let's estimate the power to be 20%'] The reduction of power from 65% to 20% is asserted rather than derived, as the text states 'Simulations could be conducted if an accurate power estimate is important, but let's estimate the power to be 20%.' No sensitivity analysis is reported for the chosen alpha (0.1) and power (0.2), even though the WoE changes from -4.1 to -0.75 across the two power values considered. The conclusion that the evidence is approximately zero is therefore highly sensitive to unvalidated input values; the paper should provide a sensitivity table over plausible ranges of alpha, power, and prior odds, or justify the 20% value with a simulation.
- [Discussion: 'zero evidence' and 'can therefore be ignored'] The conclusion that the study provides 'zero evidence' and 'can therefore be ignored, and certainly not included in meta-analyses' is not entailed by the calculation. A WoE of -0.75 dB is negative but small; it is not zero evidence. Moreover, the claim that the study can be ignored in meta-analyses does not follow from the binary WoE, because that simplification discards the precision of the adjusted estimate. A precise null estimate like OR=1.00 with CI 0.998-1.01 is informative for meta-analysis and should not be excluded on the basis of the simplified WoE value alone.
minor comments (3)
- [Designing informative experiments] The sentence 'whereas decreasing false positives to alpha=0.01 increases the WoE to 10 log10(0.95/0.05) ≈ 19' appears to be a copy-paste error: the second expression should presumably read 10 log10(0.8/0.01) ≈ 19, since the power is unchanged at 0.8.
- [Equation (3)] The notation P(H1,I)/P(H0,I) is nonstandard and could be misread; the intended quantity is the prior odds P(H1|I)/P(H0|I) conditional on background information I.
- [P(Parameter) ≠ P(Hypothesis)] The informal notation P(Parameter) and P(Hypothesis) is used without defining the events underlying either probability; defining these events explicitly would prevent confusion with the later likelihood notation.
Circularity Check
No significant circularity: the WoE framework is a direct Bayes-theorem calculation with stated inputs; the vitamin D example's conclusion follows from explicitly chosen power and false-positive values, not from circular reasoning.
full rationale
The paper's central derivation is the WoE identity in Equations 1-3, which is a direct application of Bayes' theorem: WoE = 10 log10(P(E|H1)/P(E|H0)) + prior odds. No step defines H1 or H0 in terms of the WoE output, and the prior term is set to zero in the applications by explicitly assuming equal prior probabilities. The vitamin D calculation uses Equation 5 with clearly stated inputs: specificity 1 - alpha = 0.9 and sensitivity power initially set to 65% and later hand-estimated at 20%, giving WoE values of -4.1 and -0.75. The conclusion that the study provides 'zero evidence' is a verbal interpretation of the small computed magnitude, not a separate empirical claim derived from that conclusion. There is no fitted parameter relabeled as a prediction, no load-bearing self-citation, and no uniqueness theorem imported from prior work by the author. The main limitation is that the applied result is sensitive to the hand-set power and false-positive values, so the example's conclusion reflects those assumptions more than the raw data; however, those values are stated assumptions, not outputs of the same calculation. The formal WoE framework is self-contained and non-circular, and the applied arithmetic is transparently conditional on its inputs.
Assumptions & free parameters
free parameters (5)
- False-positive probability (alpha) for vitamin D example =
0.05 nominal, adjusted to 0.10
- Initial power for vitamin D study =
65%
- Adjusted power for vitamin D study =
20%
- Prior odds of association vs no association =
1 (log prior ratio = 0)
- Clinically meaningful effect size for power calculation =
OR = 0.80
assumptions (5)
- domain assumption Scientific evidence can be dichotomized into a positive or negative test result based on p<0.05.
- domain assumption P(E|H0) equals the false-positive rate (adjustable by judgment) and P(E|H1) equals statistical power.
- domain assumption Background information can be expressed as subjective probabilities P(E*|H1,I) and P(E*|H0,I), plus a prior odds ratio.
- domain assumption An experiment has operating characteristics like a diagnostic test, with sensitivity equal to power and false positive rate equal to alpha.
- domain assumption Equal prior odds are appropriate when drawing conclusions independent of previous experiments.
Cite this review
Pith. "Pith review of The ultimate issue error in scientific inference: mistaking parameters for hypotheses." pith.science (2026). https://pith.science/paper/5J7WM665
@misc{pith2026241115398,
author = {Pith},
title = {Pith review of: The ultimate issue error in scientific inference: mistaking parameters for hypotheses},
year = {2026},
howpublished = {\url{https://pith.science/paper/5J7WM665}},
note = {Machine review of arXiv:2411.15398}
}
read the original abstract
Statistical inference often conflates the probability of a parameter with the probability of a hypothesis, a critical misunderstanding termed the ultimate issue error. This error is pervasive across the social, biological, and medical sciences, where null hypothesis significance testing (NHST) is mistakenly understood to be testing hypotheses rather than evaluating parameter estimates. Here, we advocate for using the Weight of Evidence (WoE) approach, which integrates quantitative data with qualitative background information for more accurate and transparent inference. Through a detailed example involving the relationship between vitamin D (25-hydroxy vitamin D) levels and COVID-19 risk, we demonstrate how WoE quantifies support for hypotheses while accounting for study design biases, power, and confounding factors. These findings emphasise the necessity of combining statistical metrics with contextual evaluation. This offers a structured framework to enhance reproducibility, reduce false interpretations, and foster robust scientific conclusions across disciplines.
Reference graph
Works this paper leans on
-
[1]
Aitken, Colin, Alex Biedermann, Silvia Bozza, and Franco Taroni. 2021. Statistics and the Evaluation of Evidence for Forensic Scientists. 3rd ed. Wiley & Sons, Limited, John
work page 2021
-
[2]
Allen, Ronald J, and Michael S Pardo. 2019. ``Relative Plausibility and Its Critics.'' The International Journal of Evidence & Proof 23 (1--2): 5--59. https://doi.org/10.1177/1365712718813781
-
[3]
Button, Katherine S, John P A Ioannidis, Claire Mokrysz, Brian A Nosek, Jonathan Flint, Emma S J Robinson, and Marcus R Munafo. 2013. ``Power Failure: Why Small Sample Size Undermines the Reliability of Neuroscience.'' Nat Rev Neurosci 14 (5): 365--76. https://doi.org/10.1038/nrn3475
doi:10.1038/nrn3475 2013
-
[4]
Colman, Andrew M., ed. 2009. A Dictionary of Psychology. 3. ed. Oxford Reference Online. Oxford: Oxford Univ. Press
work page 2009
-
[5]
Davies, Gareth, Richard B. Mazess, and Linda L. Benskin. 2021. ``Letter to the Editor in Response to the Article: `Vitamin D Concentrations and COVID-19 Infection in UK Biobank' (Hastie Et Al.).'' Diabetes & Metabolic Syndrome: Clinical Research & Reviews 15 (2): 643--44. https://doi.org/10.1016/j.dsx.2021.02.016
-
[6]
Earman, John, and Clark Glymour. 1980. ``Relativity and Eclipses: The British Eclipse Expeditions of 1919 and Their Predecessors.'' Historical Studies in the Physical Sciences 11 (1): 49--85. https://doi.org/10.2307/27757471
-
[7]
Edwards, A W F. 1992. Likelihood. 2nd ed. Baltimore, MD: Johns Hopkins University Press
work page 1992
-
[8]
Fairfield, Tasha, and Andrew E. Charman. 2022. Social Inquiry and Bayesian Inference: Rethinking Qualitative Research. University of Cambridge Press
work page 2022
Show all 21 references
-
[9]
Good, I J. 1950. Probability and the Weighing of Evidence. London: Charles Griffin & Company
1950
-
[10]
Mackay, Frederick Ho, Carlos A
Hastie, Claire E., Daniel F. Mackay, Frederick Ho, Carlos A. Celis-Morales, Srinivasa Vittal Katikireddi, Claire L. Niedzwiedz, Bhautesh D. Jani, et al. 2020. ``Vitamin D Concentrations and COVID-19 Infection in UK Biobank.'' Diabetes & Metabolic Syndrome: Clinical Research &;...
2020 doi
-
[11]
Higgins, J. P. T., D. G. Altman, P. C. Gotzsche, P. Juni, D. Moher, A. D. Oxman, J. Savovic, K. F. Schulz, L. Weeks, and J. A. C. Sterne. 2011. `` The Cochrane Collaboration's tool for assessing risk of bias in randomised trials .'' BMJ 343 (oct18 2): d5928--28. https://doi.or...
2011 doi
-
[12]
M., and D
Hoenig, J. M., and D. M. Heisey. 2001. ``The Abuse of Power: The Pervasive Fallacy of Power Calculations for Data Analysis.'' The American Statistician 55 (1): 19--24
2001
-
[13]
Jaynes, E. T. 2003. Probability Theory: The Logic of Science. Cambridge, UK: Cambridge University Press
2003
-
[14]
Sneve, M
Jorde, R., M. Sneve, M. Hutchinson, N. Emaus, Y. Figenschau, and G. Grimnes. 2010. ``Tracking of Serum 25-Hydroxyvitamin D Levels During 14 Years in a Population-Based Study and During 12 Months in an Intervention Study.'' American Journal of Epidemiology 171 (8): 903--8. http...
2010 doi
-
[15]
Juslin, Peter, Hakan Nilsson, Anders Winman, and Marcus Lindskog. 2011. ``Reducing Cognitive Biases in Probabilistic Reasoning by the Use of Logarithm Formats.'' Cognition 120 (2): 248--67. https://doi.org/10.1016/j.cognition.2011.05.004
2011 doi
-
[16]
Levine, M., and M. H. Ensom. 2001. ``Post Hoc Power Analysis: An Idea Whose Time Has Passed? https://www.ncbi.nlm.nih.gov/pubmed/11310512'' Pharmacotherapy 21 (4): 405--9
2001
-
[17]
Hovey, Jean Wactawski-Wende, Christopher A
Meng, Jennifer E., Kathleen M. Hovey, Jean Wactawski-Wende, Christopher A. Andrews, Michael J. LaMonte, Ronald L. Horst, Robert J. Genco, and Amy E. Millen. 2012. ``Intraindividual Variation in Plasma 25-Hydroxyvitamin D Measures 5 Years Apart Among Postmenopausal Women.'' Can...
2012 doi
-
[18]
Pardo, Michael S., and Ronald J. Allen. 2007. ``Juridical Proof and the Best Explanation.'' Law and Philosophy 27 (3): 223--68. https://doi.org/10.1007/s10982-007-9016-4
2007 doi
-
[19]
Peirce, Charles Sanders. 2014. Illustrations of the Logic of Science. Edited by Cornelis De Waal. New York: Open Court
2014
-
[20]
Polya, George. 1954. Mathematics and Plausible Reasoning. Vol. I and II. Mansfield Centre, CT: Martino Fine Books
1954
-
[21]
Senn, Stephen J. 2002. ``Power Is Indeed Irrelevant in Interpreting Completed Studies. https://www.ncbi.nlm.nih.gov/pubmed/12458264'' BMJ 325 (7375): 1304. CSLReferences document
2002
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.