{"id":"8a340ce5-ade3-4c02-a726-0acdd44aa0ea","arxiv_id":"2411.15398","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Parameter probabilities are not hypothesis probabilities, so p-values alone do not measure support for a scientific hypothesis; the paper proposes a weight-of-evidence calculation that adds study quality and priors.","lead":"This paper argues that scientists routinely mistake the probability of a statistical parameter for the probability that a scientific hypothesis is true, a confusion it calls the ultimate issue error. It offers a weight-of-evidence framework that combines likelihood ratios with qualitative judgments about study quality, then re-examines a widely cited vitamin D and COVID study.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The WoE demonstration compresses the vitamin D data into a binary significant/not event, so the 'zero evidence' conclusion is an artifact of that simplification rather than a consequence of the actual likelihood.","rationale":"The central thesis of the paper - that a p-value, Bayes factor, or posterior probability about a parameter is not automatically the probability of a scientific hypothesis - is conceptually sound and does not depend on the vitamin D numbers. The reader's conditional verdict is therefore appropriate. The concern I identify is the same one the reader flagged: the quantitative WoE examples rely on coding the evidence as significant or not, with P(E|H1)=power and P(E|H0)=alpha, and this compression is doing real work in the vitamin D application. In particular, the reported 95% CI around the adjusted OR is extremely narrow; treating this as merely 'a negative result' discards the information that the estimate is precise and near the null. A proper likelihood or Bayes-factor computation would almost certainly produce a WoE very different from -0.75 dB, and the paper's claim of 'zero evidence' overstates what its own calculation shows. Since the conceptual argument survives and the flaw is in the applied demonstration, the reader's CONDITIONAL verdict should stand unchanged. The proposed recomputation with a continuous likelihood would settle whether the applied conclusion is an artifact of binary collapsing; if it is, the paper needs to qualify the vitamin D and meta-analysis claims, not abandon the core thesis. I found no reason to accuse the authors of bad faith; the issue is a technical one about how evidence is summarized.","tokens_in":11233,"tokens_out":6404,"duration_ms":61030,"concrete_test":"Recompute the vitamin D WoE from the actual adjusted estimate rather than the binary event p>0.05. Approximate log(OR) as Normal with mean log(1.00)≈0 and SE=(log(1.01)-log(0.998))/(2*1.96)≈0.00305. Compare this likelihood under H0: logOR=0 with a prior-weighted composite H1, for example logOR ~ Normal(log(0.80), tau) for tau in {0.2, 0.5, 1}; compute 10 log10 of the resulting Bayes factor or likelihood ratio. If the continuous-data WoE differs materially from the paper's -0.75 dB and especially if it strongly favors H0, the binary compression is the load-bearing assumption behind 'zero evidence.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's conceptual distinction between parameter probabilities and hypothesis probabilities is sound, and the WoE algebra is standard. The load-bearing weak spot is the operational bridge from that thesis to the vitamin D application. Equation (4) and Equation (5), together with the statement 'For simplicity we can take p<0.05 as evidence for an effect,' collapse a study's evidence into a binary event, and the vitamin D example then sets P(E|H1) to a single power value (65%, later 20%) computed at a point alternative OR=0.80. Two problems follow. First, H1 = 'an association exists' is a composite hypothesis; P(negative result | H1) is an average over the prior distribution of true effect sizes, not the power at one chosen OR. Second, even if the alternative were fixed, binary collapsing throws away the precision of the actual data. Hastie et al. report an adjusted OR=1.00 with 95% CI 0.998 to 1.01, an extremely precise null estimate. A likelihood-based WoE evaluated at the observed estimate would give far more evidence against a clinically meaningful association than 10 log10(0.8/0.9) = -0.75 dB. The paper's own computed value is -0.75 dB of evidence for H0, which is not 'zero evidence'; the applied conclusion that the study should be excluded from meta-analyses is therefore not entailed by the WoE equations but by the binary simplification and the hand-set 20% power.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that in most biological and social-science settings, statistical inference concerns parameters, whereas scientific claims concern hypotheses; therefore p-values, Bayes factors, and posterior probabilities of parameters are not probabilities of hypotheses. It introduces the label 'ultimate issue error' and proposes a Weight of Evidence (WoE) framework based on the likelihood ratio of binary 'positive/negative test' events, extended with a prior term, to combine numerical results with study-design information. Equations (1)-(5) present the algebra. The paper illustrates the idea with a hypothetical antidepressant trial and applies it to a UK Biobank vitamin D / COVID-19 study, concluding that the study provides 'zero evidence' against an association and should be excluded from meta-analyses.","tokens_in":11541,"tokens_out":6548,"duration_ms":56731,"significance":"The conceptual distinction between parameter probabilities and hypothesis probabilities is correct and pedagogically valuable, and the WoE framework is a transparent application of Bayes' theorem. The paper's separation of parametric inference from hypothesis evaluation, its use of the diagnostic-test analogy, and its explicit acknowledgment that subjectivity and potential for misuse are inherent to the approach are strengths. However, the applied vitamin D demonstration is not yet convincing: the quantitative conclusions depend on a binary simplification of the data and on hand-selected values for power and false-positive probability. If revised to use the graded likelihood and a sensitivity analysis, the example could support the central claim; as written, the numeric illustration overreaches.","major_comments":[{"comment":"The vitamin D WoE calculation collapses the study's evidence into a binary event E = 'p > 0.05'. Equation (5) is then applied with P(E|H1)=1-power and P(E|H0)=1-alpha. This discards the graded likelihood: Hastie et al. report an adjusted OR=1.00 with 95% CI 0.998 to 1.01, an estimate that is precise and very close to the null. Evaluated at the observed estimate, a likelihood-based measure would give substantially more evidence against a clinically meaningful OR=0.80 than 10 log10(0.8/0.9) = -0.75 dB. Thus the computed WoE and the 'zero evidence' conclusion are consequences of the binary simplification, not of the actual data alone.","section":"Real world example (Eq. 5); 'For simplicity' paragraph"},{"comment":"The numerator P(E|H1) is evaluated at a single point alternative OR=0.80, but H1 = 'an association exists' is a composite hypothesis. P(negative result | H1) should be an average over the prior distribution of possible effect sizes under H1, not the power at one chosen OR. A true OR of 0.95 would produce a lower power and hence a larger numerator, while a true OR of 0.70 would produce the opposite. Without specifying a prior over effect sizes under H1, the quantity P(E|H1) is not well-defined and the resulting WoE is not identifiable.","section":"Real world example: power calculation"},{"comment":"The reduction of power from 65% to 20% is asserted rather than derived, as the text states 'Simulations could be conducted if an accurate power estimate is important, but let's estimate the power to be 20%.' No sensitivity analysis is reported for the chosen alpha (0.1) and power (0.2), even though the WoE changes from -4.1 to -0.75 across the two power values considered. The conclusion that the evidence is approximately zero is therefore highly sensitive to unvalidated input values; the paper should provide a sensitivity table over plausible ranges of alpha, power, and prior odds, or justify the 20% value with a simulation.","section":"Real world example: 'let's estimate the power to be 20%'"},{"comment":"The conclusion that the study provides 'zero evidence' and 'can therefore be ignored, and certainly not included in meta-analyses' is not entailed by the calculation. A WoE of -0.75 dB is negative but small; it is not zero evidence. Moreover, the claim that the study can be ignored in meta-analyses does not follow from the binary WoE, because that simplification discards the precision of the adjusted estimate. A precise null estimate like OR=1.00 with CI 0.998-1.01 is informative for meta-analysis and should not be excluded on the basis of the simplified WoE value alone.","section":"Discussion: 'zero evidence' and 'can therefore be ignored'"}],"minor_comments":[{"comment":"The sentence 'whereas decreasing false positives to alpha=0.01 increases the WoE to 10 log10(0.95/0.05) ≈ 19' appears to be a copy-paste error: the second expression should presumably read 10 log10(0.8/0.01) ≈ 19, since the power is unchanged at 0.8.","section":"Designing informative experiments"},{"comment":"The notation P(H1,I)/P(H0,I) is nonstandard and could be misread; the intended quantity is the prior odds P(H1|I)/P(H0|I) conditional on background information I.","section":"Equation (3)"},{"comment":"The informal notation P(Parameter) and P(Hypothesis) is used without defining the events underlying either probability; defining these events explicitly would prevent confusion with the later likelihood notation.","section":"P(Parameter) ≠ P(Hypothesis)"}],"recommendation":"major_revision","confidential_remarks":"The paper reads more as a perspective or educational piece than as a new statistical method. The central thesis is sound and the WoE algebra is standard, but the vitamin D example invites scrutiny because its numeric inputs are largely arbitrary and the binary-event simplification drives the conclusions. The editor may wish to ask the author to either reframe the example as purely illustrative with a full sensitivity analysis, or replace it with a graded-likelihood demonstration. Providing the R code referenced in the supplementary material would also strengthen reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper before reading it. First, the conceptual core is sound and clearly argued: P(parameter|data) is not P(hypothesis|data), and the WoE framework is a legitimate way to formalize the background information that researchers already weigh implicitly. Second, the applied vitamin D example is much weaker than the prose around it. The 'zero evidence' conclusion is not entailed by the WoE equations; it comes from a binary simplification and a hand-set power value.\n\nWhat is actually new here is modest. The distinction is the prosecutor's fallacy, and the likelihood-ratio/WoE math is standard—Peirce, Good, Jaynes, forensic statistics. But the paper does a genuine service by translating that into a practical framework for ordinary scientific inference, with concrete examples showing how study design, bias, and power move the WoE. The point that a study with false-positive probability above power can never support the alternative is a useful intuition. The discussion of designing experiments (prefer lowering alpha over boosting power, given diminishing returns) is also sensible and well explained. The author is upfront about subjectivity and even flags the risk of the method being abused.\n\nNow the soft spots. The load-bearing problem is Equation (4)/(5) plus the sentence \"For simplicity we can take p<0.05 as evidence for an effect.\" This collapses a continuous, graded dataset into a binary significant/not event. For the vitamin D study, the observed adjusted OR is 1.00 with CI 0.998–1.01—extremely precise null information. A likelihood computed at the actual estimate would give much stronger evidence against a clinically meaningful association than the computed -0.75 dB. The author's numbers are also chosen generously: alpha bumped to 0.1, power initially 65% based on a point alternative OR=0.80, then asserted to drop to 20% without simulation or sensitivity analysis. The prior ratio is set to 1, which is a choice, not a neutral one. And -0.75 dB is not zero evidence; it's weak evidence for H0 (LR≈0.84). The claim that the study \"provided zero evidence\" and \"certainly not included in meta-analyses\" overreaches what the framework actually yields.\n\nThat said, the central argument does not depend on those numbers. The distinction stands, and the WoE approach is a reasonable structured alternative to NHST as hypothesis testing. The paper is honest, readable, and cites its sources properly.\n\nThis is a paper for a statistics-methods audience and for applied researchers who conflate p-values with hypothesis support. It deserves a serious referee: the conceptual message is worth publishing, and the applied example can be fixed with a proper sensitivity analysis. My advice is to send it to peer review with the expectation of major revision.\n\nFor my own work, I would not cite it as a methodological contribution, but it might be a useful discussion piece in a reading group.","headline":"A clear, honest restatement of the parameter-vs-hypothesis distinction, but the vitamin D 'zero evidence' conclusion is an artifact of the author's simplifying choices, not a consequence of the WoE algebra.","tokens_in":12073,"tokens_out":1682,"would_cite":false,"duration_ms":16892,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The probability of a parameter is not the probability of a hypothesis, so p<0.05 alone cannot support a claim like 'the treatment works'.","keywords":["weight of evidence","ultimate issue error","hypothesis testing","p-values","Bayes factors","study design bias","vitamin D COVID-19","likelihood ratio"],"falsifier":"Re-run the vitamin D analysis using a likelihood ratio built from the continuous test statistic or the full logistic-regression likelihood instead of coding the result as $p>0.05$. If the resulting weight of evidence is not close to the paper's $-0.75$, then the binary $\\alpha$/power collapse is what produces the 'zero evidence' conclusion.","tokens_in":11006,"feed_emoji":"⚖️","tokens_out":6889,"duration_ms":58939,"temperature":0.7,"pith_summary":"This paper argues that standard statistical tests answer a different question from the one scientists usually pose. A p-value, Bayes factor, or posterior probability describes a parameter in a model; it does not give the probability that a scientific hypothesis such as 'this drug is effective' is true. The paper proposes a weight-of-evidence (WoE) approach, a logarithmic likelihood ratio that folds in study design, bias, power, and prior plausibility, and applies it to a widely cited vitamin D and COVID-19 study. If the argument is right, statements like 'the treatment is effective (p<0.05)' are non sequiturs, and many published negative results carry little or no evidence for the null hypothesis.","feed_headline":"p<0.05 does not make the hypothesis true","feed_subtitle":"Weight-of-evidence folds in bias, power, and priors, showing many significant or null results are uninformative.","key_machinery":"The central object is the augmented likelihood ratio, expressed as weight of evidence in decibel units: $$\\text{WoE}(H_1:E^*,I) = 10\\log_{10}\\left[\\frac{P(E^*\\mid H_1,I)}{P(E^*\\mid H_0,I)}\\right] + 10\\log_{10}\\left[\\frac{P(H_1,I)}{P(H_0,I)}\\right].$$ It carries the argument by converting a study result into support for a hypothesis, with $P(E^*\\mid H_0,I)$ the false-positive probability adjusted for bias and $P(E^*\\mid H_1,I)$ the power adjusted for design flaws. The logarithmic scale makes independent pieces of evidence additive, and the prior term lets background plausibility enter explicitly. In the worked examples, a significant result with 60% power and a 15% false-positive rate gives WoE $\\approx 6$ (0.8 probability for the hypothesis), while a negative result gives WoE $\\approx -3$ (0.67 probability for the null), showing that both outcomes are nearly uninformative.","core_discovery":"The paper's central claim is that the probability of a parameter and the probability of a hypothesis are distinct quantities, and conflating them—the 'ultimate issue error'—makes statements such as 'the treatment is effective (p<0.05)' non sequiturs. A parameter is a quantitative feature of a statistical model; a hypothesis is a testable proposition whose truth depends on background information, construct validity, competing explanations, and prior plausibility. The paper argues that no direct quantitative relationship links p-values or Bayes factors to hypothesis probabilities, and that moving from parameter evidence to hypothesis support requires the augmented likelihood ratio $$\\text{WoE}(H_1:E^*,I) = 10\\log_{10}\\left[\\frac{P(E^*\\mid H_1,I)}{P(E^*\\mid H_0,I)}\\right] + 10\\log_{10}\\left[\\frac{P(H_1,I)}{P(H_0,I)}\\right].$$ In the vitamin D example, a large negative study that concluded 'no association' yields WoE near $-0.75$ when power is corrected for proxy measurement and outcome misclassification, moving the probability of no association from 0.5 to only 0.54—essentially zero evidence for the null.","pith_inferences":["A direct implication the author leaves implicit is that the same logic applies to Bayes factors: a large Bayes factor computed from a parameter does not by itself give the probability that the hypothesis is true; it must also be adjusted for study design, bias, and prior plausibility.","The WoE framework suggests a practical audit tool for published 'negative' studies: report power and a bias-adjusted false-positive rate, and treat studies with WoE within about $\\pm3$ decibels as uninformative for meta-analyses.","A testable extension would be to calibrate WoE-derived probabilities against replication outcomes in large multi-study data sets: if hypotheses with WoE greater than 12 replicate at rates far below 0.95, the chosen power and alpha adjustments would need rethinking."],"forward_implications":["If the argument is correct, 'statistically significant' results do not by themselves support a hypothesis; support also depends on false-positive risk, power, and prior plausibility.","A study can have power and false-positive values such that it can never provide evidence for the alternative hypothesis: if the false-positive probability exceeds the power, the WoE cannot favor $H_1$.","For designing informative experiments, lowering the false-positive rate improves the WoE more than increasing power: $\\alpha=0.01$ gives WoE $\\approx 19$, while 95% power at $\\alpha=0.05$ gives WoE $\\approx 12.8$.","For evidence about the absence of an effect, minimizing false negatives matters more than the nominal alpha.","The vitamin D/COVID-19 study, despite its very large sample, provides essentially no evidence against an association once low power from proxy vitamin D measurements and outcome misclassification are accounted for."],"supporting_citations":[{"why":"supplies the term 'ultimate issue error' and the legal framing of the fallacy.","marker":"Aitken et al. 2021"},{"why":"foundational source for weight of evidence as a logarithmic likelihood measure.","marker":"Good 1950"},{"why":"early statement of the likelihood approach that the WoE formula extends.","marker":"Peirce 2014"},{"why":"supplies the point that data make a hypothesis more or less plausible without fixing its probability.","marker":"Polya 1954"},{"why":"provides the empirical estimate of low average power used to motivate adjusting power downward.","marker":"Button et al. 2013"},{"why":"the vitamin D/COVID-19 study whose conclusion the paper re-evaluates.","marker":"Hastie et al. 2020"},{"why":"gives the longitudinal correlation of vitamin D measurements used to quantify proxy noise and lower power.","marker":"Jorde et al. 2010"},{"why":"gives a second longitudinal correlation estimate used in the same way.","marker":"Meng et al. 2012"},{"why":"documents likely misclassification in the COVID-19-negative group, used to lower power further.","marker":"Davies, Mazess, and Benskin 2021"}],"fun_headline_variants":["The ultimate issue: parameters ≠ hypotheses in stats","p<0.05 doesn't confirm your hypothesis—ever","Weight-of-evidence: statistical inference with context","How to actually test hypotheses: the WoE way","Null result? WoE shows zero support for no effect"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The worked examples assume that all a study's evidence can be compressed into a single binary outcome—significant or not—with false-positive probability $\\alpha$ and power, so the numerical WoE values depend on that simplification rather than on the actual continuous data.","fun_headline_variants_meta":{"raw":{"variants":["The ultimate issue: parameters ≠ hypotheses in stats","p<0.05 doesn't confirm your hypothesis—ever","Weight-of-evidence: statistical inference with context","How to actually test hypotheses: the WoE way","Null result? WoE shows zero support for no effect"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001158,"raw_usage":{"total_tokens":4797,"prompt_tokens":943,"completion_tokens":3854,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":3777}},"tokens_in":559,"tokens_out":3854,"duration_ms":31634,"temperature":1.0,"reasoning_tokens":3777,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:21:04.743936+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the vitamin D analysis using a likelihood ratio built from the continuous test statistic or the full logistic-regression likelihood instead of coding the result as $p>0.05$. If the resulting weight of evidence is not close to the paper's $-0.75$, then the binary $\\alpha$/power collapse is what produces the 'zero evidence' conclusion.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the term 'ultimate issue error' and the legal framing of the fallacy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"foundational source for weight of evidence as a logarithmic likelihood measure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"early statement of the likelihood approach that the WoE formula extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the point that data make a hypothesis more or less plausible without fixing its probability."},{"cited_title":"Mackay, Frederick Ho, Carlos A","cited_arxiv_id":null,"evidence_quote":"the vitamin D/COVID-19 study whose conclusion the paper re-evaluates."},{"cited_title":"Sneve, M","cited_arxiv_id":null,"evidence_quote":"gives the longitudinal correlation of vitamin D measurements used to quantify proxy noise and lower power."},{"cited_title":"Hovey, Jean Wactawski-Wende, Christopher A","cited_arxiv_id":null,"evidence_quote":"gives a second longitudinal correlation estimate used in the same way."},{"cited_title":"Mazess, and Linda L","cited_arxiv_id":null,"evidence_quote":"documents likely misclassification in the COVID-19-negative group, used to lower power further."}],"review_version":1}