{"id":"14659d7b-2c34-456e-afc1-cc0d5e6dc822","arxiv_id":"2502.06628","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Random variables are measurable functions and need no randomness; the paper reframes inference as logical contradiction and proposes information-based estimator assessment.","lead":"This paper argues that random variables are just measurable functions, not sources of randomness, and that statistical inference can be framed as logical contradiction rather than repeated sampling. Its practical hook is a proposal to judge estimators by information content instead of bias and variance.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed logical-induction step is not established: Haux is an analyst-chosen tail-area threshold, making the argument a restatement of significance testing.","rationale":"The reader's weakest assumption identifies the formal status of Haux, and that is the most load-bearing point. I agree and add that Haux is not merely hard to make objective; under the paper's own tail-area definition it is equivalent to a p-value threshold, so the proposed logical mode of inference is a re-description of Fisher's disjunction rather than an extension of reductio ad absurdum. The single-randomization concern is genuine but secondary: it only connects yobs to a sampling distribution and does not remove the need for a cutoff. The information-based assessment claim is also under-derived in this manuscript, being delegated to self-cited works without definitions or proofs, but it is downstream of the inference logic. The first contribution, that random variables are measurable functions and do not require hypothetical samples, is textbook and acceptable. Overall, a conditional accept remains appropriate if the paper is positioned as an expository summary of the author's prior framework; the novelty claims about logical inference and replacing bias/variance should be softened or explicitly marked as dependent on prior work.","tokens_in":10609,"tokens_out":5909,"duration_ms":54111,"concrete_test":"In the Two Lotteries setting, for arbitrary alpha define Haux(alpha) by TA(sobs, SA) > alpha and write the reductio conclusion. Show for every alpha that the non-contradiction set is exactly {m : p-value_m(yobs) > alpha}, i.e., the acceptance region of a level-alpha test. If the equivalence holds for all alpha, FIL is a relabeling of significance testing; a choice-independent derivation of alpha would be needed to rescue the logical-induction claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is Section 5.2's auxiliary hypothesis Haux: 'the observed sum is not exceptionally rare.' The paper says this can be formalized objectively through tail areas, but the formalization requires choosing both an ordering (test statistic) and a cutoff alpha, as the paper itself shows by calibrating to 0.0002 or 0.0171. With Haux defined as TA(yobs, m) > alpha, the reductio conclusion 'not (H0 and Haux)' is exactly Fisher's disjunction: either the model is false or the observation fell in an alpha-tail. The set M_alpha in Section 7.2 is precisely the acceptance region of a level-alpha test. Thus the framework does not extend deduction to induction in a new logical sense; it restates significance testing with a free alpha. Unless Haux can be derived from the model plus logic alone, without any analyst-selected cutoff, the paper's second and linked third contributions are not established as claimed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that random variables should be understood as measurable functions and that their distributions can be developed through measure-theoretic proportions without invoking infinite hypothetical repeated sampling. It uses Penrose's three-world distinction to locate randomization in the physical world and sampling distributions in the mathematical world. The paper then proposes an 'inductive reductio' based on an auxiliary hypothesis Haux that the observed value is not exceptionally rare, yielding Fisher's disjunction that either the hypothesis is false or a rare event occurred. Section 7 introduces generalized estimators and an information-based measure Lambda(g) that is claimed to replace bias and variance as assessment tools. The three stated contributions are: a characterization of random variables without hypothetical samples, an extension of deduction to induction, and an information-based assessment of statistical procedures.","tokens_in":10800,"tokens_out":5212,"duration_ms":48908,"significance":"If the framework is correct, it would provide a clean conceptual separation between mathematical distributions and physical sampling, a logical reading of p-values, and a label-space-invariant way to compare estimators. The paper is clearly written, the lottery example is worked carefully, and the discussion engages seriously with Fisher's actual views and with the p-value debate. Its strengths include a worked two-lottery example and an explicit link to information geometry. However, the technical core of the third contribution is deferred almost entirely to self-cited earlier work, and the inductive reductio depends on an analyst-chosen rarity threshold, so the significance as a standalone contribution is conditional on those gaps being resolved.","major_comments":[{"comment":"The formalization of Haux is not objective. Section 5.2 first labels a sum of 28 as rare and 27 as not rare, corresponding to a tail area of 0.0002 for lottery A, and then labels a sum of 24 as extreme and 23 as not, corresponding to a tail area of 0.0171. No principle in the paper selects one tail-area cutoff over the other. With Haux formalized in Section 7.2 as TA(yobs, m) > alpha, the reductio conclusion not(H0 and Haux) is exactly Fisher's disjunction 'either the model is false or the observation fell in an alpha-tail,' and the set M_alpha in Section 7.2 is precisely the acceptance region of a level-alpha test. Since both the test statistic (the ordering) and alpha are analyst choices, the claimed objective extension of deduction to induction is not established; unless Haux is derived from H0 plus logic alone, the argument restates significance testing with a free alpha.","section":"5.2 and 7.2"},{"comment":"The central technical apparatus for the third contribution is not present in this manuscript. The definition of the information content Lambda(g), the generalized estimators g(y, m), and the claim in Section 8 that for an unbiased estimator Lambda(theta-hat) equals the reciprocal of its variance are all deferred to Vos (2022), Vos and Wu (2024), and Vos (2024). No theorem or proof of these properties appears here, so the central claim that information replaces bias and variance cannot be independently evaluated from this paper. The manuscript should either state the definitions and derivations or explicitly mark these as review of prior work with the relevant results reproduced.","section":"7.3 and 8"},{"comment":"The logical step from H0 and Haux to the conclusion sobs in S27 presupposes the rarity ordering and threshold rather than deriving them. Haux as introduced is the statement 'the observed sum is not exceptionally rare,' which is a property of the observed value relative to a tail-area cutoff; by itself it does not place sobs in the set S27 unless one already assumes that 28 is rare and that the cutoff is the 99.98th percentile. The reductio therefore does not supply an objective bridge from deductive to inductive logic; it assumes the bridge in the formulation of Haux.","section":"5.2"},{"comment":"The first contribution is not backed by a formal statement in the body. The abstract claims that 'random variables, properly understood as measurable functions, can be fully characterized without appealing to infinite hypothetical samples,' but Section 2 defines distributions (simple, discrete, and continuous) and never states a formal definition of a random variable as a measurable function or proves a characterization theorem. Without such a definition, the first contribution is a philosophical restatement rather than a demonstrated mathematical result.","section":"2 and Abstract"}],"minor_comments":[{"comment":"The phrase 'order pair' should be 'ordered pair' in the description of the multi-set representation.","section":"2.1"},{"comment":"In the reference for Vos (2024), 'Etimators' should be 'Estimators'.","section":"References"},{"comment":"The notation XK is used both for the label space and for the distribution on that space (e.g., 'Pr(XK = x) = mK(x)'), which is confusing; a separate symbol for the distribution would clarify the exposition.","section":"2.2"},{"comment":"The comment that the likelihood ratio is constant across the outcomes {0,1,...,7} in the lottery example is not explained; a brief derivation or an explanatory sentence would help readers see why this matters for the ordering.","section":"5.3"},{"comment":"The symbol M_alpha^complement is used without defining the complement operation; the text should state explicitly that the complement is taken within the model family M.","section":"7.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript carries most of its technical weight through self-citations to the author's prior work, and the inductive-reductio claim reduces to significance testing unless the alpha selection is solved. This may be acceptable for a philosophical/foundational venue, but for a statistics journal the paper needs to be self-contained enough for a reader to verify the central claims. I did not review the cited prior papers; my assessment is based on the present manuscript alone."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a perspective piece, not a technical advance. It restates the standard measure-theoretic definition of random variables, uses Penrose's three-worlds framing to argue that randomization lives in the physical world, and then reprises the author's own FIL program. If you already know Vos's prior work, nothing here will surprise you; if you don't, this is a clean entry point.\n\nWhat it does well: the lottery example is worked correctly and makes the logical reductio easy to follow. The distinction between label-space structure and probability assignments is useful, and the critique of bias/MSE as unit-dependent is legitimate. The paper is also refreshingly explicit that Haux is formalized through tail areas with a threshold, rather than hiding that choice.\n\nWhere it gets soft: the central claims are deferred. The information metric Lambda(g), the reciprocal-variance identity, and generalized estimators are all cited to Vos (2022), Vos & Wu (2024), and Vos & Holbert (2022), with no derivations here. That makes the paper's third contribution a summary, not a result. The bigger issue is the logical-induction step. The stress-test note is right: Haux as TA(yobs, m) > alpha is exactly a significance test with a free alpha. The paper even shows the threshold being set at 0.0002 or 0.0171, which confirms the analyst's hand. So the conclusion 'either the model is false or a rare event occurred' is Fisher's disjunction, not a new mode of inference. The paper would be more honest if it presented itself as a survey of the author's framework rather than claiming a new logical foundation.\n\nThat said, the paper is internally coherent and the math that is actually shown is correct. The self-citation load is heavy but not dishonest, since the prior results do exist. The Penrose framing is a nice pedagogical device, and the discussion of terminology is sensible.\n\nWho is this for? Someone teaching or thinking about the foundations of frequentist inference who wants a non-Bayesian, non-repeated-sampling account. It won't change the practice of statistics.\n\nRecommendation: I would send this to peer review conditionally, with the expectation that the author either derives the key information results or explicitly reframes the paper as an expository summary. Desk rejection would be too harsh for a coherent, thought-provoking piece, but unconditional acceptance would reward packaging over proof.","headline":"A readable restatement of Fisher's disjunction with the technical load carried by the author's prior papers; the novelty is packaging, not proof.","tokens_in":11292,"tokens_out":1420,"would_cite":false,"duration_ms":14731,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62A01","62B10","62F03","62F25"],"pacs":[],"model":"deepseek-v4-flash","headline":"Random variables are measurable functions, and statistical inference can proceed by pure logic without hypothetical repeated sampling.","keywords":["random variables","measurable functions","statistical inference","reductio ad absurdum","tail areas","information theory","bias-variance tradeoff","frequentist foundations"],"falsifier":"Take one dataset and one null model, and compute the paper's reductio for two thresholds, say tail area 0.05 and 0.001; if one threshold yields 'hypothesis false or rare event' and the other yields no contradiction for the same observation, the logical verdict depends on the analyst's chosen α rather than on mathematics alone. Alternatively, exhibit two physical sampling plans that could both have produced the same observed sample but yield different combinatorial distributions, showing that a single randomization does not by itself fix the sampling distribution.","tokens_in":10349,"feed_emoji":"🎲","tokens_out":7556,"duration_ms":60347,"temperature":0.7,"pith_summary":"The paper argues that randomness is not a property of the mathematical objects statisticians call random variables; they are measurable functions, and their distributions are fully specified mathematical structures. It claims that defining and using these distributions requires no appeal to infinite hypothetical samples, and that the only real randomization involved is the single physical mechanism that produced the observed sample. From this, statistical inference is recast as a logical argument: if the observed data sit in the extreme tail of a hypothesized distribution, the conclusion is that either the hypothesis is false or a rare event has occurred. The paper then replaces bias and variance as criteria for judging estimators with information-based measures that depend only on the probabilities assigned to possible data, not on the units or ordering of the data. If these claims hold, long-standing disputes over repeated-sampling interpretations of p-values and confidence intervals dissolve into a purely mathematical foundation.","feed_headline":"Random variables aren't random—inference can be pure logic","feed_subtitle":"One physical draw links data to a distribution; bias and variance give way to information-based evaluation.","key_machinery":"The central mechanism is the Fisher-Information-Logic (FIL) approach: a modified reductio ad absurdum in which the probability that the observed sample falls in a tail of a hypothesized distribution supplies the contradiction. The objects doing the work are simple distributions—frequencies normalized on a finite label space—extended to discrete and continuous distributions by measure theory, and tail areas computed from them. 'Rare' is defined by a tail-area threshold α, which enters through an auxiliary hypothesis Haux, the statement that the observed sample is not exceptionally rare; the argument then partitions the model family into the set Mα where no contradiction arises and its complement where it does. These tail-based contradictions yield confidence regions, while generalized estimators and their information content replace bias and variance as the tools for comparing statistical procedures.","core_discovery":"The paper's central claim is that random variables, properly understood as measurable functions, are fully characterized by their associated distributions on label spaces, and that this characterization needs no reference to randomization or repeated sampling. It develops a hierarchy of distributions—simple distributions on finite label spaces, then discrete and continuous generalizations via measure theory—and shows how percentiles and tail areas locate an observed sample within a sampling distribution. From this it extends the reductio ad absurdum of classical deduction to induction: given a null hypothesis and an auxiliary hypothesis that the observed value is not exceptionally rare, a value in the extreme tail produces the logical conclusion 'either the hypothesis is false or a rare event has occurred.' The discovery is that this logical scaffolding, together with single-instance physical randomization, supports confidence regions and point estimation without hypothetical infinite samples, and that estimator assessment should be based on the information content of the probability assignments rather than on bias and variance defined through label-space structure.","pith_inferences":["A natural extension is to re-read established frequentist procedures—likelihood ratios, p-values, confidence sets—as logical location statements in model space; the formulas would survive while the repeated-sampling story would be optional.","One testable extension would compare orderings of estimators by information content against finite-sample prediction error in cases where bias and MSE disagree, to see whether the information criterion tracks practical performance.","The framework's attribution of randomness to physical sampling mechanisms alone could sharpen the justification of randomized experiments in causal inference, where the design feature that matters is the single realized randomization, not hypothetical re-randomizations.","Because the auxiliary rarity threshold α remains a choice, the purely logical status of induction could be tested by asking whether a principled, for example minimax or decision-theoretic, choice of α removes the residue of convention."],"forward_implications":["P-values become statements about where an observed sample sits in a fixed sampling distribution, so their justification no longer depends on what would happen across repeated datasets.","A single physical randomization—the one that produced the data—is enough to connect the observed sample to the mathematical sampling distribution; infinite hypothetical samples are unnecessary.","Confidence regions inherit the logical reading that either the true model lies in the region or the observed sample is in the tail of every excluded model's sampling distribution.","Estimator performance is evaluated by the information content of probability assignments, so conclusions do not change when the label space is re-parameterized, as with km/L versus L/km units.","In exponential families, point estimates can be treated as distributions in the model family itself, generalizing least squares through Kullback-Leibler divergence."],"supporting_citations":[{"why":"Pages 42–44 are the explicit source for extending reductio ad absurdum from deductive to inductive inference.","marker":"Fisher (1959)"},{"why":"Supplies the logical disjunction 'Either the hypothesis is not true, or an exceptionally rare outcome has occurred' that the FIL argument reproduces.","marker":"Fisher (1960)"},{"why":"Provides the three-worlds distinction (mathematical, physical, mental) used to separate distributions as mathematics from randomization as a physical process.","marker":"Penrose (2007)"},{"why":"The repeated-sampling reading of p-values the paper argues against; supplies the target interpretation.","marker":"Gelman & Loken (2014)"},{"why":"The 'always' smaller-MSE estimator result that motivates replacing bias and variance with information-based assessment.","marker":"Efron (2024)"},{"why":"Supports the adequacy of a single random sample for connecting data to the sampling distribution.","marker":"Vos & Holbert (2022)"},{"why":"Source of the generalized estimation framework that orders models by consistency with the observed sample.","marker":"Vos & Wu (2024)"},{"why":"Argues why information is superior to mean square error for assessing estimators.","marker":"Vos (2024)"}],"fun_headline_variants":["Random variables are deterministic, inference is logical","Statistics without random sampling: pure logic inference","Random variables defined by distribution, not randomness","Induction as reductio ad absurdum: new statistical logic","Bias and variance out, information in: new inference"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that one physical randomization is enough to connect the observed sample to its mathematical sampling distribution, together with the premise that 'exceptionally rare' has an objective, threshold-independent meaning as a tail area; if either premise fails, the induced disjunction is just an ordinary significance test.","fun_headline_variants_meta":{"raw":{"variants":["Random variables are deterministic, inference is logical","Statistics without random sampling: pure logic inference","Random variables defined by distribution, not randomness","Induction as reductio ad absurdum: new statistical logic","Bias and variance out, information in: new inference"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000186,"raw_usage":{"total_tokens":1294,"prompt_tokens":885,"completion_tokens":409,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":501,"completion_tokens_details":{"reasoning_tokens":336}},"tokens_in":501,"tokens_out":409,"duration_ms":4249,"temperature":1.0,"reasoning_tokens":336,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T14:51:25.115976+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one dataset and one null model, and compute the paper's reductio for two thresholds, say tail area 0.05 and 0.001; if one threshold yields 'hypothesis false or rare event' and the other yields no contradiction for the same observation, the logical verdict depends on the analyst's chosen α rather than on mathematics alone. Alternatively, exhibit two physical sampling plans that could both have produced the same observed sample but yield different combinatorial distributions, showing that a single randomization does not by itself fix the sampling distribution.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The 'always' smaller-MSE estimator result that motivates replacing bias and variance with information-based assessment."}],"review_version":1}