REVIEW 1 major objections 2 minor
Identifying Information from Observations with Uncertainty and Novelty
T0 review · 1 major / 2 minor · reviewed 2026-05-23 · grok-4.3
Pith's one-line read Observations supply identifying information that verifies or falsifies hypotheses as the data-generating process.
desk verdict The paper defines identifying information via an indicator function over a fixed finite hypothesis set and proves that PAC-Bayes sample complexity distributions are determined by their moments under ergodic stationary processes. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
An indicator function over a hypothesis set that determines if observations verify or falsify each hypothesis, computing the identifying information and sample complexity.
What would settle it
A counterexample where the sample complexity distribution of a PAC-Bayes learner with finite hypotheses cannot be determined or approximated from the moments of the prior probability distribution, or where novelty is not detected despite a misspecified hypothesis set.
Extended reading notes
Core claim
Identifying information are bits that verify or falsify a hypothesis as the data-generating process. Hypothesis identification and sample complexity are defined via the computation of an indicator function over a set of hypotheses. This bridges algorithmic and probabilistic information. The sample complexity properties are detailed for deterministic to ergodic stationary stochastic processes, connecting finite-step identification to asymptotic statistics and PAC-learning. Novel information is formalized as detection of a misspecified hypothesis set. A computable PAC-Bayes learner's sample complexity distribution is determined by its moments from the prior over a fixed finite hypothesis set,
Load-bearing premise
The set of hypotheses is fixed and finite, and the data-generating processes are ergodic and stationary.
Editorial extensions
If this is right
- The information theoretic characteristics of hypothesis identification computation are provable.
- Sample complexity connects finite observations to asymptotic statistics for ergodic processes.
- Novel information detection identifies when the hypothesis set is misspecified.
- Sample complexity distributions for PAC-Bayes learners are computable via moments of the prior.
Reading between the lines
- Learners could use this to adaptively decide when they have enough data for reliable identification.
- The approach might generalize to settings with infinite or continuous hypothesis spaces through suitable approximations.
- It provides a foundation for quantifying uncertainty reduction in sequential decision making under novelty.
- Connections to change detection in streaming data could be explored by monitoring shifts in identifying information.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces 'identifying information' as the bits from observations that verify or falsify a hypothesis as the data-generating process. It synthesizes prior algorithmic and probabilistic information theory, defines hypothesis identification and sample complexity via an indicator function over a hypothesis set, analyzes sample complexity for deterministic through ergodic stationary processes, formalizes novelty detection for misspecified sets, and proves that the sample complexity distribution of computable PAC-Bayes learners is determined by moments of the prior over a fixed finite hypothesis set (making approximations computable to arbitrary precision).
Significance. If the central derivations hold, the work offers a coherent bridge between algorithmic information and PAC-learning that directly addresses uncertainty and novelty detection. The explicit scoping to fixed finite hypothesis sets and ergodic processes, together with the computability result for the sample-complexity distribution, constitutes a concrete, falsifiable contribution that can be checked within those bounds. The parameter-free character of the moment-based approximation is a notable strength.
major comments (1)
- [PAC-Bayes sample complexity section] Abstract and § on PAC-Bayes result: the claim that the sample-complexity distribution 'is determined by its moments in terms of the prior' is load-bearing for the computability conclusion; the manuscript must exhibit the explicit moment-to-distribution mapping (or the relevant theorem) to confirm it does not tacitly re-introduce fitted parameters or infinite-dimensional objects.
minor comments (2)
- [Definition section] Notation for the indicator function and 'identifying information' should be introduced with a short comparison table to mutual information and Kolmogorov complexity to clarify the synthesis of prior works.
- [Sample complexity for ergodic processes] The transition from finite-step identification to asymptotic statistics for ergodic processes would benefit from an explicit statement of the ergodicity assumption in the theorem statement rather than only in the surrounding text.
Simulated Author's Rebuttal
We thank the referee for their careful review and constructive feedback on the manuscript. We address the major comment below and agree that an explicit mapping will strengthen the presentation.
read point-by-point responses
-
Referee: [PAC-Bayes sample complexity section] Abstract and § on PAC-Bayes result: the claim that the sample-complexity distribution 'is determined by its moments in terms of the prior' is load-bearing for the computability conclusion; the manuscript must exhibit the explicit moment-to-distribution mapping (or the relevant theorem) to confirm it does not tacitly re-introduce fitted parameters or infinite-dimensional objects.
Authors: We agree the explicit mapping is required for full rigor. The manuscript proves the result for fixed finite hypothesis sets but presents the moment-to-distribution link implicitly rather than via a dedicated theorem. In revision we will add an explicit statement: for finite H the sample-complexity random variable takes values in a finite set of identification times whose probabilities are polynomials in the prior masses p(h); these probabilities (hence the full distribution) are recoverable from the sequence of moments of the prior via the moment-generating function evaluated at the finite support. No auxiliary parameters or infinite-dimensional objects are introduced. This will be inserted into the PAC-Bayes section. revision: yes
Circularity Check
No significant circularity detected in derivation chain
full rationale
The paper's central contributions are a formalization of identifying information via an indicator function over a fixed finite hypothesis set, plus a proof that a computable PAC-Bayes sample-complexity distribution is determined by its moments with respect to the prior. These steps are presented as direct consequences of the definitions and the stated assumptions (fixed finite hypothesis set, ergodic stationary processes), without any reduction of the claimed results to fitted parameters, self-referential quantities, or load-bearing self-citations. The abstract and reader's summary explicitly scope the results to these assumptions rather than smuggling them in, and no quoted derivation equates an output to its input by construction. The work therefore remains self-contained.
Assumptions & free parameters
assumptions (2)
- standard math Standard axioms of probability theory and information measures
- domain assumption PAC learning framework assumptions including finite hypothesis class
invented entities (1)
-
Identifying information
Cite this review
Pith. "Pith review of Identifying Information from Observations with Uncertainty and Novelty." pith.science (2026). https://pith.science/paper/2501.09331
@misc{pith2026250109331,
author = {Pith},
title = {Pith review of: Identifying Information from Observations with Uncertainty and Novelty},
year = {2026},
howpublished = {\url{https://pith.science/paper/2501.09331}},
note = {Machine review of arXiv:2501.09331}
}
read the original abstract
A machine that learns a task from observations must encounter and process uncertainty and novelty, especially when it is to maintain performance when observing new information and to select the hypothesis that best fits the current observations. In this context, some key questions arise: what and how much information did the observations provide, how much information is required to identify the data-generating process, how many observations remain to get that information, and how does a predictor determine that it has observed novel information? We formalize identifying information to answer these questions and synthesize prior works. Identifying information are bits that verify or falsify a hypothesis as the data-generating process. In this formalization, we prove the information theoretic characteristics of the computation of hypothesis identification and the resulting sample complexity. We define hypothesis identification and sample complexity via the computation of an indicator function over a set of hypotheses, bridging algorithmic and probabilistic information. We detail the sample complexity and its properties for data-generating processes ranging from deterministic processes to ergodic stationary stochastic processes, which connect the notion of identifying information in finite steps with asymptotic statistics and PAC-learning. The indicator function's computation naturally formalizes novel information and its identification from observations with respect to a hypothesis set, which detects a misspecified hypothesis set. We also proved that a computable PAC-Bayes learners' sample complexity distribution is determined by its moments in terms of the prior probability distribution over a fixed finite hypothesis set, and thus an approximation of the sample complexity distribution is always computable within the desired precision that resources allow.
Reviewed May 23, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.