REVIEW 1 cited by
Predictive Information
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Observations on the past provide some hints about what will happen in the future, and this can be quantified using information theory. The ``predictive information'' defined in this way has connections to measures of complexity that have been proposed both in the study of dynamical systems and in mathematical statistics. In particular, the predictive information diverges when the observed data stream allows us to learn an increasingly precise model for the dynamics that generate the data, and the structure of this divergence measures the complexity of the model. We argue that divergent contributions to the predictive information provide the only measure of complexity or richness that is consistent with certain plausible requirements.
Forward citations
Cited by 1 Pith paper
-
Next-token pretraining implies in-context learning
A well-trained next-token predictor's in-context loss equals the conditional entropy of the data process, which must decrease with context for stationary data.
Discussion (0). Continue with ORCID to comment.