A well-trained next-token predictor's in-context loss equals the conditional entropy of the data process, which must decrease with context for stationary data.
Signatures of Infinity: Nonergodicity and Resource Scaling in Prediction, Complexity, and Learning
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We introduce a simple analysis of the structural complexity of infinite-memory processes built from random samples of stationary, ergodic finite-memory component processes. Such processes are familiar from the well known multi-arm Bandit problem. We contrast our analysis with computation-theoretic and statistical inference approaches to understanding their complexity. The result is an alternative view of the relationship between predictability, complexity, and learning that highlights the distinct ways in which informational and correlational divergences arise in complex ergodic and nonergodic processes. We draw out consequences for the resource divergences that delineate the structural hierarchy of ergodic processes and for processes that are themselves hierarchical.
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Next-token pretraining implies in-context learning
A well-trained next-token predictor's in-context loss equals the conditional entropy of the data process, which must decrease with context for stationary data.