The sample complexity of exact-trace learning for autoregressive Chain-of-Thought is O((DSdim(H) + log(1/δ))/ε), matching the local next-token class with no dependence on rollout length.
The simulation below is oracle-assisted: it uses a halting oracle to decide whether the current local prediction map terminates from the queried start state
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
The Optimal Sample Complexity of Learning Autoregressive Chain-of-Thought
The sample complexity of exact-trace learning for autoregressive Chain-of-Thought is O((DSdim(H) + log(1/δ))/ε), matching the local next-token class with no dependence on rollout length.