For stochastic next-token generators, chain-of-thought supervision costs at most base learning at scale epsilon/M^2, while end-to-end learning costs at most about M/epsilon times chain-of-thought, with tight worst-case constructions.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Stochastic Autoregressive Learning
For stochastic next-token generators, chain-of-thought supervision costs at most base learning at scale epsilon/M^2, while end-to-end learning costs at most about M/epsilon times chain-of-thought, with tight worst-case constructions.