Multi-token prediction helps transformers plan by enabling a reverse-reasoning circuit, but the proof that next-token prediction cannot learn this omits a key part of the next-token loss.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2representative citing papers
Adding a next-latent prediction loss to next-token training makes transformer hidden states more predictive of future tokens and improves planning/reasoning on small benchmarks.
citing papers explorer
-
How Transformers Learn to Plan via Multi-Token Prediction
Multi-token prediction helps transformers plan by enabling a reverse-reasoning circuit, but the proof that next-token prediction cannot learn this omits a key part of the next-token loss.
-
Next-Latent Prediction Transformers Learn Compact World Models
Adding a next-latent prediction loss to next-token training makes transformer hidden states more predictive of future tokens and improves planning/reasoning on small benchmarks.