REVIEW 4 cited by
Learning to Plan for Language Modeling from Unlabeled Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
By training to predict the next token in an unlabeled corpus, large language models learn to perform many tasks without any labeled data. However, their next-token-prediction objective arguably limits their performance in scenarios that require planning, such as writing a coherent article. In this paper, we train a module for planning the future writing process via a self-supervised learning objective. Given the textual context, this planning module learns to predict future abstract writing actions, which correspond to centroids in a clustered text embedding space. By conditioning on these actions, our model extends the successful language model formula to more abstract planning in an unsupervised way. Empirically, we demonstrate that our method improves language modeling performance in general, particularly with respect to the text structure. Because our framework uses a planner module that is unsupervised and external to the language model, new planner modules can be trained at large scale and easily be shared with the community.
Forward citations
Cited by 4 Pith papers
-
Large Concept Models: Language Modeling in a Sentence Representation Space
A sentence-level language model trained to autoregressively predict SONAR sentence embeddings can summarize, expand, and generate text in unseen languages.
-
LoopMTP: A looped transformer guided by latent multi-token prediction
Aligning each loop iteration's hidden state with a future token's embedding improves looped transformer accuracy by up to 8.1% relative over a non-looped baseline.
-
Emergent Response Planning in LLMs
Hidden representations of LLM prompts encode global attributes of the upcoming response, and simple probes can predict length, content choices, and answer confidence before generation begins.
-
Temporal horizons in forecasting: a performance-learnability trade-off
Longer training horizons improve forecast quality but worsen learnability, with loss-landscape roughness growing exponentially for chaotic dynamics and linearly for limit cycles.
Discussion (0). Continue with ORCID to comment.