A decoder-only model trained to predict the second-to-last token can slightly improve next-token predictions when used to re-rank a GPT's top-k candidates.
In: Proceedings of the Con- ference on Empirical Methods in Natural Language Processing (2021)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
A decoder-only model trained to predict the second-to-last token can slightly improve next-token predictions when used to re-rank a GPT's top-k candidates.