The paper claims that no polynomial-time decoder is optimal for all distributions under N-gram Hamming loss, that random sampling is consistent for sequence cross-entropy, and that temperature scaling only works at temperature one; the central lower-bound proof is invalid as written.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms
The paper claims that no polynomial-time decoder is optimal for all distributions under N-gram Hamming loss, that random sampling is consistent for sequence cross-entropy, and that temperature scaling only works at temperature one; the central lower-bound proof is invalid as written.