Under a new structural assumption (DDMC), token-level bandit algorithms achieve sublinear regret and greedy LLM decoding is shown to be near-optimal.
Balancing work and self-care is
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Tokenized Bandit for LLM Decoding and Alignment
Under a new structural assumption (DDMC), token-level bandit algorithms achieve sublinear regret and greedy LLM decoding is shown to be near-optimal.