Presents CQB-η-2 algorithm achieving 𝒪̃(T^{-1/2}) queue length regret in contextual queueing bandits under stochastic contexts, with matching Ω(T^{-1/2}) lower bound.
Llm routing with dueling feedback
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
Average token log-probability provides a zero-shot confidence signal for small LLMs that matches supervised baselines in-distribution and outperforms them out-of-distribution, with a new retrieval-conditional variant improving further at lower latency.
A dueling bandit algorithm with belief-aware upper confidence bound is introduced for efficient, interaction-based selection of LLMs matching user latent preferences.
citing papers explorer
-
Algorithm for Contextual Queueing Bandits with Rate-Optimal Queue Length Regret
Presents CQB-η-2 algorithm achieving 𝒪̃(T^{-1/2}) queue length regret in contextual queueing bandits under stochastic contexts, with matching Ω(T^{-1/2}) lower bound.
-
Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training
Average token log-probability provides a zero-shot confidence signal for small LLMs that matches supervised baselines in-distribution and outperforms them out-of-distribution, with a new retrieval-conditional variant improving further at lower latency.
-
CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM
A dueling bandit algorithm with belief-aware upper confidence bound is introduced for efficient, interaction-based selection of LLMs matching user latent preferences.