Active context sampling algorithm for contextual linear bandits achieves instance-dependent guarantees improving over minimax rate by up to sqrt(d) and reduces samples needed in empirical tasks.
Sample-efficient alignment for llms
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2verdicts
UNVERDICTED 2representative citing papers
ActiveDPO is a theoretically grounded active data selection method for sample-efficient LLM alignment that parameterizes the reward model directly with the LLM being aligned.
citing papers explorer
-
Active Learning for Stochastic Contextual Linear Bandits
Active context sampling algorithm for contextual linear bandits achieves instance-dependent guarantees improving over minimax rate by up to sqrt(d) and reduces samples needed in empirical tasks.
-
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
ActiveDPO is a theoretically grounded active data selection method for sample-efficient LLM alignment that parameterizes the reward model directly with the LLM being aligned.