Soft Best-of-n sampling provably approaches the optimal tilted reward distribution at O(1/n) KL divergence and relative reward error, with sample complexity that grows exponentially in sequence length for blockwise sampling.
Nonsymmetrical distance between probability distribu- tions, entropy and the theorem of pythagoras,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IT 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Soft Best-of-n Sampling for Model Alignment
Soft Best-of-n sampling provably approaches the optimal tilted reward distribution at O(1/n) KL divergence and relative reward error, with sample complexity that grows exponentially in sequence length for blockwise sampling.