Verbalized Rejection Sampling reduces bias in LLM Bernoulli sampling by prompting the model to reason about and accept or reject proposed samples.
Do llms play dice? exploring probability distribution sampling in large language models for behavioral simulation
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
verdicts
UNVERDICTED 2representative citing papers
DynamicPO adds dynamic boundary-negative selection and dual-margin beta adjustment to multi-negative DPO to avoid gradient suppression and improve recommendation accuracy.
citing papers explorer
-
Flipping Against All Odds: Reducing LLM Coin Flip Bias via Verbalized Rejection Sampling
Verbalized Rejection Sampling reduces bias in LLM Bernoulli sampling by prompting the model to reason about and accept or reject proposed samples.
-
DynamicPO: Dynamic Preference Optimization for Recommendation
DynamicPO adds dynamic boundary-negative selection and dual-margin beta adjustment to multi-negative DPO to avoid gradient suppression and improve recommendation accuracy.