Hard-thresholding next tokens by a safety value function yields higher expected safety, an explicit type-I intervention bound controlled by one threshold, and better empirical safety-helpfulness-similarity trade-offs than Gibbs-style or ARGS steering.
Value augmented sampling for language model alignment and personalization
3 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
UniR is a composable reasoning module trained with verifiable rewards and added to frozen LLMs via logit summation, enabling modular composition and weak-to-strong generalization across tasks and model sizes.
Maximizing joint conditional mutual information I(Y; C_Z, W, Z | X) decomposes multi-objective LLM alignment into preference-specific DPO terms plus an I(Y;W|X) exploration term that reduces reward-distribution overlap.
citing papers explorer
-
Selective Safety Steering via Value-Filtered Decoding
Hard-thresholding next tokens by a safety value function yields higher expected safety, an explicit type-I intervention bound controlled by one threshold, and better empirical safety-helpfulness-similarity trade-offs than Gibbs-style or ARGS steering.
-
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
UniR is a composable reasoning module trained with verifiable rewards and added to frozen LLMs via logit summation, enabling modular composition and weak-to-strong generalization across tasks and model sizes.
-
Multi-Objective Exploration and Preference Optimization via Mutual Information
Maximizing joint conditional mutual information I(Y; C_Z, W, Z | X) decomposes multi-objective LLM alignment into preference-specific DPO terms plus an I(Y;W|X) exploration term that reduces reward-distribution overlap.