A single linear transformation aligns concept representations between different LLMs, so steering vectors transfer across models and even from small to large models.
Aligning large language models with human preferences through representation engineering
5 Pith papers cite this work, alongside 4 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Direct distribution alignment (DCO) consistently improves cross-lingual consistency across models and datasets, while cross-domain transfer is limited and culture-diverse knowledge is largely preserved in closed-form evaluation.
The paper introduces the Proxy Compression Hypothesis as a unifying framework explaining reward hacking in RLHF as an emergent result of compressing high-dimensional human objectives into proxy reward signals under optimization pressure.
NTIL replaces cross-entropy with an exponential-weighted Earth Mover's Distance plus sequence-level relative and magnitude deviation losses, and improves numerical prediction accuracy in several LLM and MLLM fine-tuning experiments.
Proxy RL produces a staged proxy-internalization capability that emerges before and predicts reward hacking in coding environments.
citing papers explorer
-
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
The paper introduces the Proxy Compression Hypothesis as a unifying framework explaining reward hacking in RLHF as an emergent result of compressing high-dimensional human objectives into proxy reward signals under optimization pressure.