Aligning self and other-referencing activations during fine-tuning reduced deceptive responses on tested LLM and RL benchmarks, with small capability costs.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Towards Safe and Honest AI Agents with Neural Self-Other Overlap
Aligning self and other-referencing activations during fine-tuning reduced deceptive responses on tested LLM and RL benchmarks, with small capability costs.