Aligning self and other-referencing activations during fine-tuning reduced deceptive responses on tested LLM and RL benchmarks, with small capability costs.
Constitutional ai: Harmlessness from ai feedback
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AI 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Towards Safe and Honest AI Agents with Neural Self-Other Overlap
Aligning self and other-referencing activations during fine-tuning reduced deceptive responses on tested LLM and RL benchmarks, with small capability costs.