← back to paper
arxiv: 2508.09019 · 2 revisions
Activation Steering for Bias Mitigation: An Interpretable Approach to Safer LLMs