Pith. sign in

REVIEW 1 cited by

Investigating Bias Representations in Llama 2 Chat via Activation Steering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.00402 v1 pith:4TLNOFZX submitted 2024-02-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords biasmodelsteeringactivationbiaseschatgenderllama
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We address the challenge of societal bias in Large Language Models (LLMs), focusing on the Llama 2 7B Chat model. As LLMs are increasingly integrated into decision-making processes with substantial societal impact, it becomes imperative to ensure these models do not reinforce existing biases. Our approach employs activation steering to probe for and mitigate biases related to gender, race, and religion. This method manipulates model activations to direct responses towards or away from biased outputs, utilizing steering vectors derived from the StereoSet dataset and custom GPT4 generated gender bias prompts. Our findings reveal inherent gender bias in Llama 2 7B Chat, persisting even after Reinforcement Learning from Human Feedback (RLHF). We also observe a predictable negative correlation between bias and the model's tendency to refuse responses. Significantly, our study uncovers that RLHF tends to increase the similarity in the model's representation of different forms of societal biases, which raises questions about the model's nuanced understanding of different forms of bias. This work also provides valuable insights into effective red-teaming strategies for LLMs using activation steering, particularly emphasizing the importance of integrating a refusal vector.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions

    cs.CL 2025-09 conditional novelty 6.0 of 10

    KAPPA reduces the knowledge-prediction gap in LLMs by aligning a prediction-direction coordinate to a knowledge-direction coordinate in the residual stream, yielding accuracy gains on binary-choice MCQs and modest gai...

Pith tools