One epoch of LoRA on a QueerNews corpus reduces WinoQueer anti-LGBTQIA+ bias scores by up to 50 points in Llama 3 8B, Mistral 7B, and Gemma 7B; soft-prompt tuning does not.
For soft-prompt tuning, we train 10 virtual tokens with random initialization, which corresponds to only about 0.0005% of the model's parameters
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
One epoch of LoRA on a QueerNews corpus reduces WinoQueer anti-LGBTQIA+ bias scores by up to 50 points in Llama 3 8B, Mistral 7B, and Gemma 7B; soft-prompt tuning does not.