An attention-alignment fine-tuning objective (KL, CE, and triplet losses) reduces some bias scores on BBQ and BOLD for Llama-3.2-3B, but not consistently across models and with notable accuracy drops.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
An attention-alignment fine-tuning objective (KL, CE, and triplet losses) reduces some bias scores on BBQ and BOLD for Llama-3.2-3B, but not consistently across models and with notable accuracy drops.