REVIEW 6 cited by
Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) have shown remarkable advances in language generation and understanding but are also prone to exhibiting harmful social biases. While recognition of these behaviors has generated an abundance of bias mitigation techniques, most require modifications to the training data, model parameters, or decoding strategy, which may be infeasible without access to a trainable model. In this work, we leverage the zero-shot capabilities of LLMs to reduce stereotyping in a technique we introduce as zero-shot self-debiasing. With two approaches, self-debiasing via explanation and self-debiasing via reprompting, we show that self-debiasing can significantly reduce the degree of stereotyping across nine different social groups while relying only on the LLM itself and a simple prompt, with explanations correctly identifying invalid assumptions and reprompting delivering the greatest reductions in bias. We hope this work opens inquiry into other zero-shot techniques for bias mitigation.
Forward citations
Cited by 6 Pith papers
-
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
The paper introduces DiffAware and CtxtAware metrics, an 8-benchmark suite with 16,000 questions, and shows ten LLMs are less able to recognize legitimate group differences than current fairness benchmarks suggest.
-
BiasFilter: An Inference-Time Debiasing Framework for Large Language Models
BiasFilter filters low-fairness segments during LLM generation using a reward model trained on a GPT-4-scored preference dataset, cutting bias on CEB and FairMT.
-
FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
FairSteer uses a linear probe to detect biased activations and adds a contrastively computed steering vector to shift generation toward unbiased answers, cutting bias across six LLMs without retraining.
-
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
LLM bias scores change depending on whether 'unbiased' means equal treatment across groups or close alignment with US workforce statistics.
-
A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms
A survey that organizes foundation-model recommender systems into feature-based, generative, and agentic paradigms and reviews tasks, empirical results, and open challenges.
-
Smoothie-Qwen: Post-Hoc Smoothing to Reduce Language Bias in Multilingual LLMs
Smoothie-Qwen reduces Chinese-language output in Qwen2.5-Coder by scaling down lm head weights of Chinese tokens, reaching 95% suppression on a synthetic elicitation set while keeping Korean MMLU accuracy roughly stable.
Discussion (0). Continue with ORCID to comment.