Pith. sign in

REVIEW 6 cited by

Reinforcement Learning from Multi-role Debates as Feedback for Bias Mitigation in LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.10160 v6 pith:Q46LMPYG submitted 2024-04-15 cs.AI

classification cs.AI
keywords biasllmsdebatesfeedbackmitigationmulti-roleapproachlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Bias in LLMs can harm user experience and societal outcomes. However, current bias mitigation methods often require intensive human feedback, lack transferability to other topics or yield overconfident and random outputs. We find that involving LLMs in role-playing scenario boosts their ability to recognize and mitigate biases. Based on this, we propose Reinforcement Learning from Multi-role Debates as Feedback (RLDF), a novel approach for bias mitigation replacing human feedback in traditional RLHF. We utilize LLMs in multi-role debates to create a dataset that includes both high-bias and low-bias instances for training the reward model in reinforcement learning. Our approach comprises two modes: (1) self-reflection, where the same LLM participates in multi-role debates, and (2) teacher-student, where a more advanced LLM like GPT-3.5-turbo guides the LLM to perform this task. Experimental results across different LLMs on BBQ and our datasets demonstrate the effectiveness of our approach in bias mitigation. Our source code and datasets are available at \texttt{https://anonymous.4open.science/r/RLDF-E344}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges

    cs.SE 2026-07 accept novelty 6.0 of 10

    A systematic review of 141 papers derives a three-axis taxonomy of multi-agent debate design (participants, interaction, agreement) and shows the field has converged on a narrow default pattern.

  2. Mitigation of Gender and Ethnicity Bias in AI-Generated Stories through Model Explanations

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Feeding a model's own explanation of its biased story output back into a rewritten prompt improves demographic parity by 2% to 20%.

  3. MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion

    cs.CL 2025-07 conditional novelty 6.0 of 10

    MPF aligns LLM output sentiment with a target distribution by fitting weights over five perspective prompts and sampling responses accordingly.

  4. BiasFilter: An Inference-Time Debiasing Framework for Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    BiasFilter filters low-fairness segments during LLM generation using a reward model trained on a GPT-4-scored preference dataset, cutting bias on CEB and FairMT.

  5. A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming

    cs.CR 2025-05 reject novelty 5.0 of 10

    A reward-driven, PPO-finetuned LLM pipeline claims to generate diverse, evasive webshell payloads with higher escape rates than prompt-engineering baselines.

  6. Detection, Classification, and Mitigation of Gender Bias in Large Language Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A Chinese gender-bias system using SFT, chain-of-thought, and DPO with GPT-4-generated preference pairs reports top validation scores and first place on all three NLPCC 2025 subtasks.

Pith tools