Pith. sign in

REVIEW 2 cited by

Promoting Equality in Large Language Models: Identifying and Mitigating the Implicit Bias based on Bayesian Theory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.10608 v1 pith:5RURQ7W6 submitted 2024-08-20 cs.CL cs.AI

classification cs.CLcs.AI
keywords biasbiasesllmsbtbrimplicitbayesianbiasedextensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are trained on extensive text corpora, which inevitably include biased information. Although techniques such as Affective Alignment can mitigate some negative impacts of these biases, existing prompt-based attack methods can still extract these biases from the model's weights. Moreover, these biases frequently appear subtly when LLMs are prompted to perform identical tasks across different demographic groups, thereby camouflaging their presence. To address this issue, we have formally defined the implicit bias problem and developed an innovative framework for bias removal based on Bayesian theory, Bayesian-Theory based Bias Removal (BTBR). BTBR employs likelihood ratio screening to pinpoint data entries within publicly accessible biased datasets that represent biases inadvertently incorporated during the LLM training phase. It then automatically constructs relevant knowledge triples and expunges bias information from LLMs using model editing techniques. Through extensive experimentation, we have confirmed the presence of the implicit bias problem in LLMs and demonstrated the effectiveness of our BTBR approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Language models show negligible persona-based differences on MMLU benchmarks but large, income-relevant differences when asked for salary negotiation advice.

  2. Large Language Models and Provenance Metadata for Determining the Relevance of Images and Videos in News Stories

    cs.CL 2025-02 conditional novelty 5.0 of 10

    A prototype combines LLM reasoning with C2PA provenance metadata to classify news images and videos as relevant or not, without any benchmark evaluation.

Pith tools