Pith. sign in

REVIEW 2 cited by

Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.03537 v1 pith:COOA6RAR submitted 2024-12-04 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords modelsfairnesspre-trainedwhenbiaslanguagellmstransfer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are increasingly being adapted to achieve task-specificity for deployment in real-world decision systems. Several previous works have investigated the bias transfer hypothesis (BTH) by studying the effect of the fine-tuning adaptation strategy on model fairness to find that fairness in pre-trained masked language models have limited effect on the fairness of models when adapted using fine-tuning. In this work, we expand the study of BTH to causal models under prompt adaptations, as prompting is an accessible, and compute-efficient way to deploy models in real-world systems. In contrast to previous works, we establish that intrinsic biases in pre-trained Mistral, Falcon and Llama models are strongly correlated (rho >= 0.94) with biases when the same models are zero- and few-shot prompted, using a pronoun co-reference resolution task. Further, we find that bias transfer remains strongly correlated even when LLMs are specifically prompted to exhibit fair or biased behavior (rho >= 0.92), and few-shot length and stereotypical composition are varied (rho >= 0.97). Our findings highlight the importance of ensuring fairness in pre-trained LLMs, especially when they are later used to perform downstream tasks via prompt adaptation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Guiding LLM Decision-Making with Fairness Reward Models

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A single process-level reward model, trained on weakly labeled biased versus unbiased reasoning, transfers across tasks and models to reduce equalized odds gaps in LLM decision-making.

  2. Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    UCerF scores LLM fairness by both correctness and confidence, and SynthBias provides 31,756 gender-occupation coreference samples for benchmark testing.

Pith tools