REVIEW 21 cited by
A Survey on Fairness in Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) have shown powerful performance and development prospects and are widely deployed in the real world. However, LLMs can capture social biases from unprocessed training data and propagate the biases to downstream tasks. Unfair LLM systems have undesirable social impacts and potential harms. In this paper, we provide a comprehensive review of related research on fairness in LLMs. Considering the influence of parameter magnitude and training paradigm on research strategy, we divide existing fairness research into oriented to medium-sized LLMs under pre-training and fine-tuning paradigms and oriented to large-sized LLMs under prompting paradigms. First, for medium-sized LLMs, we introduce evaluation metrics and debiasing methods from the perspectives of intrinsic bias and extrinsic bias, respectively. Then, for large-sized LLMs, we introduce recent fairness research, including fairness evaluation, reasons for bias, and debiasing methods. Finally, we discuss and provide insight on the challenges and future directions for the development of fairness in LLMs.
Forward citations
Cited by 21 Pith papers
-
More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing
Lifelong Normalization combined with ridge-regularized regression produces asymptotically orthogonal and bounded parameter updates that mitigate forgetting and collapse in lifelong model editing.
-
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
StereoTales shows that all tested LLMs emit harmful stereotypes in open-ended stories, with associations adapting to prompt language and targeting locally salient groups rather than transferring uniformly across languages.
-
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
StereoTales shows that LLMs produce harmful, culturally adapted stereotypes in open-ended multilingual stories, with patterns consistent across providers and aligned human-LLM harm judgments.
-
SCOPE: A Dataset of Stereotyped Prompts for Counterfactual Fairness Assessment of LLMs
SCOPE is a new large-scale dataset of counterfactual prompt pairs for evaluating fairness and stereotype sensitivity in LLMs across 1,438 topics, nine bias dimensions, 1,536 groups, and four communicative intents.
-
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
A human-in-the-loop audit of system prompts from 88 commercial AI products finds protective instructions nearly universal yet incomplete, with ~40% of products containing at least one user-harmful directive.
-
Whose fairness? Structural concentration in AI bias research
Bibliometric and semantic analysis of 692 AI-bias papers shows US-led structural concentration is strongest in the general-fairness domain that supplies definitions and benchmarks to the rest of the field.
-
Estimating Grammatical Gender Directions in Contextual Embeddings under Controlled and Natural Contexts
A framework estimates grammatical gender directions in contextual embeddings via controlled and natural contexts, finding unweighted controlled contexts and centroid estimators yield the purest directions.
-
The Wrong Kind of Right: Quantifying and Localizing Misfired Alignment in LLMs
LLMs exhibit misfired alignment on stereotype questions at 4.7-18.9% rates on the new VETO benchmark of 2,032 contrastive pairs, unlike humans at 0%, due to overgeneralized safety cues after instruction tuning.
-
More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing
Online value-gradient normalization in lifelong LLM editing produces bounded, asymptotically orthogonal parameter updates; an explicit warm-up and full whitening (StableEdit) strengthen this effect and improve long-ho...
-
FairNVT: Improving Fairness via Noise Injection in Vision Transformers
FairNVT injects calibrated noise into sensitive embeddings of transformer encoders to jointly improve representation-level and prediction-level fairness metrics without degrading task performance.
-
Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
Benign PEFT fine-tuning changes LLM safety and fairness: adapter-based methods (LoRA, IA3) preserve alignment better than prompt-based methods, and the base model strongly moderates outcomes.
-
Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
LLMs exhibit identity-dependent hedging on human rights questions, with group identity as the strongest predictor among tested factors, and group steering mitigates the disparity.
-
Pareto-Guided Teacher Alignment for Fair Personalized Text Generation
Fairness mitigation in personalized text generation is objective-dependent with methods occupying different regions of the fairness-personalization Pareto frontier rather than any single strategy dominating all objectives.
-
Resume-ing Control: (Mis)Perceptions of Agency Around GenAI Use in Recruiting Workflows
Recruiters perceive themselves as retaining agency over GenAI in hiring pipelines, yet GenAI invisibly architects core evaluation inputs, producing only marginal efficiency gains at the cost of deskilling.
-
Intersectional Fairness in Large Language Models
LLMs are more accurate when answers match stereotypes in clear contexts, especially for race-gender combinations, and no tested model shows consistent fairness or reliability across intersectional groups.
-
A Close Reading Approach to Gender Narrative Biases in AI-Generated Stories
A close reading of 15 AI-generated stories finds that even when character counts are balanced, narrative roles, descriptions, and plot dynamics remain gender-stereotyped (e.g., every villain is male).
-
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
The abstract claims LLMs show up to 40% coreference confidence disparities across intersectional identities, but the article body is an unrelated paper on robotic fruit handling.
-
Fairness Testing of Large Language Models in Role-Playing
Generates 550 roles and 33,000 questions to evaluate 10 LLMs in role-playing, finding 107,580 biased responses.
-
Examining Agents' Bias Amplification versus Suppression in Multi-Agent Systems
Empirical tests show that uniformly biased agents in multi-agent LLM systems produce system-wide bias exceeding the sum of individual biases, quantified via a new Favor Bias Strength metric.
-
Ethical Medical Image Synthesis
A submission whose abstract and full text are two different papers on unrelated topics.
-
A Survey on the Memory Mechanism of Large Language Model based Agents
A systematic review of memory designs, evaluation methods, applications, limitations, and future directions for LLM-based agents.
Discussion (0). Sign in to comment.