REVIEW 14 cited by
StereoSet: Measuring stereotypical bias in pretrained language models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or Asians are bad drivers. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on large real world data, they are known to capture stereotypical biases. In order to assess the adverse effects of these models, it is important to quantify the bias captured in them. Existing literature on quantifying bias evaluates pretrained language models on a small set of artificially constructed bias-assessing sentences. We present StereoSet, a large-scale natural dataset in English to measure stereotypical biases in four domains: gender, profession, race, and religion. We evaluate popular models like BERT, GPT-2, RoBERTa, and XLNet on our dataset and show that these models exhibit strong stereotypical biases. We also present a leaderboard with a hidden test set to track the bias of future language models at https://stereoset.mit.edu
Forward citations
Cited by 14 Pith papers
-
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
Audit format (rate vs rank vs allocate; transparent vs disguised) reverses the apparent direction of LLM demographic bias, while causal framing of need dominates allocations by roughly an order of magnitude.
-
Dutch CrowS-Pairs: Adapting a Challenge Dataset for Measuring Social Biases in Language Models for Dutch
The paper presents a Dutch adaptation of the CrowS-Pairs bias benchmark and reports bias scores for seven masked and two autoregressive language models across nine demographic categories.
-
McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
A new Chinese bias benchmark with 4,077 instances and five tasks indicates larger language models are less biased than smaller ones when bias is measured through understanding tasks.
-
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
A retrieval-augmented generation system can be poisoned with reward-optimized biased documents and vector-space manipulation to substantially increase biased LLM outputs.
-
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
BiasLens uses concept activation vectors and sparse autoencoders to estimate LLM bias from internal representations, reporting moderate to strong agreement with behavioral bias metrics in a small evaluation.
-
Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques
A text-based AI model predicts which standardized test items will be permanently rejected with AUC 0.80 overall and 0.86 for math, though it misses most bias-related rejections.
-
Domyn-Small: A European 10B Reasoning Language Model
Domyn-Small is a 10B reasoning LLM that claims to deliver roughly one-third the inference tokens of Qwen3.5-9B at competitive accuracy, though results are marked as preliminary.
-
BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models
BioPro uses orthogonal projection on a gender-variation subspace to selectively debias vision-language models, reducing gender bias in neutral contexts while preserving explicit gender cues.
-
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
The abstract claims LLMs show up to 40% coreference confidence disparities across intersectional identities, but the article body is an unrelated paper on robotic fruit handling.
-
PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
One epoch of LoRA on a QueerNews corpus reduces WinoQueer anti-LGBTQIA+ bias scores by up to 50 points in Llama 3 8B, Mistral 7B, and Gemma 7B; soft-prompt tuning does not.
-
Fairness Dynamics During Training
Gender bias in Pythia-6.9b grows sharply after about 80k training steps even as general performance improves, and stopping earlier could trade 1.7% LAMBADA accuracy for a large fairness gain.
-
Probing AI Safety with Source Code
Code-style prompts such as make_more_toxic('text') consistently elicit more toxic output from seven current LLMs than equivalent natural-language instructions, and recursive application amplifies the effect.
-
The Science of Evaluating Foundation Models
A survey-and-checklist proposal that organizes LLM evaluation into an ABCD framework (Algorithm, Big Data, Computation, Domain Expertise) for context-aware, documented assessment.
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
Discussion (0). Continue with ORCID to comment.