REVIEW 10 cited by
Measuring and Reducing Gendered Correlations in Pre-trained Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Pre-trained models have revolutionized natural language understanding. However, researchers have found they can encode artifacts undesired in many applications, such as professions correlating with one gender more than another. We explore such gendered correlations as a case study for how to address unintended correlations in pre-trained models. We define metrics and reveal that it is possible for models with similar accuracy to encode correlations at very different rates. We show how measured correlations can be reduced with general-purpose techniques, and highlight the trade offs different strategies have. With these results, we make recommendations for training robust models: (1) carefully evaluate unintended correlations, (2) be mindful of seemingly innocuous configuration differences, and (3) focus on general mitigations.
Forward citations
Cited by 10 Pith papers
-
From Measurement to Mitigation: Exploring the Transferability of Debiasing Approaches to Gender Bias in Maltese Language Models
English-style debiasing methods only partially transfer to Maltese language models, with Counterfactual Data Augmentation most effective but hampered by grammatical errors.
-
FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes
A new India-focused benchmark shows that popular LLMs exhibit measurable negative bias against marginalized Indian identities and frequently reinforce caste, religion, region, and tribe stereotypes.
-
Position: It's Time to Optimize LLMs for Self-Consistency
The paper proposes self-consistency, a mathematical framework that treats relationships between model outputs across related inputs as the primary training target, unifying many existing alignment and robustness methods.
-
CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models
A training method that injects token-level causal labels into attention improves out-of-distribution accuracy on a synthetic benchmark and slightly on math/reasoning tasks.
-
KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
An attention-alignment fine-tuning objective (KL, CE, and triplet losses) reduces some bias scores on BBQ and BOLD for Llama-3.2-3B, but not consistently across models and with notable accuracy drops.
-
Paying Alignment Tax with Contrastive Learning
A contrastive learning framework with positive and negative example pairs improves faithfulness and slightly reduces toxicity on Reddit TL;DR summarization, but the central claim of avoiding the alignment tax is not e...
-
No LLM Solved Yu Tsumura's 554th Problem
A benchmark problem that is IMO-adjacent and whose solution is public was not solved by any tested commercial or open-source LLM.
-
Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
LLM-based three-class gender labeling agrees with human annotations better than the lexical NFaiRR score, and the proposed CWEx metric combines neutral exposure with male-female exposure disparity for ranking fairness...
-
Mitigating Confounding in Speech-Based Dementia Detection through Weight Masking
Masking weights that react to gender in a fine-tuned BERT reduces gender gaps in dementia predictions while keeping most of the detection accuracy.
-
Advertising in AI systems: Society must be vigilant
Generative AI outputs will likely carry embedded commercial content, and the paper proposes design principles, provenance tracking, and two debiasing strategies to preserve transparency.
Discussion (0). Sign in to comment.