Pith. sign in

REVIEW 8 cited by

Disclosure and Mitigation of Gender Bias in LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.11190 v1 pith:VH2BRNGG submitted 2024-02-17 cs.CL

classification cs.CL
keywords genderbiasllmsexplicitevenstereotypesdiscloseimplicit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) can generate biased responses. Yet previous direct probing techniques contain either gender mentions or predefined gender stereotypes, which are challenging to comprehensively collect. Hence, we propose an indirect probing framework based on conditional generation. This approach aims to induce LLMs to disclose their gender bias even without explicit gender or stereotype mentions. We explore three distinct strategies to disclose explicit and implicit gender bias in LLMs. Our experiments demonstrate that all tested LLMs exhibit explicit and/or implicit gender bias, even when gender stereotypes are not present in the inputs. In addition, an increased model size or model alignment amplifies bias in most cases. Furthermore, we investigate three methods to mitigate bias in LLMs via Hyperparameter Tuning, Instruction Guiding, and Debias Tuning. Remarkably, these methods prove effective even in the absence of explicit genders or stereotypes.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Across six LLMs, 9.8 to 35.1 percent of statements receive different veracity labels when only the speaker's gender-coded job title changes.

  2. DeFrame: Debiasing Large Language Models Against Framing Effects

    cs.CL 2026-02 conditional novelty 6.0 of 10

    LLM fairness scores shift substantially with positive vs negative framing of the same question, and DeFrame—a three-step self-revision prompt—reduces both average bias and this framing gap.

  3. Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases

    cs.CL 2025-09 conditional novelty 6.0 of 10

    LLMs shift their gendered pronoun choices, toward 'they' and away from 'he', when prompts signal a gender-bias evaluation, so measured bias is highly dependent on prompt framing.

  4. Gender Inclusivity Fairness Index (GIFI): A Multilevel Framework for Evaluating Gender Diversity in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    GIFI is a seven-part fairness index showing that LLMs handle 'he' and 'she' far better than neutral and neopronouns, with GPT-4o ranking highest among 22 models.

  5. DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Caste-based stereotypes are measurably present in widely used LLMs, with the largest bias appearing when Dalits and Shudras are compared with dominant castes.

  6. Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A Chinese gender-bias corpus and shared-task benchmark show detection and classification are feasible, while automatic mitigation remains weak.

  7. Relative Bias: A Comparative Framework for Quantifying Bias in LLMs

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A model is 'relatively biased' when its responses deviate from the consensus of a baseline LLM set, and this deviation can be scored by embedding distances or LLM judges plus equivalence tests.

  8. LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models

    cs.CL 2025-05 reject novelty 4.0 of 10

    A block-localizing fine-tuning method for gender debiasing is presented, but its stated loss is inconsistent with its reported behavior and the evaluation tables contain duplicate rows.

Pith tools