Pith. sign in

REVIEW 9 cited by

Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.03907 v1 pith:YA2PYQLL submitted 2024-08-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords evaluationbiasllmsmethodsmetricshumanlanguagemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have excelled at language understanding and generating human-level text. However, even with supervised training and human alignment, these LLMs are susceptible to adversarial attacks where malicious users can prompt the model to generate undesirable text. LLMs also inherently encode potential biases that can cause various harmful effects during interactions. Bias evaluation metrics lack standards as well as consensus and existing methods often rely on human-generated templates and annotations which are expensive and labor intensive. In this work, we train models to automatically create adversarial prompts to elicit biased responses from target LLMs. We present LLM- based bias evaluation metrics and also analyze several existing automatic evaluation methods and metrics. We analyze the various nuances of model responses, identify the strengths and weaknesses of model families, and assess where evaluation methods fall short. We compare these metrics to human evaluation and validate that the LLM-as-a-Judge metric aligns with human judgement on bias in response generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLMs on Trial: Evaluating Judicial Fairness for Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new 177,100-case benchmark shows that 16 LLMs systematically vary criminal sentences based on extra-legal demographic and procedural details, revealing pervasive judicial unfairness.

  2. Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Gender-diverse users perceive ChatGPT's gender bias differently, with non-binary/transgender participants reporting condescending and stereotypical responses, and men reporting higher trust.

  3. Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

    cs.CY 2026-06 conditional novelty 5.0 of 10

    Across 25,000 stories from five LLMs, an LLM judge rated stories mentioning intellectual disabilities as more infantile, paternalistic, dependent, and inspirational than stories without the label.

  4. A Close Reading Approach to Gender Narrative Biases in AI-Generated Stories

    cs.HC 2025-08 conditional novelty 5.0 of 10

    A close reading of 15 AI-generated stories finds that even when character counts are balanced, narrative roles, descriptions, and plot dynamics remain gender-stereotyped (e.g., every villain is male).

  5. Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives

    cs.CL 2025-06 reject novelty 5.0 of 10

    A multi-hop QA probe of four LLMs claims intersectional mental-health bias and 66-94% debiasing, but the bias metric is undefined and no conventional baseline is tested.

  6. Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead

    cs.SE 2025-06 conditional novelty 4.0 of 10

    A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.

  7. Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement

    cs.IR 2025-06 conditional novelty 4.0 of 10

    LLM-based three-class gender labeling agrees with human annotations better than the lexical NFaiRR score, and the proposed CWEx metric combines neutral exposure with male-female exposure disparity for ranking fairness...

  8. LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models

    cs.CL 2025-05 reject novelty 4.0 of 10

    A block-localizing fine-tuning method for gender debiasing is presented, but its stated loss is inconsistent with its reported behavior and the evaluation tables contain duplicate rows.

  9. Do Biased Models Have Biased Thoughts?

    cs.CL 2025-08 unverdicted novelty 3.0 of 10

    The manuscript is internally inconsistent: the abstract describes an LLM fairness experiment while the body is a different paper on pilot-wave quantum mechanics, so no coherent result can be assessed.

Pith tools