Pith. sign in

REVIEW 2 cited by

Evaluating LLMs for Gender Disparities in Notable Persons

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.09148 v1 pith:MH4KACPO submitted 2024-03-14 cs.CL cs.IR

classification cs.CLcs.IR
keywords responsesdisparitiesgenderevaluatingllmsfactualmodelsacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study examines the use of Large Language Models (LLMs) for retrieving factual information, addressing concerns over their propensity to produce factually incorrect "hallucinated" responses or to altogether decline to even answer prompt at all. Specifically, it investigates the presence of gender-based biases in LLMs' responses to factual inquiries. This paper takes a multi-pronged approach to evaluating GPT models by evaluating fairness across multiple dimensions of recall, hallucinations and declinations. Our findings reveal discernible gender disparities in the responses generated by GPT-3.5. While advancements in GPT-4 have led to improvements in performance, they have not fully eradicated these gender disparities, notably in instances where responses are declined. The study further explores the origins of these disparities by examining the influence of gender associations in prompts and the homogeneity in the responses.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A name-substituted variant of the BBQ benchmark shows that LLMs retain nationality stereotypes even when explicit labels are removed, with smaller models showing more bias and lower accuracy.

  2. LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models

    cs.CL 2025-05 reject novelty 4.0 of 10

    A block-localizing fine-tuning method for gender debiasing is presented, but its stated loss is inconsistent with its reported behavior and the evaluation tables contain duplicate rows.

Pith tools