Pith. sign in

REVIEW 4 cited by

Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.15585 v1 pith:EVK6RRMH submitted 2024-01-28 cs.CL

Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting

classification cs.CL
keywords llmspredictionsreasoningtasksunscalablewordsmodelbias
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

There exist both scalable tasks, like reading comprehension and fact-checking, where model performance improves with model size, and unscalable tasks, like arithmetic reasoning and symbolic reasoning, where model performance does not necessarily improve with model size. Large language models (LLMs) equipped with Chain-of-Thought (CoT) prompting are able to make accurate incremental predictions even on unscalable tasks. Unfortunately, despite their exceptional reasoning abilities, LLMs tend to internalize and reproduce discriminatory societal biases. Whether CoT can provide discriminatory or egalitarian rationalizations for the implicit information in unscalable tasks remains an open question. In this study, we examine the impact of LLMs' step-by-step predictions on gender bias in unscalable tasks. For this purpose, we construct a benchmark for an unscalable task where the LLM is given a list of words comprising feminine, masculine, and gendered occupational words, and is required to count the number of feminine and masculine words. In our CoT prompts, we require the LLM to explicitly indicate whether each word in the word list is a feminine or masculine before making the final predictions. With counting and handling the meaning of words, this benchmark has characteristics of both arithmetic reasoning and symbolic reasoning. Experimental results in English show that without step-by-step prediction, most LLMs make socially biased predictions, despite the task being as simple as counting words. Interestingly, CoT prompting reduces this unconscious social bias in LLMs and encourages fair predictions.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

    cs.AI 2026-05 unverdicted novelty 6.0

    Causality provides a unifying framework for resolving trade-offs in trustworthy AI by managing invariance conflicts under changes to the data-generating process.

  2. Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation

    cs.CL 2025-10 unverdicted novelty 5.0

    Reasoning LLMs aggregate social biases through stereotype repetition and irrelevant information injection in their thinking processes, and a self-review prompt mitigates this on BBQ, StereoSet, and BOLD benchmarks.

  3. Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

    cs.AI 2026-05 unverdicted novelty 4.0

    Causality resolves trade-offs in trustworthy AI by treating them as invariance conflicts under different data-generating process changes.

  4. SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures

    cs.CL 2026-05 unverdicted novelty 4.0

    SemEval-2026 Task 7 presents a benchmark and two evaluation tracks for assessing LLMs on everyday knowledge in diverse languages and cultures without allowing training on the test data.