Pith. sign in

REVIEW 2 cited by

The power of Prompts: Evaluating and Mitigating Gender Bias in MT with LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.18786 v1 pith:QGTPJITM submitted 2024-07-26 cs.CL

classification cs.CL
keywords biasgenderllmsmodelstranslationbasecomparedenglish
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This paper studies gender bias in machine translation through the lens of Large Language Models (LLMs). Four widely-used test sets are employed to benchmark various base LLMs, comparing their translation quality and gender bias against state-of-the-art Neural Machine Translation (NMT) models for English to Catalan (En $\rightarrow$ Ca) and English to Spanish (En $\rightarrow$ Es) translation directions. Our findings reveal pervasive gender bias across all models, with base LLMs exhibiting a higher degree of bias compared to NMT models. To combat this bias, we explore prompting engineering techniques applied to an instruction-tuned LLM. We identify a prompt structure that significantly reduces gender bias by up to 12% on the WinoMT evaluation dataset compared to more straightforward prompts. These results significantly reduce the gender bias accuracy gap between LLMs and traditional NMT systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Mitigating Prompt-Induced Cognitive Biases in General-Purpose AI for Software Engineering

    cs.SE 2026-04 unverdicted novelty 6.0 of 10

    A prompting method that forces GPAI models to state SE best practices before deciding reduces prompt-induced cognitive biases by 51% on average across eight tested biases.

  2. Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation

    cs.CL 2025-09 conditional novelty 6.0 of 10

    LLMs show significant demographic bias in decision-making, favoring women, younger ages, and certain minority backgrounds; summarization shows little bias, and bias patterns largely transfer from English to Dutch.

Pith tools