Pith. sign in

REVIEW 3 cited by

CORGI-PM: A Chinese Corpus For Gender Bias Probing and Mitigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.00395 v1 pith:NFQKXUSI submitted 2023-01-01 cs.CL cs.AIcs.CYcs.LG

classification cs.CLcs.AIcs.CYcs.LG
keywords biasgenderchinesecorpusmitigationcorgi-pmlanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As natural language processing (NLP) for gender bias becomes a significant interdisciplinary topic, the prevalent data-driven techniques such as large-scale language models suffer from data inadequacy and biased corpus, especially for languages with insufficient resources such as Chinese. To this end, we propose a Chinese cOrpus foR Gender bIas Probing and Mitigation CORGI-PM, which contains 32.9k sentences with high-quality labels derived by following an annotation scheme specifically developed for gender bias in the Chinese context. Moreover, we address three challenges for automatic textual gender bias mitigation, which requires the models to detect, classify, and mitigate textual gender bias. We also conduct experiments with state-of-the-art language models to provide baselines. To our best knowledge, CORGI-PM is the first sentence-level Chinese corpus for gender bias probing and mitigation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

    cs.CL 2025-10 conditional novelty 5.0 of 10

    Chinese LLMs generate more positive continuations after 'we' prompts and more negative after 'they' prompts; the feminine 'they' intensifies negativity in several pretrained models.

  2. Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A Chinese gender-bias corpus and shared-task benchmark show detection and classification are feasible, while automatic mitigation remains weak.

  3. Detection, Classification, and Mitigation of Gender Bias in Large Language Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A Chinese gender-bias system using SFT, chain-of-thought, and DPO with GPT-4-generated preference pairs reports top validation scores and first place on all three NLPCC 2025 subtasks.

Pith tools