Pith. sign in

REVIEW 1 cited by

Towards Controllable Biases in Language Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.00268 v2 pith:NNLSJGVL submitted 2020-05-01 cs.CL

classification cs.CL
keywords biasesdemographicdemographicsgenerationtextapproachcontrollableeffectiveness
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a general approach towards controllable societal biases in natural language generation (NLG). Building upon the idea of adversarial triggers, we develop a method to induce societal biases in generated text when input prompts contain mentions of specific demographic groups. We then analyze two scenarios: 1) inducing negative biases for one demographic and positive biases for another demographic, and 2) equalizing biases between demographics. The former scenario enables us to detect the types of biases present in the model. Specifically, we show the effectiveness of our approach at facilitating bias analysis by finding topics that correspond to demographic inequalities in generated text and comparing the relative effectiveness of inducing biases for different demographics. The second scenario is useful for mitigating biases in downstream applications such as dialogue generation. In our experiments, the mitigation technique proves to be effective at equalizing the amount of biases across demographics while simultaneously generating less negatively biased text overall.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bias Unveiled: Investigating Social Bias in LLM-Generated Code

    cs.SE 2024-11 conditional novelty 6.0 of 10

    An evaluation framework and 343-task benchmark show that four code LLMs produce socially biased code, and iterative bias feedback reduces measured bias substantially.

Pith tools