Pith. sign in

REVIEW 3 cited by

Revealing Persona Biases in Dialogue Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.08728 v2 pith:BQFMUE7I submitted 2021-04-18 cs.CL

classification cs.CL
keywords dialoguepersonapersonassystemsbiasesadoptingdifferentharmful
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Dialogue systems in the form of chatbots and personal assistants are being increasingly integrated into people's lives. Modern dialogue systems may consider adopting anthropomorphic personas, mimicking societal demographic groups to appear more approachable and trustworthy to users. However, the adoption of a persona can result in the adoption of biases. In this paper, we present the first large-scale study on persona biases in dialogue systems and conduct analyses on personas of different social classes, sexual orientations, races, and genders. We define persona biases as harmful differences in responses (e.g., varying levels of offensiveness, agreement with harmful statements) generated from adopting different demographic personas. Furthermore, we introduce an open-source framework, UnitPersonaBias, to explore and aggregate persona biases in dialogue systems. By analyzing the Blender and DialoGPT dialogue systems, we observe that adopting personas can actually decrease harmful responses, compared to not using any personas. Additionally, we find that persona choices can affect the degree of harms in generated responses and thus should be systematically evaluated before deployment. We also analyze how personas can result in different amounts of harm towards specific demographics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Localizing Persona Representations in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.

  2. Adaptive Prompting: Ad-hoc Prompt Composition for Social Bias Detection

    cs.CL 2025-02 conditional novelty 6.0 of 10

    On StereoSet and SBIC, a trained DeBERTa encoder that selects per-input prompt compositions from 64 options raises macro F1 above every fixed composition, but on CobraFrames it falls below the best fixed composition.

  3. Exploring the Impact of Occupational Personas on Domain-Specific QA

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Profession-based personas slightly improve LLM accuracy on science QA, while occupational personality personas often reduce it, even when semantically related.

Pith tools