Pith. sign in

REVIEW 2 cited by

Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16523 v1 pith:LTJ5XP4U submitted 2023-10-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords diversityllmsresponseslargemodelspeopleculturedemographic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A crucial challenge for generative large language models (LLMs) is diversity: when a user's prompt is under-specified, models may follow implicit assumptions while generating a response, which may result in homogenization of the responses, as well as certain demographic groups being under-represented or even erased from the generated responses. In this paper, we formalize diversity of representation in generative LLMs. We present evaluation datasets and propose metrics to measure diversity in generated responses along people and culture axes. We find that LLMs understand the notion of diversity, and that they can reason and critique their own responses for that goal. This finding motivated a new prompting technique called collective-critique and self-voting (CCSV) to self-improve people diversity of LLMs by tapping into its diversity reasoning capabilities, without relying on handcrafted examples or prompt tuning. Extensive empirical experiments with both human and automated evaluations show that our proposed approach is effective at improving people and culture diversity, and outperforms all baseline methods by a large margin.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis

    cs.CR 2025-10 conditional novelty 6.0 of 10

    Angry natural phrases selected by PCA and k-means medoids act as generalizable backdoor triggers in RLHF, improving attack success on unseen phrasings.

  2. AI in Mental Health: Emotional and Sentiment Analysis of Large Language Models' Responses to Depression, Anxiety, and Stress Queries

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Eight LLMs show measurably different emotional tones in mental-health answers: anxiety prompts produced near-saturated fear scores, depression prompts the most sadness, and stress prompts the most optimism.

Pith tools