Pith. sign in

REVIEW 1 cited by

Voices in a Crowd: Searching for Clusters of Unique Perspectives

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.14259 v1 pith:W6DPKR4X submitted 2024-07-19 cs.CL cs.LG

classification cs.CLcs.LG
keywords clustersannotatorperspectivesframeworkmetadataminoritymodelsresulting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models have been shown to reproduce underlying biases existing in their training data, which is the majority perspective by default. Proposed solutions aim to capture minority perspectives by either modelling annotator disagreements or grouping annotators based on shared metadata, both of which face significant challenges. We propose a framework that trains models without encoding annotator metadata, extracts latent embeddings informed by annotator behaviour, and creates clusters of similar opinions, that we refer to as voices. Resulting clusters are validated post-hoc via internal and external quantitative metrics, as well a qualitative analysis to identify the type of voice that each cluster represents. Our results demonstrate the strong generalisation capability of our framework, indicated by resulting clusters being adequately robust, while also capturing minority perspectives based on different demographic factors throughout two distinct datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Explaining Matters: Leveraging Definitions and Semantic Expansion for Sexism Detection

    cs.CL 2025-06 conditional novelty 6.0 of 10

    On the EDOS benchmark, definition-based augmentation and context expansion with a Mistral-7B tie-breaker reach macro F1 0.8819 (binary) and 0.6018 (fine-grained).

Pith tools