REVIEW 3 cited by
Unintended Impacts of LLM Alignment on Global Representation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Before being deployed for user-facing applications, developers align Large Language Models (LLMs) to user preferences through a variety of procedures, such as Reinforcement Learning From Human Feedback (RLHF) and Direct Preference Optimization (DPO). Current evaluations of these procedures focus on benchmarks of instruction following, reasoning, and truthfulness. However, human preferences are not universal, and aligning to specific preference sets may have unintended effects. We explore how alignment impacts performance along three axes of global representation: English dialects, multilingualism, and opinions from and about countries worldwide. Our results show that current alignment procedures create disparities between English dialects and global opinions. We find alignment improves capabilities in several languages. We conclude by discussing design decisions that led to these unintended impacts and recommendations for more equitable preference tuning. We make our code and data publicly available on Github.
Forward citations
Cited by 3 Pith papers
-
CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries
Leading LLMs produce culturally appropriate health responses only 20-30% of the time on a new continuum-of-norm-adherence benchmark, with a strong bias toward Western defaults.
-
People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe
Across 10 LLMs and the European Social Survey, AI alignment favors wealthier, more educated, less religious, and more politically interested groups, with country of residence explaining as much variance as all sociode...
-
Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
A pilot probe finds weak, non-robust evidence that Qwen2.5-7B internally represents Colombian identity from a single implicit cue; the only nominally significant effect is driven by unrestricted, confabulated national...
Discussion (0). Continue with ORCID to comment.