Pith. sign in

REVIEW 6 cited by

Unintended Impacts of LLM Alignment on Global Representation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15018 v2 pith:7TZ3NQL5 submitted 2024-02-22 cs.CL cs.CYcs.LG

Unintended Impacts of LLM Alignment on Global Representation

classification cs.CL cs.CYcs.LG
keywords alignmentglobalimpactspreferenceproceduresunintendedcurrentdialects
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Before being deployed for user-facing applications, developers align Large Language Models (LLMs) to user preferences through a variety of procedures, such as Reinforcement Learning From Human Feedback (RLHF) and Direct Preference Optimization (DPO). Current evaluations of these procedures focus on benchmarks of instruction following, reasoning, and truthfulness. However, human preferences are not universal, and aligning to specific preference sets may have unintended effects. We explore how alignment impacts performance along three axes of global representation: English dialects, multilingualism, and opinions from and about countries worldwide. Our results show that current alignment procedures create disparities between English dialects and global opinions. We find alignment improves capabilities in several languages. We conclude by discussing design decisions that led to these unintended impacts and recommendations for more equitable preference tuning. We make our code and data publicly available on Github.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries

    cs.CY 2026-06 conditional novelty 7.0

    Leading LLMs produce culturally appropriate health responses only 20-30% of the time on a new continuum-of-norm-adherence benchmark, with a strong bias toward Western defaults.

  2. AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction

    cs.CL 2023-05 unverdicted novelty 6.0

    LLM embeddings enable strong retrodiction of masked GSS opinions via cross-validation and external validation but only modest performance on entirely unasked opinions.

  3. Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders

    cs.CL 2026-07 reject novelty 5.0

    A pilot probe finds weak, non-robust evidence that Qwen2.5-7B internally represents Colombian identity from a single implicit cue; the only nominally significant effect is driven by unrestricted, confabulated national...

  4. Representational Harms in LLM-Generated Narratives Against Global Majority Nationalities

    cs.CL 2026-04 unverdicted novelty 5.0

    LLMs generate narratives containing persistent stereotypes, erasure, and one-dimensional portrayals of Global Majority national identities, with minoritized groups overrepresented in subordinated roles by more than fi...

  5. Attributing Culture-Conditioned Generations to Pretraining Corpora

    cs.CL 2024-12 unverdicted novelty 5.0

    MEMOed framework attributes LLM generations about cultures to pretraining memorization and finds frequency-based biases across 110 cultures for food and clothing.

  6. SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures

    cs.CL 2026-05 unverdicted novelty 4.0

    SemEval-2026 Task 7 presents a benchmark and two evaluation tracks for assessing LLMs on everyday knowledge in diverse languages and cultures without allowing training on the test data.