Pith. sign in

REVIEW 7 cited by

Investigating Cultural Alignment of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13231 v2 pith:OGWGCAYV submitted 2024-02-20 cs.CL cs.CY

classification cs.CLcs.CY
keywords alignmentculturallanguagemodelsculturedifferentsurveyanthropological
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The intricate relationship between language and culture has long been a subject of exploration within the realm of linguistic anthropology. Large Language Models (LLMs), promoted as repositories of collective human knowledge, raise a pivotal question: do these models genuinely encapsulate the diverse knowledge adopted by different cultures? Our study reveals that these models demonstrate greater cultural alignment along two dimensions -- firstly, when prompted with the dominant language of a specific culture, and secondly, when pretrained with a refined mixture of languages employed by that culture. We quantify cultural alignment by simulating sociological surveys, comparing model responses to those of actual survey participants as references. Specifically, we replicate a survey conducted in various regions of Egypt and the United States through prompting LLMs with different pretraining data mixtures in both Arabic and English with the personas of the real respondents and the survey questions. Further analysis reveals that misalignment becomes more pronounced for underrepresented personas and for culturally sensitive topics, such as those probing social values. Finally, we introduce Anthropological Prompting, a novel method leveraging anthropological reasoning to enhance cultural alignment. Our study emphasizes the necessity for a more balanced multilingual pretraining dataset to better represent the diversity of human experience and the plurality of different cultures with many implications on the topic of cross-lingual transfer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LKValues: Aligning Large Language Models with Sri Lankan Societal Values

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A survey-derived Sri Lankan value alignment suite (LKValues) with 150k instruction instances and a 1k benchmark improves Qwen-family LLMs' Sri Lankan value judgment in Sinhala and English.

  2. The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs

    cs.CL 2025-10 reject novelty 6.0 of 10

    The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—...

  3. Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation

    cs.CL 2025-08 conditional novelty 6.0 of 10

    An evaluation of five VLMs on culturally-prompted multimodal story generation finds measurable cultural adaptation alongside metric bias and inverse alignment in some models.

  4. Prompt Programming for Cultural Bias and Alignment of Large Language Models

    cs.AI 2026-03 conditional novelty 5.0 of 10

    Automatically optimized prompts (DSPy) reduce survey-measured cultural distance for open-weight LLMs more often than manual cultural prompting, with MIPROv2 and a large proposer model giving the most consistent gains.

  5. Speaking images. A novel framework for the automated self-description of artworks

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A four-stage open-source AI pipeline turns a digitized artwork into a short video where a depicted person animates and narrates the scene.

  6. Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise

    cs.AI 2025-05 reject novelty 4.0 of 10

    A multi-agent router that selects culturally specialized LLM personas reports a jump in self-scored cultural alignment from 0.208 to 0.820, but the metric and the claimed method are not independently validated.

  7. LLM Web Dynamics: Tracing Model Collapse in a Network of LLMs

    cs.LG 2025-05 conditional novelty 4.0 of 10

    Under a shared retrieval-augmented memory, multiple LLMs' outputs converge to near-identical semantic answers, and the analogous Gaussian mixture system is proven to collapse.

Pith tools