Pith. sign in

REVIEW 4 cited by

Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.12342 v2 pith:BZO33UNB submitted 2023-08-25 cs.CY cs.CLcs.LG

classification cs.CYcs.CLcs.LG
keywords culturalllmsalignmentmodelsdimensionsexplanatoryresearchanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The deployment of large language models (LLMs) raises concerns regarding their cultural misalignment and potential ramifications on individuals and societies with diverse cultural backgrounds. While the discourse has focused mainly on political and social biases, our research proposes a Cultural Alignment Test (Hoftede's CAT) to quantify cultural alignment using Hofstede's cultural dimension framework, which offers an explanatory cross-cultural comparison through the latent variable analysis. We apply our approach to quantitatively evaluate LLMs, namely Llama 2, GPT-3.5, and GPT-4, against the cultural dimensions of regions like the United States, China, and Arab countries, using different prompting styles and exploring the effects of language-specific fine-tuning on the models' behavioural tendencies and cultural values. Our results quantify the cultural alignment of LLMs and reveal the difference between LLMs in explanatory cultural dimensions. Our study demonstrates that while all LLMs struggle to grasp cultural values, GPT-4 shows a unique capability to adapt to cultural nuances, particularly in Chinese settings. However, it faces challenges with American and Arab cultures. The research also highlights that fine-tuning LLama 2 models with different languages changes their responses to cultural questions, emphasizing the need for culturally diverse development in AI for worldwide acceptance and ethical use. For more details or to contribute to this research, visit our GitHub page https://github.com/reemim/Hofstedes_CAT/

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation

    cs.CL 2025-08 conditional novelty 6.0 of 10

    An evaluation of five VLMs on culturally-prompted multimodal story generation finds measurable cultural adaptation alongside metric bias and inverse alignment in some models.

  2. MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A dynamic multilingual cultural evaluation framework shows that LLM cultural performance depends on both training data distribution and language-culture alignment, and that English-only evaluations hide severe cultura...

  3. Against 'softmaxing' culture

    cs.HC 2025-06 unverdicted novelty 5.0 of 10

    A position paper arguing that AI evaluations should shift from defining culture to understanding when culture becomes relationally valid.

  4. A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities

    cs.CL 2025-07 conditional novelty 2.0 of 10

    A review that categorizes methods for adapting LLMs to discipline-specific research and surveys applications across five broad academic fields.

Pith tools