REVIEW 14 cited by
Assessing Cross-Cultural Alignment between ChatGPT and Human Societies: An Empirical Study
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The recent release of ChatGPT has garnered widespread recognition for its exceptional ability to generate human-like responses in dialogue. Given its usage by users from various nations and its training on a vast multilingual corpus that incorporates diverse cultural and societal norms, it is crucial to evaluate its effectiveness in cultural adaptation. In this paper, we investigate the underlying cultural background of ChatGPT by analyzing its responses to questions designed to quantify human cultural differences. Our findings suggest that, when prompted with American context, ChatGPT exhibits a strong alignment with American culture, but it adapts less effectively to other cultural contexts. Furthermore, by using different prompts to probe the model, we show that English prompts reduce the variance in model responses, flattening out cultural differences and biasing them towards American culture. This study provides valuable insights into the cultural implications of ChatGPT and highlights the necessity of greater diversity and cultural awareness in language technologies.
Forward citations
Cited by 14 Pith papers
-
Toward a Theory of Value in AI Alignment
A systematic annotation of 94 AI alignment papers shows the field largely equates human values with measurable preferences, rarely defines values, and is increasingly removing humans from alignment evaluation.
-
Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation
An evaluation of five VLMs on culturally-prompted multimodal story generation finds measurable cultural adaptation alongside metric bias and inverse alignment in some models.
-
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
Value representations from token logits, sequence perplexity, and text generation are all sensitive to prompt and option changes, and their correlation with model behavior in value scenarios is weak.
-
Autonomy by Design: Preserving Human Autonomy in AI Decision-Support
AI decision support can erode domain-specific autonomy through opaque failures and unconscious value shifts, but socio-technical design patterns such as defeaters and positive friction can mitigate these risks.
-
A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs
A multilingual two-phase evaluation shows LLMs lean on query language for factual questions and on training-country perspective for territorial and historical disputes.
-
Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models
Across six country-pair comparisons, GPT-4o-mini and GigaChat-Max side with US positions 64-81% of the time, Qwen2.5 and Llama-4 lean neutral more often, and a debias prompt shifts these numbers by only a few points.
-
Is It Bad to Work All the Time? Cross-Cultural Evaluation of Social Norm Biases in GPT-4
GPT-4 writes accurate but generic social norms for non-US cultures, defaults to US-style judgments, and still holds recoverable stereotypes about China, India, and Iran.
-
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.
-
LaQual: An Automated Framework for LLM App Quality Evaluation
LaQual automates LLM app-store quality evaluation through scenario classification, static indicator filtering, and LLM-generated dynamic metrics, with Spearman correlations of about 0.6 against human ratings.
-
A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies
Across U.S. and Chinese survey questions, DeepSeek, GPT-4o, Qwen2.5, and Llama-3.3 all show demographic overgeneralization, with no consistent home-field advantage for the Chinese model.
-
Normative Evaluation of Large Language Models with Everyday Moral Dilemmas
Seven LLMs give different moral verdicts on AITA dilemmas, differ from Redditors, and only in an ensemble approximate human consensus.
-
AI and Cultural Context: An Empirical Investigation of Large Language Models' Performance on Chinese Social Work Professional Standards
Most Chinese and Western LLMs pass the Chinese social work exam's knowledge sections, with Chinese models ahead on jurisprudence, and both groups showing cultural bias in practice scenarios.
-
One world, one opinion? The superstar effect in LLM responses
Across ten languages, LLMs consistently name a small set of figures such as Einstein, Shakespeare, and Turing for each profession, revealing a 'superstar effect' that may narrow cultural representation.
-
Building Bridges across Papua New Guinea's Digital Divide in Growing the ICT Industry
A Papua New Guinea-focused workshop report identifies open source software, education reform, and local context as key levers for closing the digital divide, and lists twelve research questions.
Discussion (0). Continue with ORCID to comment.