Pith. sign in

REVIEW 7 cited by

Assessing Cross-Cultural Alignment between ChatGPT and Human Societies: An Empirical Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.17466 v2 pith:DHRBTPFZ submitted 2023-03-30 cs.CL

classification cs.CL
keywords culturalchatgptamericanresponsesalignmentculturedifferenceshuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent release of ChatGPT has garnered widespread recognition for its exceptional ability to generate human-like responses in dialogue. Given its usage by users from various nations and its training on a vast multilingual corpus that incorporates diverse cultural and societal norms, it is crucial to evaluate its effectiveness in cultural adaptation. In this paper, we investigate the underlying cultural background of ChatGPT by analyzing its responses to questions designed to quantify human cultural differences. Our findings suggest that, when prompted with American context, ChatGPT exhibits a strong alignment with American culture, but it adapts less effectively to other cultural contexts. Furthermore, by using different prompts to probe the model, we show that English prompts reduce the variance in model responses, flattening out cultural differences and biasing them towards American culture. This study provides valuable insights into the cultural implications of ChatGPT and highlights the necessity of greater diversity and cultural awareness in language technologies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 23 citations worldwide. Full citation record

  1. Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation

    cs.CL 2025-08 conditional novelty 6.0 of 10

    An evaluation of five VLMs on culturally-prompted multimodal story generation finds measurable cultural adaptation alongside metric bias and inverse alignment in some models.

  2. Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Value representations from token logits, sequence perplexity, and text generation are all sensitive to prompt and option changes, and their correlation with model behavior in value scenarios is weak.

  3. Autonomy by Design: Preserving Human Autonomy in AI Decision-Support

    cs.HC 2025-06 accept novelty 6.0 of 10

    AI decision support can erode domain-specific autonomy through opaque failures and unconscious value shifts, but socio-technical design patterns such as defeaters and positive friction can mitigate these risks.

  4. A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A multilingual two-phase evaluation shows LLMs lean on query language for factual questions and on training-country perspective for territorial and historical disputes.

  5. Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Across six country-pair comparisons, GPT-4o-mini and GigaChat-Max side with US positions 64-81% of the time, Qwen2.5 and Llama-4 lean neutral more often, and a debias prompt shifts these numbers by only a few points.

  6. LaQual: An Automated Framework for LLM App Quality Evaluation

    cs.SE 2025-08 reject novelty 5.0 of 10

    LaQual automates LLM app-store quality evaluation through scenario classification, static indicator filtering, and LLM-generated dynamic metrics, with Spearman correlations of about 0.6 against human ratings.

  7. A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Across U.S. and Chinese survey questions, DeepSeek, GPT-4o, Qwen2.5, and Llama-3.3 all show demographic overgeneralization, with no consistent home-field advantage for the Chinese model.

Pith tools