Pith. sign in

REVIEW 19 cited by

NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.12464 v10 pith:PTQS4CSE submitted 2024-04-18 cs.CL

classification cs.CL
keywords culturalllmsadaptabilitysocialframeworkglobalmodelsnorms
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

To be effectively and safely deployed to global user populations, large language models (LLMs) may need to adapt outputs to user values and cultures, not just know about them. We introduce NormAd, an evaluation framework to assess LLMs' cultural adaptability, specifically measuring their ability to judge social acceptability across varying levels of cultural norm specificity, from abstract values to explicit social norms. As an instantiation of our framework, we create NormAd-Eti, a benchmark of 2.6k situational descriptions representing social-etiquette related cultural norms from 75 countries. Through comprehensive experiments on NormAd-Eti, we find that LLMs struggle to accurately judge social acceptability across these varying degrees of cultural contexts and show stronger adaptability to English-centric cultures over those from the Global South. Even in the simplest setting where the relevant social norms are provided, the best LLMs' performance (< 82\%) lags behind humans (> 95\%). In settings with abstract values and country information, model performance drops substantially (< 60\%), while human accuracy remains high (> 90\%). Furthermore, we find that models are better at recognizing socially acceptable versus unacceptable situations. Our findings showcase the current pitfalls in socio-cultural reasoning of LLMs which hinder their adaptability for global audiences.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Portugal's 9B-parameter national language model AMALIA agrees with human annotators on coding moral authority but fails a construct-validity test showing it reaches correct codes via surface correlates rather than the...

  2. PLURAL: A Global Dataset for Value Alignment

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Synthetic preference data generated from the Integrated Values Survey preserves cross-country value differences and enables DPO fine-tuning that improves LLM cultural alignment across five countries.

  3. XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad

    cs.CL 2026-01 conditional novelty 6.0 of 10

    XCR-Bench provides 4,100+ parallel sentences with 1,098 culture-specific items mapped to Hall's Triad, and shows LLMs struggle most with deeper, semi-visible cultural norms.

  4. Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Across nine Asian languages, multilingual LLMs favor Western cultural entities in 30-40% of culturally grounded contexts, with model-specific sentiment biases and extraction accuracy gaps.

  5. CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A taxonomy-guided retrieval-augmented framework generates CultureSynth-7, a multilingual cultural QA benchmark, and its evaluation of 14 LLMs suggests cultural competence emerges around 3B parameters.

  6. EtiCor++: Towards Understanding Etiquettical Bias in LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new English etiquette corpus and bias metrics show that LLMs over-prefer Western norms and under-predict etiquettes from low-resource regions.

  7. CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A multilingual critique-data training paradigm with a knowledge-unit reward improves LLM cultural alignment on several benchmarks, but its headline benchmark is evaluated with the same LLM-judged metric used to select...

  8. Fair-PP: A Synthetic Dataset for Aligning LLM with Personalized Preferences of Social Equity

    cs.AI 2025-05 conditional novelty 6.0 of 10

    Fair-PP contributes a synthetic persona-anchored preference dataset for social equity and a reweighted DPO/SFT alignment method that outperforms baselines on LLM-similarity tests.

  9. PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian

    cs.CL 2025-02 conditional novelty 6.0 of 10

    PerCul is a Persian cultural story-comprehension benchmark on which the best LLMs lag human performance by 11.3 to 21.3 percentage points.

  10. Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness

    cs.CY 2025-02 conditional novelty 6.0 of 10

    The paper argues that LLMs should be evaluated and built for meta-cultural competence rather than static knowledge of specific cultures, and gives a first, illustrative measurement of one component.

  11. When One LLM Drools, Multi-LLM Collaboration Rules

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A position paper that introduces a four-level taxonomy of multi-LLM collaboration (API, text, logit, weight) and argues it is essential for reliability, pluralism, and democratization.

  12. On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Arab cultural entities that double as everyday Arabic words are harder for language models to recognize, especially when tokenized as single tokens.

  13. CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries

    cs.AI 2025-01 conditional novelty 6.0 of 10

    CultureVerse is a 188-country, 19k-concept visual QA benchmark, and fine-tuning open VLMs on it improves cultural accuracy, but the main evaluation shares concepts between training and test sets.

  14. SafeWorld: Geo-Diverse Safety Alignment

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper introduces a geo-diverse cultural and legal safety benchmark and shows that a DPO-trained 7B model can outperform GPT-4o on it, with caveats about the GPT-4-based evaluation loop.

  15. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.

  16. INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

    cs.CL 2024-11 conditional novelty 6.0 of 10

    INCLUDE is a multilingual benchmark of 197,243 exam questions from local sources that evaluates how well LLMs handle regional and cultural knowledge.

  17. Do Large Language Models Know Folktales? A Case Study of Yokai in Japanese Folktales

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A benchmark of 809 yokai questions shows Japanese-centric LLMs, particularly Llama-3-based continual pretraining models, outperform English-centric models on Japanese folktale knowledge.

  18. Musical ethnocentrism in Large Language Models

    cs.CL 2025-01 conditional novelty 4.0 of 10

    LLMs like ChatGPT and Mixtral show a strong Western bias when asked to name top musical contributors and to rate the musical cultures of countries.

  19. A Survey on Human-Centric LLMs

    cs.CL 2024-11 conditional novelty 1.0 of 10

    A review that sorts existing evidence on how well large language models imitate individual human skills and collective social dynamics into one taxonomy.

Pith tools