Pith. sign in

REVIEW 8 cited by

Having Beer after Prayer? Measuring Cultural Bias in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14456 v4 pith:BCDXGY6X submitted 2023-05-23 cs.CL cs.AIcs.LG

Having Beer after Prayer? Measuring Cultural Bias in Large Language Models

classification cs.CL cs.AIcs.LG
keywords culturalcamelarabicmodelsappropriatearabbiascontexts
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

As the reach of large language models (LMs) expands globally, their ability to cater to diverse cultural contexts becomes crucial. Despite advancements in multilingual capabilities, models are not designed with appropriate cultural nuances. In this paper, we show that multilingual and Arabic monolingual LMs exhibit bias towards entities associated with Western culture. We introduce CAMeL, a novel resource of 628 naturally-occurring prompts and 20,368 entities spanning eight types that contrast Arab and Western cultures. CAMeL provides a foundation for measuring cultural biases in LMs through both extrinsic and intrinsic evaluations. Using CAMeL, we examine the cross-cultural performance in Arabic of 16 different LMs on tasks such as story generation, NER, and sentiment analysis, where we find concerning cases of stereotyping and cultural unfairness. We further test their text-infilling performance, revealing the incapability of appropriate adaptation to Arab cultural contexts. Finally, we analyze 6 Arabic pre-training corpora and find that commonly used sources such as Wikipedia may not be best suited to build culturally aware LMs, if used as they are without adjustment. We will make CAMeL publicly available at: https://github.com/tareknaous/camel

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MultiLinguahah : A New Unsupervised Multilingual Acoustic Laughter Segmentation Method

    cs.CL 2026-05 unverdicted novelty 6.0

    An unsupervised multilingual laughter segmentation method using Isolation Forest on BYOL-A audio representations outperforms existing supervised methods on non-English datasets.

  2. The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs

    cs.CL 2025-10 reject novelty 6.0

    The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—...

  3. Understanding and Debugging Failures in N-Gram-Based Generative Retrieval

    cs.IR 2026-06 unverdicted novelty 5.0

    Presents a taxonomy of generative retrieval failures, empirically identifies issues such as ambiguous docids and low diversity in n-gram methods, and introduces a web-based debugging tool.

  4. MultiLinguahah : A New Unsupervised Multilingual Acoustic Laughter Segmentation Method

    cs.CL 2026-05 unverdicted novelty 5.0

    An unsupervised multilingual laughter segmentation technique using Isolation Forest on BYOL-A representations outperforms state-of-the-art supervised detectors on non-English audio datasets.

  5. Representational Harms in LLM-Generated Narratives Against Global Majority Nationalities

    cs.CL 2026-04 unverdicted novelty 5.0

    LLMs generate narratives containing persistent stereotypes, erasure, and one-dimensional portrayals of Global Majority national identities, with minoritized groups overrepresented in subordinated roles by more than fi...

  6. Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?

    cs.AI 2025-03 conditional novelty 5.0

    LLMs show accuracy drops of 0.3% to 5.9% on GSM8K math problems when culturally adapted to six countries while keeping math operations identical, with statistical significance confirmed by McNemar tests.

  7. Attributing Culture-Conditioned Generations to Pretraining Corpora

    cs.CL 2024-12 unverdicted novelty 5.0

    MEMOed framework attributes LLM generations about cultures to pretraining memorization and finds frequency-based biases across 110 cultures for food and clothing.

  8. SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures

    cs.CL 2026-05 unverdicted novelty 4.0

    SemEval-2026 Task 7 presents a benchmark and two evaluation tracks for assessing LLMs on everyday knowledge in diverse languages and cultures without allowing training on the test data.