REVIEW 3 major objections 5 minor 13 references
Is It Bad to Work All the Time? Cross-Cultural Evaluation of Social Norm Biases in GPT-4
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that GPT-4's default, culturally unmarked social-norm judgments align most closely with the United States, that its generated norms are accurate but generic, and that its cultural stereotypes are hidden rather than…
desk verdict A genuinely better bottom-up protocol for probing cultural norms in LLMs, with a solid generic-norm finding, but the stereotype-recovery headline rests on cherry-picked examples and the US-alignment JSD is a within-model comparison that the paper overstates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument rests on the Rules of Thumb (RoT) format—short statements of the form 'it is [judgment] to [action]'—applied to 80 English Wikipedia movie plots, 20 per country, as culturally grounded narratives. Norms come from three sources: human annotators from each country, GPT-4 with default prompting, and GPT-4 with cultural prompting ('As someone with a [country X] cultural background...'). Cultural alignment is measured by sampling 30 ratings per norm at temperature 0.8 and computing the Jensen-Shannon divergence between the default rating distribution and each country-prompted distribution. The stereotype claim is carried by reversing the generation task into an ordinal prediction task: the model rates how people from the target country would judge human-rated stereotypical norms, surfacing associations that its text-generation guardrails suppress.
What would settle it
Re-run the rating task with country cues randomly swapped (for example, asking for the Chinese perspective on US norms and vice versa), or compute the Jensen-Shannon divergence between every pair of country-prompted distributions. If the default distribution tracks whichever country is named, or if the pairwise country distances do not reproduce the same ordering, the US-default claim is a prompt artifact rather than a latent cultural prior; near-zero pairwise distances would mean the prompts all restate the same default.
Extended reading notes
Core claim
GPT-4 reasons about social norms in culturally grounded narratives from a default stance that is closer to the United States than to India, Iran, or China, and its avoidance of stereotypes in generated text masks stereotypes that remain in its judgments. The model's default rating distribution over a norm differs from its US-prompted distribution by an average Jensen-Shannon divergence of 0.12, versus 0.20 for Iran, 0.24 for India, and 0.33 for China. The norms it generates are rated about as accurate as human-written ones but less culture-specific, and prompting it 'as someone with a Chinese cultural background' does not significantly improve its generation. When the task is reversed from generating norms to predicting agreement, the model assigns distributions that match the stereotype—for example, predicting that Chinese raters would disagree that it is bad to work all of the time—so the stereotypes are hidden rather than removed.
Load-bearing premise
The load-bearing premise is that prompting GPT-4 to rate norms 'as someone with a [country X] cultural background' genuinely invokes that country's perspective, so the divergence ordering in Figure 4 measures cultural alignment rather than the model's reaction to added wording; the paper itself reports that cultural prompting gave no significant advantage in norm generation, which leaves this premise open.
Editorial extensions
If this is right
- Any use of GPT-4 that consults an unmarked default judgment—advice, moderation, decision support—will carry a US-aligned normative stance even when no cultural reference appears in the prompt.
- Accuracy ratings of generated norms cannot be read as cultural competence, because the model achieves accuracy partly by producing generic norms that people across cultures can agree with.
- A model that avoids voicing stereotypes can still encode them in its predictions, so safety evaluations limited to generated text will miss persistent biases.
- Explicitly naming a country in the prompt does not by itself give the model that country's perspective in norm generation, which limits naive cultural prompting as a mitigation.
- There is a trade-off the paper calls a key tension: the knowledge needed to be culture-specific is entangled with stereotypical associations, so improving cultural specificity may amplify stereotypicality.
Reading between the lines
- If the divergence ordering reflects a latent cultural prior rather than prompt wording, the same movie-plot protocol could serve as a survey-free benchmark for ranking any model by whose norms it defaults to.
- The stereotype-recovery probe—predicting agreement from a cultural persona over human-rated stereotypical norms—is a cheap adversarial test that model developers could run before deployment; it would be worth seeing whether newer models or stronger refusal tuning change the recovered distributions.
- Because the study is deliberately English-only and relies on English Wikipedia summaries, part of the US-alignment gap could be a language-pipeline artifact; running the same task with source-language plots or translated prompts is the natural replication that would test this.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a bottom-up approach to evaluating the cultural alignment of GPT-4: rather than asking the model to answer value surveys, the authors present movie plots from the US, China, India, and Iran, ask human annotators and GPT-4 to write social-norm rules-of-thumb (RoTs), and then have human annotators rate the resulting RoTs for accuracy, culture-specificity, and stereotypicality. They report that GPT-4's RoTs are as accurate as human-written ones but significantly more generic, that GPT-4's default rating distributions are closer to its own US-prompted distributions than to its China-, India-, or Iran-prompted distributions, and that although GPT-4 avoids generating overt stereotypes, stereotypical associations can be recovered by asking the model to predict how people from a target culture would rate stereotypical RoTs. The central claims are that GPT-4's default normative representation is US-aligned and that its stereotype avoidance is superficial.
Significance. If confirmed, the findings would be a valuable contribution to cross-cultural LLM evaluation: the paper introduces a narrative-grounded, bottom-up protocol that goes beyond direct survey questions, and it collects original human annotations from four cultural groups. The distinction between overt stereotype avoidance and recoverable latent stereotypes is an important and under-explored phenomenon. The paper also ships its data/code (footnote 1), which supports reproducibility. However, the strength of the contribution depends on two load-bearing empirical choices that are not fully supported: the US-alignment claim is measured by a self-comparison of GPT-4's distributions rather than against human ground truth, and the stereotype-recovery claim is demonstrated only with selected examples rather than an aggregate analysis. These issues are fixable within the manuscript's scope.
major comments (3)
- [§4.2, Figure 4] The JSD-based alignment measure compares GPT-4's default rating distribution with its own rating distribution under a country-prompted condition. This is a self-similarity measure, not a comparison to any external cultural ground truth. The paper already collects human rating distributions in §4.1 for RoTs, but §4.2 does not validate that the culturally prompted GPT-4 distributions are actually closer to human raters from those countries than the default distribution is. Without that validation, the ordering US 0.12, Iran 0.20, India 0.24, China 0.33 can be interpreted as the model's sensitivity to country labels in the prompt, rather than as evidence that GPT-4's default stance is US-aligned. Please add an analysis that compares the culturally prompted distributions to the human-annotator distributions for the same RoTs (e.g., JSD or rank correlation), and report the origin of the 20 sampled RoTs per country (human-written, GPT-generated, or both), since the composition could systematically affect the result.
- [§4.2 vs. Appendix B.2, Table 7] The task description in §4.2 says the model is asked to rate RoTs "as someone with a [country X] cultural background," but the actual prompt in Table 7 asks the model to "estimate how people with a X cultural background would rate" the statement. These are different tasks: the first elicits the model's own culturally situated judgment, while the second elicits a prediction about a demographic group's opinions. The alignment claim depends on which task is actually performed, because predicting how others rate can activate statistical stereotypes about the group instead of reflecting the model's own cultural perspective. Please reconcile the description with the prompt, and discuss the third-person framing as a possible confound.
- [§4.3, Figure 5] The headline claim that "stereotypes are merely hidden rather than suppressed" and can be "easily recovered" is supported only by four hand-selected example RoTs from the 338-item set of highly stereotypical norms. No aggregate statistic is reported for this set, such as the fraction of RoTs in which the culturally prompted distribution shifts stereotypically relative to default, or the proportion of cases where GPT-4's culturally prompted modal rating matches the human modal rating. As written, the recovery claim is anecdotal rather than demonstrated. Please add a quantitative summary over the 338 RoTs, together with an appropriate statistical test (e.g., comparing default vs. culturally prompted rating distributions on stereotype-relevant items, or comparing predicted vs. human modal ratings).
minor comments (5)
- [§4.1, Figure 3A] The sentence reporting US accuracy says "the difference is not statistically significant (β = −0.31, p = .003)", but p = .003 conventionally indicates significance; please check the coefficient, standard error, and intended comparison, and report the correct regression output.
- [Title, §3.3] The title and most of the text say "GPT-4," but §3.3 specifies "GPT-4o (OpenAI, 2024)." Please be consistent about the exact model version throughout, as results may not transfer across model versions.
- [§4.1, cultural prompting result] The claim that cultural prompting showed no significant advantage is reported only qualitatively; please report the regression coefficient and p-value for the cultural-prompting covariate in the accuracy and specificity models, so that the reader can assess the size of the null effect.
- [Figure 3] The culture-specificity values (1.7–2.8) cluster in a narrow range near the bottom of the 1–5 scale, making differences hard to see; consider zooming into the relevant range or adding confidence intervals.
- [Limitations] The limitations section notes that the annotators are bicultural (living in English-speaking countries) and that cultural priming was used; please connect this explicitly to the human-rating baselines used in §4.1 and Figure 5, since these baselines are the only external ground truth in the study.
Circularity Check
No circular derivation found: the central claims are anchored to external human judgments, and the Section 4.2 self-comparison is a measurement-validity limitation, not an input-output identity.
full rationale
The paper's core claims are anchored to external human annotations rather than to GPT-4's own outputs. Human-written RoTs and human ratings of accuracy, culture-specificity, and stereotypicality are the reference points for the 'more generic' and 'hidden stereotypes' findings (Sections 3.2, 4.1, 4.3), so those results are not defined in terms of the model's responses. Section 4.2 measures alignment as Jensen-Shannon Divergence between GPT-4's default rating distribution and its own culturally-prompted rating distributions; this is a self-consistency measure and the US JSD of 0.12 is an empirical quantity, not an identity. The concern that country labels may only change output due to instruction-following rather than cultural knowledge is a construct-validity limitation, not circularity, and the paper explicitly acknowledges related threats (English Wikipedia plots possibly written from a Western editor's perspective; bicultural annotators; cultural prompts showing no significant advantage in Section 4.1). No parameter is fitted to the quantity it is said to predict, and no load-bearing claim depends on a self-citation: citations to Forbes et al. (2020) and Bhatia et al. (2024) are methodological conventions with independent foundations. Accordingly, no circular step is exhibited.
Assumptions & free parameters
free parameters (5)
- GPT-4o model instance =
OpenAI GPT-4o (proprietary)
- Top-quartile stereotypicality cutoff =
Top 25% within each culture (338 RoTs)
- Manually added negation RoTs =
A small number, e.g., 'It is bad to work all of the time' (China)
- Rating sampling scheme =
Temperature 0.8, 30 samples per RoT
- Movie plot sampling window =
20 movies per country, 40th to 60th percentile of length
assumptions (6)
- domain assumption Country can serve as a proxy for culture, with China, India, Iran, and the US as distinct cultural groups.
- domain assumption English Wikipedia movie plots provide culturally grounded narratives that elicit each culture's social norms.
- domain assumption Bicultural annotators (living in English-speaking countries, primed with cultural images) provide valid in-culture ground truth.
- domain assumption Social norms are adequately captured by the rules-of-thumb format with a judgment adjective and a verb-anchored action.
- ad hoc to paper Cultural prompting makes GPT-4 simulate the target country's perspective in the rating tasks.
- domain assumption Likert ratings (1 to 5) can be treated as interval data in OLS regressions.
Cite this review
Pith. "Pith review of Is It Bad to Work All the Time? Cross-Cultural Evaluation of Social Norm Biases in GPT-4." pith.science (2026). https://pith.science/paper/EJC6ZALM
@misc{pith2026250518322,
author = {Pith},
title = {Pith review of: Is It Bad to Work All the Time? Cross-Cultural Evaluation of Social Norm Biases in GPT-4},
year = {2026},
howpublished = {\url{https://pith.science/paper/EJC6ZALM}},
note = {Machine review of arXiv:2505.18322}
}
read the original abstract
LLMs have been demonstrated to align with the values of Western or North American cultures. Prior work predominantly showed this effect through leveraging surveys that directly ask (originally people and now also LLMs) about their values. However, it is hard to believe that LLMs would consistently apply those values in real-world scenarios. To address that, we take a bottom-up approach, asking LLMs to reason about cultural norms in narratives from different cultures. We find that GPT-4 tends to generate norms that, while not necessarily incorrect, are significantly less culture-specific. In addition, while it avoids overtly generating stereotypes, the stereotypical representations of certain cultures are merely hidden rather than suppressed in the model, and such stereotypes can be easily recovered. Addressing these challenges is a crucial step towards developing LLMs that fairly serve their diverse user base.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
For example, It is rude to curse at people
Judgment+action: Each RoT is in a single sentence with a straightforward structure: it is [the judgment] of [an action]. For example, It is rude to curse at people
-
[2]
curse; It is rude to curse at people
Verb-Centric: Anchor each RoT to a specific verb from the story (tense doesn’t matter). For example: “curse; It is rude to curse at people.”, “likes; It is good to like your relatives.”, “invited; It is devastating to be excluded from a wedding you were invited to.”
-
[3]
Specificity: Avoid overly generic statements
- [4]
-
[5]
arXiv preprint arXiv:2502.12057
Culture is not trivia: Sociocultural theory for cultural nlp. arXiv preprint arXiv:2502.12057. Ying Zhu and Junqi Peng. 2023. From diaosi to sang to tangping: The chinese dst youth subculture online. Global Storytelling: Journal of Digital and Moving Images, 3(2):13–38. Caleb Ziems, Jane Dwivedi-Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2023. NormBank...
arXiv 2023
- [10]
- [11]
-
[12]
(requires original story context)
cut; It’s ok to cut off contact... (requires original story context)
Show all 13 references
-
[13]
It is rude to curse at people
helped; It’s kind to help people. (too vague) Table 6: Prompt for the RoT collection task. 13 Figure 6: Cultural activation task interface. Annotators are presented with five culturally relevant images (e.g., national flag, historical figures, landmarks, daily life, and festiv...
-
[2010]
The weirdest people in the world? Behavioral and Brain Sciences, 33(2-3):61–83. 10 Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Pi- queras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierr...
2022 arXiv
-
[2017]
Science, 356(6334):183–186
Semantics derived automatically from lan- guage corpora contain human-like biases. Science, 356(6334):183–186. Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. 2023a. Assessing cross-cultural alignment between chatgpt and hu- man societies: An e...
2023 arXiv
-
[2024]
arXiv preprint arXiv:2406.03930
Culturally aware and adapted nlp: A taxonomy and a survey of the state of the art. arXiv preprint arXiv:2406.03930. Chen Cecilia Liu, Fajri Koto, Timothy Baldwin, and Iryna Gurevych. 2023. Are multilingual llms culturally-diverse reasoners? an investigation into multicultural ...
2023 arXiv
-
[2025]
Social norms in cinema: A cross-cultural anal- ysis of shame, pride and prejudice. In Proceedings of the 2025 Conference of the Nations of the Amer- icas Chapter of the Association for Computational Linguistics: Human Language Technologies (V olume 1: Long Papers), pages 11396...
2025 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.