REVIEW 3 major objections 2 minor 25 references
Against 'softmaxing' culture
T0 review · 3 major / 2 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper argues that machine-learning and HCI evaluations of culture fail because they open with 'what is culture?' and that shifting to 'when is culture?' would make them more responsive to cultural complexity.
desk verdict A well-written position paper with a memorable metaphor and a fresh philosophical source, but its central 'when is culture?' reframing is never operationalized, so the practical payoff is asserted rather than shown. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a question shift, from 'what is culture?' to 'when is culture?', paired with a philosophical distinction between cultural universals and particulars. In the paper, universals are the human capabilities—reflective perception, abstraction, deduction, induction, and sympathetic impartiality—that make intercultural communication and translation possible; particulars are the situated forms those universals take. The question 'when is culture?' does the work: it directs evaluators to ask under which relational conditions a cultural value or practice is valid and tenable, instead of trying to define culture in advance. The named target phenomenon, 'softmaxing culture,' names the failure mode this machinery is meant to correct: the probabilistic logic of softmax that privileges dominant, statistically frequent expressions and conflates statistical probability with cultural appropriateness.
What would settle it
Run two evaluation studies on the same model outputs, one scoped by 'what is culture?' and one by 'when is culture?'; if they identify the same failures and rank the same systems best, the paper's claim that the framing question is decisive is falsified.
Extended reading notes
Core claim
The paper's central claim is that a failure mode it names 'softmaxing culture'—large AI models amplifying statistically frequent linguistic and cultural expressions while suppressing non-frequent ones—is worsened by how evaluations are scoped. ML and HCI researchers, despite different methods, share a reliance on macro-level definitions such as Hofstede's cultural dimensions or benchmarks built from static snapshots of language. These capture narrow aspects and can stereotype or essentialize people. The author's proposed correction is conceptual: replace 'what is culture?' with 'when is culture?' as the scoping question, and treat cultural universals, the shared capabilities that make intercultural communication possible, only in relation to their particulars. The intended result is evaluation practice that asks about relational validity and tenability rather than identifying fixed cultural traits.
Load-bearing premise
The whole position depends on the premise that the opening question is the main obstacle to better cultural evaluation, rather than data, funding, tooling, or other practical constraints.
Editorial extensions
If this is right
- Evaluations that begin with 'when is culture?' would be judged by their relational validity, not by how well outputs match pre-defined cultural dimensions.
- Culture benchmarks would need to include particulars—localized, low-resource, underrepresented expressions—rather than aiming at universal descriptions alone.
- Current LLM evaluations should measure suppression of non-frequent expressions, not just average accuracy or benchmark scores.
- ML and HCI evaluation practices would converge on a shared conceptual framing, reducing the divide the paper identifies.
- 'Softmaxing' would become a first-class failure mode in cultural alignment, alongside bias and stereotyping.
Reading between the lines
- A testable extension: build two evaluation protocols for the same model outputs, one scoped by 'what is culture?' and one by 'when is culture?', and compare which decisions they produce; the paper predicts they diverge in ways that matter for low-resource languages.
- The 'when is culture?' framing may connect naturally to time-sensitive and context-sensitive evaluation methods such as longitudinal or in-the-wild studies, though the paper only gestures at real-world contexts.
- The argument implies that purely technical fixes—more data, bigger benchmarks—will not resolve cultural flattening because the problem is in the framing question itself.
- If the conceptual shift is adopted, 'cultural universals' risk becoming a new form of flattening unless the particulars relation is operationalized; universals should never be measured alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that current ML and HCI evaluations of culture in large language models are limited because they rely on static, macro-level definitions of culture, exemplified by Hofstede's cultural dimensions. The author identifies a phenomenon called 'softmaxing culture,' in which models homogenize diverse cultural expressions into statistically dominant generic forms. The paper proposes two conceptual shifts: (1) begin evaluations with the question 'when is culture?' rather than 'what is culture?', and (2) situate cultural universals in relation to their particulars, drawing on Wiredu's philosophical work. The paper claims these shifts will make evaluations more responsive to the complexity of culture.
Significance. If the proposed reframing were made precise and shown to have practical import, the paper would contribute a useful conceptual corrective to the growing literature on cultural evaluation in AI. It connects AI evaluation practice to an established philosophical tradition (Wiredu) and to recent critiques such as Zhou et al. (2025) and Wallach et al. (2025). The paper also names a real and important phenomenon--the homogenization of culture by LLMs--and points to relevant prior work. However, the central claim is under-specified: the 'when is culture?' question is never operationalized, and the paper explicitly limits itself to conceptual aspects. As a result, the argument's practical significance remains undemonstrated.
major comments (3)
- [Section 3 (paragraph introducing 'when is culture?')] The central proposal, 'when is culture?', is never operationalized. The paper does not specify what counts as an answer to this question: whether 'when' refers to a temporal index, a context of use, a relational occasion, or all three. No worked example is provided, and no existing evaluation is shown to be ruled in or out by this reframing. Since the paper's stated goal is to invite evaluation approaches that are more responsive to culture's complexity, the absence of any concrete illustration makes the central claim unfalsifiable and lacking operational content. At minimum, the paper should provide one example of an evaluation that would be designed differently under the 'when' framing, and one example of an evaluation that would not.
- [Section 1 and Section 5] The paper explicitly limits itself to 'the conceptual aspects of culture and its formulation in evaluation' (Section 1), yet Section 5 claims that the reframing will lead to 'deep and meaningful integration of theoretical and empirical strategies' and implies practical progress. The paper offers no mechanism by which asking 'when is culture?' changes evaluation practice, and it does not engage with alternative explanations for the limitations it identifies, such as data availability, tooling, or institutional incentives. The practical significance claimed in the abstract and Section 5 therefore depends on an assumption that is neither argued nor tested. The authors should either narrow the paper's claims to the purely conceptual level or provide a plausible pathway from the reframing to concrete changes in evaluation design.
- [Section 4] The paper's use of Wiredu's argument is not logically connected to the proposed reframing. Section 4 establishes that Wiredu posits cultural universals and particulars, and that universals enable intercultural communication. It does not establish that asking 'when is culture?' is the right way to relate universals and particulars. The paper asserts this connection without argument. To make the central claim load-bearing, the author should articulate why 'when' is the appropriate relational frame, as distinct from other possible frames such as 'how', 'where', or 'with whom'. As written, Wiredu's philosophy supports the existence of universals and particulars but not the specific 'when' question.
minor comments (2)
- [Throughout] There are several typos and formatting issues: 'such asMasakhane and Deep Learning Indabahave' is missing spaces (Section 3), 'modes' should likely be 'models' (Section 2), 'Humanites' should be 'Humanities' in the Bajohr reference, and the footer '© 2010 D. Mwesigwa' and the venue header 'Proceedings of Machine Learning Research 1:1–7, 2010' appear to have the wrong year. These should be corrected.
- [Section 1, footnote 1] The footnote distinguishing 'softmaxing' from 'softmaxxing' is helpful, but the main text could more clearly define the scope of the metaphor relative to the softmax function in transformer models, since the analogy is central to the paper's framing.
Circularity Check
No circularity: the paper is a position piece with no fitted parameters, derived quantities, or self-citation chains; its central reframing argument rests on external philosophical and empirical citations.
full rationale
This is an opinion/position paper, not a derivation. The central proposal—replacing 'what is culture?' with 'when is culture?' and situating universals in relation to particulars—is asserted and argued via external sources (Hofstede 2001/2011; Zhou et al. 2025; Wiredu 1995/1996). There are no equations, no fitted values, and no predicted quantities that could reduce to inputs. The author introduces the term 'softmaxing culture' as a metaphor and explicitly notes the phenomenon is distinct from a beauty trend; the metaphor is not used to derive evaluation outcomes. No load-bearing step depends on a citation from the present author, and the sole author cites no self-authored prior work. The most that can be said is that the practical payoff of the reframing is not operationalized, but that is a question of empirical support and impact, not circularity. Per the scoring rules, lack of operational specificity is a correctness concern, not a circularity concern. Therefore the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Cultural universals exist, per Wiredu.
- domain assumption Culture is dynamic and relational, not a static set of attributes.
- ad hoc to paper The question "what is culture?" is the wrong starting point for evaluations.
Cite this review
Pith. "Pith review of Against 'softmaxing' culture." pith.science (2026). https://pith.science/paper/EIE3FDEH
@misc{pith2026250622968,
author = {Pith},
title = {Pith review of: Against 'softmaxing' culture},
year = {2026},
howpublished = {\url{https://pith.science/paper/EIE3FDEH}},
note = {Machine review of arXiv:2506.22968}
}
read the original abstract
AI is flattening culture. Evaluations of "culture" are showing the myriad ways in which large AI models are homogenizing language and culture, averaging out rich linguistic differences into generic expressions. I call this phenomenon "softmaxing culture,'' and it is one of the fundamental challenges facing AI evaluations today. Efforts to improve and strengthen evaluations of culture are central to the project of cultural alignment in large AI systems. This position paper argues that machine learning (ML) and human-computer interaction (HCI) approaches to evaluation are limited. I propose two key conceptual shifts. First, instead of asking "what is culture?" at the start of system evaluations, I propose beginning with the question: "when is culture?" Second, while I acknowledge the philosophical claim that cultural universals exist, the challenge is not simply to describe them, but to situate them in relation to their particulars. Taken together, these conceptual shifts invite evaluation approaches that move beyond technical requirements toward perspectives that are more responsive to the complexities of culture.
Reference graph
Works this paper leans on
-
[1]
David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime, Jian Yun Zhuang, Jesujoba O. Alabi, Xuanli He, Millicent Ochieng, Sara Hooker, Andiswa Bukula, En-Shiun Annie Lee, Chiamaka Chukwuneke, Happy Buzaaba, Blessing Sibanda, Godson Kalipe, Jonathan Mukiibi, Salomon Kabongo, Foutse Yuehgoh, Mmasibidi Setaka, Lolwethu Ndolela, Nkiruka Odu, Rooweither Mabuy...
arXiv 2025
-
[2]
Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Alham Fikri Aji, Jacki O'Neill, Ashutosh Modi, and Monojit Choudhury. Towards Measuring and Modeling " Culture " in LLMs : A Survey , September 2024. URL http://arxiv.org/abs/2403.15412. arXiv:2403.15412 [cs]
arXiv 2024
-
[3]
AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances
Dhruv Agarwal, Mor Naaman, and Aditya Vashistha. AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances . In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems , CHI '25, pages 1--21, New York, NY, USA, April 2025. Association for Computing Machinery. ISBN 9798400713941. doi:10.1145/3706598.3713564....
arXiv 2025
-
[4]
Thinking with AI : Machine Learning the Humanities
Hannes Bajohr, editor. Thinking with AI : Machine Learning the Humanities . Open Humanites Press, 2025. ISBN 978-1-78542-141-9 978-1-78542-140-2. URL https://www.openhumanitiespress.org/books/titles/thinking-with-ai/
work page 2025
-
[5]
Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the Dangers of Stochastic Parrots : Can Language Models Be Too Big ? In Proceedings of the 2021 ACM Conference on Fairness , Accountability , and Transparency , FAccT '21, pages 610--623, New York, NY, USA, March 2021. Association for Computing Machinery. ISBN 978-1-4503...
arXiv 2021
-
[6]
Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stef...
arXiv 2022
-
[7]
Myra Cheng, Esin Durmus, and Dan Jurafsky. Marked Personas : Using Natural Language Prompts to Measure Stereotypes in Language Models , May 2023. URL http://arxiv.org/abs/2305.18189. arXiv:2305.18189 [cs]
arXiv 2023
-
[8]
Rémy Decoupes, Roberto Interdonato, Mathieu Roche, Maguelonne Teisseire, and Sarah Valentin. Evaluation of Geographical Distortions in Language Models : A Crucial Step Towards Equitable Representations . volume 15243, pages 86--100. 2025. doi:10.1007/978-3-031-78977-9_6. URL http://arxiv.org/abs/2404.17401. arXiv:2404.17401 [cs]
Show all 25 references
-
[9]
Large AI models are cultural and social technologies
Henry Farrell, Alison Gopnik, Cosma Shalizi, and James Evans. Large AI models are cultural and social technologies. Science, 387 0 (6739): 0 1153--1156, March 2025. doi:10.1126/science.adt9819. URL https://www.science.org/doi/10.1126/science.adt9819. Publisher: American Associ...
2025 doi
-
[10]
The Unreasonable Effectiveness of Data
Alon Halevy, Peter Norvig, and Fernando Pereira. The Unreasonable Effectiveness of Data . IEEE Intelligent Systems, 24 0 (2): 0 8--12, March 2009. ISSN 1541-1672. doi:10.1109/MIS.2009.36. URL http://ieeexplore.ieee.org/document/4804817/
2009
-
[11]
Culture's Consequences : Comparing Values , Behaviors , Institutions and Organizations Across Nations
Geert Hofstede. Culture's Consequences : Comparing Values , Behaviors , Institutions and Organizations Across Nations . SAGE, second edition, 2001. ISBN 978-0-8039-7324-4. Google-Books-ID: w6z18LJ\_1VsC
2001
-
[12]
Dimensionalizing Cultures : The Hofstede Model in Context
Geert Hofstede. Dimensionalizing Cultures : The Hofstede Model in Context . Online Readings in Psychology and Culture, 2 0 (1), December 2011. ISSN 2307-0919. doi:10.9707/2307-0919.1014. URL https://scholarworks.gvsu.edu/orpc/vol2/iss1/8
2011
-
[13]
Lilly Irani, Janet Vertesi, Paul Dourish, Kavita Philip, and Rebecca E. Grinter. Postcolonial computing: a lens on design and development. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems , CHI '10, pages 1311--1320, New York, NY, USA, April 2010. ...
2010
-
[14]
Culturally Aware and Adapted NLP : A Taxonomy and a Survey of the State of the Art , March 2025
Chen Cecilia Liu, Iryna Gurevych, and Anna Korhonen. Culturally Aware and Adapted NLP : A Taxonomy and a Survey of the State of the Art , March 2025. URL http://arxiv.org/abs/2406.03930. arXiv:2406.03930 [cs]
2025 arXiv
-
[15]
Masoud, Ziquan Liu, Martin Ferianc, Philip Treleaven, and Miguel Rodrigues
Reem I. Masoud, Ziquan Liu, Martin Ferianc, Philip Treleaven, and Miguel Rodrigues. Cultural Alignment in Large Language Models : An Explanatory Analysis Based on Hofstede 's Cultural Dimensions , May 2024. URL http://arxiv.org/abs/2309.12342. arXiv:2309.12342 [cs]
2024 arXiv
-
[16]
Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbohungbe, Solomon Oluwole Akinola, Shamsuddeen Muhammad, Salomon Kabongo Kabenamualu, Salomey Osei, Freshia Sackey, Rubungo Andre Niyongabo, Ricky Macharm, Perez Ogayo, Orevaoghene Ahia, Musie Mere...
2020
-
[17]
Beyond Metrics : Evaluating LLMs ' Effectiveness in Culturally Nuanced , Low - Resource Real - World Scenarios , June 2024
Millicent Ochieng, Varun Gumma, Sunayana Sitaram, Jindong Wang, Vishrav Chaudhary, Keshet Ronen, Kalika Bali, and Jacki O'Neill. Beyond Metrics : Evaluating LLMs ' Effectiveness in Culturally Nuanced , Low - Resource Real - World Scenarios , June 2024. URL http://arxiv.org/abs...
2024 arXiv
-
[18]
Cultural Incongruencies in Artificial Intelligence , November 2022
Vinodkumar Prabhakaran, Rida Qadri, and Ben Hutchinson. Cultural Incongruencies in Artificial Intelligence , November 2022. URL http://arxiv.org/abs/2211.13069. arXiv:2211.13069 [cs]
2022 arXiv
-
[19]
CultureBank : An Online Community - Driven Knowledge Base Towards Culturally Aware Language Technologies , April 2024
Weiyan Shi, Ryan Li, Yutong Zhang, Caleb Ziems, Chunhua yu, Raya Horesh, Rogério Abreu de Paula, and Diyi Yang. CultureBank : An Online Community - Driven Knowledge Base Towards Culturally Aware Language Technologies , April 2024. URL http://arxiv.org/abs/2404.15238. arXiv:240...
2024 arXiv
-
[20]
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David I. Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Wei-Yin Ko, Sebastian Ruder, Madeline Smith, Antoine Bosselut, Alic...
2025 arXiv
-
[21]
Feder Cooper, Angelina Wang, Chad Atalla, Solon Barocas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P
Hanna Wallach, Meera Desai, A. Feder Cooper, Angelina Wang, Chad Atalla, Solon Barocas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Alexandra Olteanu, Nicholas Pangakis, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vaugha...
2025 arXiv
-
[22]
Are There Cultural Universals ? The Monist, 78 0 (1): 0 52--64, 1995
Kwasi Wiredu. Are There Cultural Universals ? The Monist, 78 0 (1): 0 52--64, 1995. ISSN 0026-9662. URL https://www.jstor.org/stable/27903418. Publisher: Oxford University Press
1995
-
[23]
Cultural universals and particulars: An African perspective
Kwasi Wiredu. Cultural universals and particulars: An African perspective . African Systems of Thought . Indiana University Press, Bloomington, January 1996. ISBN 978-0-253-21080-7. URL https://research.ebsco.com/linkprocessor/plink?id=b733d0e7-ab78-3554-b888-33529bb83cf5
1996
-
[24]
Chawla, and Xiangliang Zhang
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, Nitesh V. Chawla, and Xiangliang Zhang. Justice or Prejudice ? Quantifying Biases in LLM -as-a- Judge , October 2024. URL http://arxiv.org/abs/2410.02736. ...
2024 arXiv
-
[25]
Naitian Zhou, David Bamman, and Isaac L. Bleaman. Culture is Not Trivia : Sociocultural Theory for Cultural NLP , February 2025. URL http://arxiv.org/abs/2502.12057. arXiv:2502.12057 [cs]
2025 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.