REVIEW 3 major objections 4 minor 27 references
Minding the Politeness Gap in Cross-cultural Communication
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read British and American English speakers interpret 'quite' differently because of both different literal thresholds and different utterance-cost weights, not politeness norms alone.
desk verdict New UK/US intensifier data worth having, but the model comparison claiming 'semantic plus pragmatic' is not robust once you look at the paper's own BIC column. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a recursive listener model built on a speaker utility $U(w|s,\varphi_i,\varphi_s) = \varphi_i U_i(w|s) + \varphi_s U_s(w) - C(w)$, where the informativity term $U_i$ is the log-likelihood that a literal listener recovers the intended state, $U_s$ is the perceived social appropriateness of the utterance (taken from politeness ratings), and $C(w)$ penalizes adding a modifier. Culture enters through twelve double-threshold parameters fixing the literal denotation of each modifier and the unmodified baseline, plus pragmatic weights $\varphi_i,\varphi_s$ and cost $C(w)$. A pragmatic listener inverts the speaker's softmax production model via Bayes' rule, and the authors compare models allowing different subsets of these parameters to vary between UK and US to see which cultural variation is necessary and sufficient.
What would settle it
Fit the same model to simulated datasets generated from a known model with only threshold differences between two cultures and no cost differences; if the optimizer attributes the difference to cost or informativity weights instead, the semantic/pragmatic split is not identifiable from this design. A simpler version: compute profile likelihood or bootstrap confidence intervals for the threshold and cost parameters and check whether the culture differences survive the uncertainty.
Extended reading notes
Core claim
In the authors' terms, cross-cultural variation in modifier interpretation arises through both semantic variation and pragmatic variation: British and American participants differ in the literal threshold values they assign to intensifiers, most clearly for 'quite', and they differ in the perceived cost of producing a modifier and the weight placed on being informative. These conclusions come from comparing models in which different subsets of parameters are allowed to vary across cultures: the best-fitting model allows both semantic thresholds and pragmatic weights to differ, and constraining either the utterance cost or the informativity weight to be shared across cultures significantly degrades fit. The model's social-utility term contributed little to fit, and perceived politeness ratings did not linearly explain the cross-cultural interpretation differences, even though politeness ratings did predict interpretation overall. A robustness check shows the model generalizes reasonably to narrator-framed data when the social term is removed, with a moderate loss increase.
Load-bearing premise
The clean separation of semantic from pragmatic causes assumes that the twelve threshold parameters and the pragmatic weights are jointly identifiable from the pooled z-scored ratings, so a fitted threshold difference is not just absorbing what is really a cost difference (or vice versa).
Editorial extensions
If this is right
- A purely semantic account is insufficient: models that let only literal thresholds vary miss the improvement gained from also letting cost and informativity weights vary.
- A purely politeness-based account is insufficient: the 'quite' gap persists in narrator-framed utterances where politeness pressure is removed, and politeness ratings do not explain cross-cultural differences.
- Culture-specific utterance cost means listeners infer different amounts of extra meaning from the same act of modification depending on the culture of the speaker.
- Practical cross-cultural communication and machine translation should treat intensifiers as carrying both culture-specific denotations and culture-specific pragmatic weights.
Reading between the lines
- If the identified cost difference is real, one testable prediction is that UK and US listeners should differ in how much extra strength they infer from the mere presence of a modifier even for novel or invented modifiers with no established semantic difference.
- The joint estimation of twelve thresholds and three pragmatic weights from z-scored ratings leaves a possible trade-off between threshold shifts and cost shifts; a synthetic recoverability analysis would tell whether the semantic/pragmatic split is an artifact of the optimizer.
- The paper's own note that a culture-wide social utility parameter may average over heterogeneous subcommunities suggests a hierarchical extension with regional or dialect-level politeness norms could revive the role of politeness in the model.
- The model's clustering by modifier regardless of predicate, and the outliers for marked expressions like 'extremely exhausted', point toward predicate-specific social meaning rather than a single modifier-level utility as the next modeling step.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports three behavioral experiments comparing how British and American English speakers interpret intensifiers such as "quite" and "very," embedded in dialogue, narrator, and politeness-rating contexts. The experiments show robust cross-cultural differences, particularly for "quite" and "very," and show that politeness ratings alone do not explain the interpretation differences. The authors then develop an RSA-style computational model in which culture-specific semantic thresholds and culture-specific pragmatic weights (informativity, social utility, utterance cost) are fitted to the z-scored rating data. The central claim is that cross-cultural differences in interpretation arise from a combination of different literal meanings and different weights on utterance cost/informativity, rather than from politeness norms alone. The behavioral findings are well supported, but the modeling conclusion is weakened by model-selection ambiguity and the absence of identifiability checks.
Significance. The paper addresses a genuinely important question in cross-cultural pragmatics: whether UK-US differences in intensifier interpretation are semantic, pragmatic, or both. The behavioral experiments are a real contribution: they document specific, replicable differences for "quite" and "very," and the narrator manipulation is a sensible control for politeness-related pragmatics. The computational framework extends the RSA politeness model of Yoon et al. (2020) to a cross-cultural setting, and the authors provide code and materials. If the central modeling claim were established, the paper would push the field toward models that jointly incorporate culture-specific lexical semantics and culture-specific utterance costs. However, as it stands, the model-comparison evidence for the "both semantic and pragmatic" conclusion is not robust: the paper's own BIC values favor a simpler semantic-only model, and the parameter separation is not tested for identifiability. These issues are load-bearing for the paper's main claim, but they are addressable with additional analysis.
major comments (3)
- [Results / Model comparison (Table 2)] The paper states that the best-fitting model (M9) integrates cross-cultural differences in both semantic and pragmatic factors, but Table 2 does not uniformly support this. M9 has BIC 22631, while M6 (only "quite" threshold varies, no pragmatic parameters) has BIC 22590, and M8 (all thresholds vary, semantic only) has BIC 22656. BIC therefore favors M6 by 41 points over M9, while AIC favors M9 (22448 vs. 22485 for M8). The authors do not justify choosing AIC over BIC, nor do they report cross-validated log loss. Because the central conclusion depends on preferring M9, the paper needs to report a held-out predictive comparison or otherwise justify the model-selection criterion. Without that, the alternative reading under BIC—that only "quite" has a different literal threshold—is equally compatible with the reported numbers.
- [Computational Model and Appendix: Optimization Process] The separation of semantic from pragmatic causes assumes that the twelve semantic threshold parameters and the pragmatic weights (φ_i, φ_s, C(w)) are jointly identifiable from the pooled z-scored ratings. No parameter-recovery simulations, confidence intervals, or profile likelihoods are reported. If threshold shifts and cost shifts can trade off, a fitted UK-US difference in thresholds could be partially or wholly an artifact of the optimizer. The manuscript should include synthetic recoverability checks (e.g., simulate data from known parameters and confirm the fitting procedure recovers them) and some measure of parameter uncertainty before attributing differences cleanly to "different literal meanings" versus "different pragmatics." This issue is central to the paper's main claim.
- [Appendix: Robustness to dropping modifiers (Table 4)] The robustness check shows that dropping "extremely" changes the log loss from 11250 to 11287 on the original data, which is a larger change than dropping "impressive" or "difficult." The authors acknowledge this sensitivity, but it is not incorporated into the model-comparison conclusion. In addition, Figure 2's "model predictions" are in-sample fits from parameters optimized on the same data, so the visual agreement is not independent evidence. The paper would be substantially strengthened by an out-of-sample evaluation, such as cross-validation by predicate or by participant, to demonstrate that the M9 parameterization generalizes rather than overfitting.
minor comments (4)
- [Experiment 3: Politeness Ratings] There is a typo: "supprting" should be "supporting."
- [Table 2] The column labeled "Log Loss" is not defined in the text; it appears to be the negative log likelihood, and the relationship between this column and AIC/BIC should be stated explicitly.
- [Figure 1 caption] The caption "Experiment 1: Dialogue Experiment 2: Narrator Experiment 3: Politeness" lacks punctuation between panels and would be clearer as "Experiment 1: Dialogue; Experiment 2: Narrator; Experiment 3: Politeness."
- [Experiment 1: Dialogue Context] The valence analysis reports p=.053 for British participants' stronger interpretation of "very" with positive predicates and describes this as "most pronounced"; given the marginal p-value, the wording should be hedged.
Circularity Check
In-sample 'model predictions' and a z-scored Gaussian prior are the only circular elements; the central model comparison is not circular but is fragile.
-
fitted input called prediction
[Computational Model / Results, Figure 2 caption; Discussion]
"The predictions generated are from the model parameters specifically optimized to fit each country's data separately. ... Figure 2 suggests that despite this limitation, the best-fitting models for each culture successfully capture the overall pattern of behavioral responses, with model predictions closely tracking empirical data for most modifier-predicate combinations."
The parameters were optimized to minimize prediction loss on the same Experiment 1 ratings that Figure 2 plots, so 'model predictions closely tracking empirical data' is an in-sample fit guaranteed to improve with more parameters, not an independent test of the model. The figure is therefore presented as validation of the semantic/pragmatic decomposition while actually re-expressing the fitted values; the paper reports no held-out loss, cross-validation, or parameter-recovery check that would let the plot discriminate models.
-
self definitional
[Appendix, 'Modeling state prior as Gaussian']
"We modeled the state prior as a normal distribution with mean 0 and variance 1. As seen in Figure 3, this assumption matches relatively well with the distribution of z-scored responses from participants suggesting that participants expect states to be distributed normally a priori."
The fit is partly constructed: Experiment 1 z-scored all responses within each participant, so the empirical distribution has mean 0 and variance 1 by definition. A N(0,1) prior therefore matches the first two moments of the data regardless of the true state distribution; only the shape comparison is non-tautological, and no shape test is reported. The step is ancillary, but as stated it presents a preprocessing artifact as empirical support for the Gaussian prior.
full rationale
The main derivation—comparing RSA models with culture-specific semantic thresholds versus pragmatic weights via AIC/BIC—is not circular: the model parameters are fit to Experiment 1 data, and the model comparisons are genuine, if in-sample, statistical selections. The central claim that both semantic and pragmatic variation are needed rests on Table 2, not on the Figure 2 plot. However, two supporting steps do reduce to their inputs by construction: the Figure 2 'model predictions' are in-sample fits (parameters optimized on the very data plotted), and the N(0,1) prior 'match' is partially guaranteed by z-scoring. Also note that Table 2's BIC column favors the semantic-only M6 (BIC 22590) over the full M9 (BIC 22631), while the paper's choice of AIC over BIC is unargued; this is a robustness/correctness concern, not circularity. The paper transparently discloses that the predictions come from per-country fits and discusses outliers, which limits the severity. Net: partial circularity in the validation figures, but not in the formal model comparison.
Assumptions & free parameters
free parameters (4)
- Semantic double thresholds for modifiers and baseline =
not reported in paper
- Informativity weight phi_i per culture =
not reported in paper
- Social utility weight phi_s per culture =
not reported in paper
- Utterance cost C(w) per culture =
not reported in paper
assumptions (4)
- domain assumption RSA recursive speaker-listener model is the correct process model for modifier interpretation
- domain assumption State prior is Gaussian N(0,1)
- ad hoc to paper Literal meaning of each modifier is a smooth double threshold
- domain assumption Politeness ratings from Experiment 3 are a valid proxy for social utility
Cite this review
Pith. "Pith review of Minding the Politeness Gap in Cross-cultural Communication." pith.science (2026). https://pith.science/paper/WPUTZP5W
@misc{pith2026250615623,
author = {Pith},
title = {Pith review of: Minding the Politeness Gap in Cross-cultural Communication},
year = {2026},
howpublished = {\url{https://pith.science/paper/WPUTZP5W}},
note = {Machine review of arXiv:2506.15623}
}
read the original abstract
Misunderstandings in cross-cultural communication often arise from subtle differences in interpretation, but it is unclear whether these differences arise from the literal meanings assigned to words or from more general pragmatic factors such as norms around politeness and brevity. In this paper, we report three experiments examining how speakers of British and American English interpret intensifiers like "quite" and "very." To better understand these cross-cultural differences, we developed a computational cognitive model where listeners recursively reason about speakers who balance informativity, politeness, and utterance cost. Our model comparisons suggested that cross-cultural differences in intensifier interpretation stem from a combination of (1) different literal meanings, (2) different weights on utterance cost. These findings challenge accounts based purely on semantic variation or politeness norms, demonstrating that cross-cultural differences in interpretation emerge from an intricate interplay between the two.
Figures
Reference graph
Works this paper leans on
-
[1]
Brown, P., & Levinson, S. C. (1987).Politeness: Some uni- versals in language usage. Cambridge University Press
work page 1987
-
[2]
Chandra, K., Chen, T., Tenenbaum, J. B., & Ragan-Kelley, J. (2025). A domain-specific probabilistic programming language for reasoning about reasoning (or: a memo on memo).PsyArXiv preprint
work page 2025
-
[3]
Culpeper, J., Haugh, M., & K´ad´ar, D. Z. (2017).The palgrave handbook of linguistic (im) politeness. Springer
work page 2017
-
[4]
Desagulier, G. (2014). Visualizing distances in a set of near-synonyms: Rather, quite, fairly, and pretty. InCor- pus methods for semantics(pp. 145–178). John Benjamins Publishing Company
work page 2014
-
[5]
Fowler, H. W. (2015).Fowler’s dictionary of modern english usage. Oxford University Press
work page 2015
-
[6]
Goddard, C. (2012). ‘Early interactions’ in Australian En- glish, American English, and English English: Cultural dif- ferences and cultural scripts.Journal of Pragmatics,44, 1038-1050
work page 2012
-
[7]
Goddard, C., & Wierzbicka, A. (2004). Cultural scripts: What are they and what are they good for?Intercultural Pragmatics,1(2), 153–166
work page 2004
-
[8]
(2013).Words and mean- ings: Lexical semantics across domains, languages, and cultures
Goddard, C., & Wierzbicka, A. (2013).Words and mean- ings: Lexical semantics across domains, languages, and cultures. Oxford University Press
work page 2013
Show all 27 references
-
[9]
Grice, H. P. (1975). Logic and conversation.Syntax and Semantics,3, 41-58
1975
-
[10]
(2023).The CMA Evolution Strategy: A Tutorial
Hansen, N. (2023).The CMA Evolution Strategy: A Tutorial
2023
-
[11]
Haugh, M., & Bousfield, D. (2012). Mock impoliteness, jocular mockery and jocular abuse in Australian and British English.Journal of Pragmatics,44(9), 1099-1114
2012
-
[12]
Haugh, M., & Schneider, K. P. (2012). Im/politeness across englishes.Journal of Pragmatics,44(9), 1017-1021
2012
-
[13]
Kecskes, I. (2023). Language variation and temporary norm development in intercultural interactions.Pragmatics & Cognition,30(2), 235–257
2023
-
[14]
Lumer, E., & Buschmeier, H. (2022). Modeling social in- fluences on indirectness in a rational speech act approach to politeness. InProceedings of the Annual Meeting of the Cognitive Science Society
2022
-
[15]
(2012).This is so cool! a comparative corpus study on intensifiers in British and American English.(Un- published manuscript, University of Tampere) Ruzait˙e, J
Romero, S. (2012).This is so cool! a comparative corpus study on intensifiers in British and American English.(Un- published manuscript, University of Tampere) Ruzait˙e, J. (2007). Vague references to quantities as a face- saving strategy in teacher-student interaction.Lodz Pa...
2012
-
[16]
(2017).Pragmatic aspects of scalar modifiers: The semantics-pragmatics interface(V ol
Sawada, O. (2017).Pragmatic aspects of scalar modifiers: The semantics-pragmatics interface(V ol. 69). Oxford Uni- versity Press
2017
-
[17]
Schneider, K. P. (2012). Appropriate behaviour across vari- eties of english.Journal of Pragmatics,44(9), 1022-1037
2012
-
[18]
Schneider, K. P. (2024).Pragmatic variation within lan- guages(V ol. 232). Elsevier
2024
-
[19]
Rieser, V
Evans, G., . . . Rieser, V . (2025). Value profiles for encod- ing human variation.arXiv preprint arXiv:2503.15484
2025
-
[20]
Spencer-Oatey, H., & K´ad´ar, D. Z. (2021).Intercultural po- liteness: Managing relations across cultures. Cambridge University Press
2021
-
[21]
That’s proper cool
Stratton, J. M. (2021). “That’s proper cool”: The emerging intensifier proper in British English.English Today,37(4), 206–213
2021
-
[22]
Su, Y . (2016). Corpus-based comparative study of inten- sifiers: quite, pretty, rather and fairly.Journal of World Languages,3, 224-236
2016
-
[23]
Thomas, J. (1983). Cross-cultural pragmatic failure.Applied linguistics,4(2), 91–112
1983
-
[24]
Troutman, D. (2022). Sassy sasha?: The intersectionality of (im) politeness and sociolinguistics.Journal of Politeness Research,18(1), 121–149. van Dorst, I., Gillings, M., & Culpeper, J. (2024). Socioprag- matic variation in britain: A corpus-based study of polite- ness.Journ...
2022
-
[25]
It’s rude to VP
Waters, S. (2012). “It’s rude to VP”: The cultural semantics of rudeness.Journal of Pragmatics,44(9), 1051-1062
2012
-
[26]
White, I., Pandey, S., & Pan, M. (2024). Communicate to play: Pragmatic reasoning for efficient cross-cultural com- munication. In Y . Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.),Findings of the Association for Computational Lin- guistics: EMNLP 2024(pp. 12201–12216)
2024
-
[27]
J., Tessler, M
Yoon, E. J., Tessler, M. H., Goodman, N. D., & Frank, M. C. (2020). Polite speech emerges from competing so- cial goals.Open Mind : Discoveries in Cognitive Science, 4, 71 - 87. Appendix Model ImplementationWe used the memo language (Chandra, Chen, Tenenbaum, & Ragan-Kelley, 2...
2020
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.