{"id":"47714213-3839-4ab4-b2e8-157df2c56d37","arxiv_id":"2506.15623","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"UK/US differences in intensifier interpretation are best explained by a combination of different literal meanings, especially for 'quite', and different cultural weights on utterance cost and informativity.","lead":"Researchers compared how British and American English speakers interpret words like 'quite' and 'very', and built a computer model of polite conversation to explain the differences. The model suggests the gap comes from a mix of different word meanings and different cultural views on how costly it is to add extra words, rather than from politeness alone.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that both semantic and pragmatic variation are needed is not robust: Table 2's BIC favors the semantic-only model M6 (22590) over M9 (22631), while AIC favors M9; without cross-validation the 'both' conclusion is an artifact of model-selection criterion.","rationale":"The reader focused on identifiability of threshold vs cost parameters. I agree that is a risk, but the most direct threat is the model-selection step: the paper's own BIC values favor a purely semantic model, so the central 'combination' claim depends on an unstated preference for AIC. This is checkable and fixable. The behavioral experiments and code availability are real assets, and the limitation discussions are candid, so a conditional acceptance remains appropriate. The paper should be revised to include cross-validated model comparison and report fitted parameter values with uncertainty.","tokens_in":10082,"tokens_out":5301,"duration_ms":62441,"concrete_test":"Run leave-one-predicate-out cross-validation comparing M6 (semantic only) and M9 (all parameters) with the same CMA-ES optimization, reporting held-out log loss and bootstrap CIs. If M9 does not outperform M6 (or M8) out-of-sample, the pragmatic component of the headline claim is unsupported; if it does, the BIC objection is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that UK-US intensifier differences arise from a combination of different literal thresholds and different pragmatic weights on cost/informativity—rests entirely on the model comparison in Table 2. The paper reports that M9 ('all') is 'the best-fitting model,' but the table's own BIC column contradicts this: M6 ('quite' only, semantic) has BIC 22590, M8 (all thresholds, semantic only) has BIC 22656, and M9 has BIC 22631. M6 is the best BIC model by 41 points. AIC prefers M9 (22448 vs 22485), but the authors never justify choosing AIC over BIC or report cross-validated log loss. The marginal contribution of the three pragmatic parameters is ΔLL=25 relative to M8; with N≈3400, that is strong by AIC but negative by BIC. Since the paper's own robustness section shows sensitivity to dropping 'extremely' and no parameter estimates or confidence intervals are reported, the 'both semantic and pragmatic' conclusion is not yet identifiable from this comparison. The alternative reading—only 'quite' has a different literal threshold—is equally compatible with the reported numbers under BIC.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports three behavioral experiments comparing how British and American English speakers interpret intensifiers such as \"quite\" and \"very,\" embedded in dialogue, narrator, and politeness-rating contexts. The experiments show robust cross-cultural differences, particularly for \"quite\" and \"very,\" and show that politeness ratings alone do not explain the interpretation differences. The authors then develop an RSA-style computational model in which culture-specific semantic thresholds and culture-specific pragmatic weights (informativity, social utility, utterance cost) are fitted to the z-scored rating data. The central claim is that cross-cultural differences in interpretation arise from a combination of different literal meanings and different weights on utterance cost/informativity, rather than from politeness norms alone. The behavioral findings are well supported, but the modeling conclusion is weakened by model-selection ambiguity and the absence of identifiability checks.","tokens_in":10316,"tokens_out":2845,"duration_ms":35939,"significance":"The paper addresses a genuinely important question in cross-cultural pragmatics: whether UK-US differences in intensifier interpretation are semantic, pragmatic, or both. The behavioral experiments are a real contribution: they document specific, replicable differences for \"quite\" and \"very,\" and the narrator manipulation is a sensible control for politeness-related pragmatics. The computational framework extends the RSA politeness model of Yoon et al. (2020) to a cross-cultural setting, and the authors provide code and materials. If the central modeling claim were established, the paper would push the field toward models that jointly incorporate culture-specific lexical semantics and culture-specific utterance costs. However, as it stands, the model-comparison evidence for the \"both semantic and pragmatic\" conclusion is not robust: the paper's own BIC values favor a simpler semantic-only model, and the parameter separation is not tested for identifiability. These issues are load-bearing for the paper's main claim, but they are addressable with additional analysis.","major_comments":[{"comment":"The paper states that the best-fitting model (M9) integrates cross-cultural differences in both semantic and pragmatic factors, but Table 2 does not uniformly support this. M9 has BIC 22631, while M6 (only \"quite\" threshold varies, no pragmatic parameters) has BIC 22590, and M8 (all thresholds vary, semantic only) has BIC 22656. BIC therefore favors M6 by 41 points over M9, while AIC favors M9 (22448 vs. 22485 for M8). The authors do not justify choosing AIC over BIC, nor do they report cross-validated log loss. Because the central conclusion depends on preferring M9, the paper needs to report a held-out predictive comparison or otherwise justify the model-selection criterion. Without that, the alternative reading under BIC—that only \"quite\" has a different literal threshold—is equally compatible with the reported numbers.","section":"Results / Model comparison (Table 2)"},{"comment":"The separation of semantic from pragmatic causes assumes that the twelve semantic threshold parameters and the pragmatic weights (φ_i, φ_s, C(w)) are jointly identifiable from the pooled z-scored ratings. No parameter-recovery simulations, confidence intervals, or profile likelihoods are reported. If threshold shifts and cost shifts can trade off, a fitted UK-US difference in thresholds could be partially or wholly an artifact of the optimizer. The manuscript should include synthetic recoverability checks (e.g., simulate data from known parameters and confirm the fitting procedure recovers them) and some measure of parameter uncertainty before attributing differences cleanly to \"different literal meanings\" versus \"different pragmatics.\" This issue is central to the paper's main claim.","section":"Computational Model and Appendix: Optimization Process"},{"comment":"The robustness check shows that dropping \"extremely\" changes the log loss from 11250 to 11287 on the original data, which is a larger change than dropping \"impressive\" or \"difficult.\" The authors acknowledge this sensitivity, but it is not incorporated into the model-comparison conclusion. In addition, Figure 2's \"model predictions\" are in-sample fits from parameters optimized on the same data, so the visual agreement is not independent evidence. The paper would be substantially strengthened by an out-of-sample evaluation, such as cross-validation by predicate or by participant, to demonstrate that the M9 parameterization generalizes rather than overfitting.","section":"Appendix: Robustness to dropping modifiers (Table 4)"}],"minor_comments":[{"comment":"There is a typo: \"supprting\" should be \"supporting.\"","section":"Experiment 3: Politeness Ratings"},{"comment":"The column labeled \"Log Loss\" is not defined in the text; it appears to be the negative log likelihood, and the relationship between this column and AIC/BIC should be stated explicitly.","section":"Table 2"},{"comment":"The caption \"Experiment 1: Dialogue Experiment 2: Narrator Experiment 3: Politeness\" lacks punctuation between panels and would be clearer as \"Experiment 1: Dialogue; Experiment 2: Narrator; Experiment 3: Politeness.\"","section":"Figure 1 caption"},{"comment":"The valence analysis reports p=.053 for British participants' stronger interpretation of \"very\" with positive predicates and describes this as \"most pronounced\"; given the marginal p-value, the wording should be hedged.","section":"Experiment 1: Dialogue Context"}],"recommendation":"major_revision","confidential_remarks":"The behavioral portion of the paper is solid and the research question is well motivated. The main modeling claim is not yet established by the reported analyses because the model-selection evidence is ambiguous (BIC versus AIC) and no identifiability or out-of-sample checks are provided. I see this as addressable within the manuscript's scope: cross-validated model comparison and parameter-recovery simulations would make the central claim credible. I would not recommend rejection, but the current version should not be accepted without these additions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I've read the Machino et al. paper on UK/US intensifier interpretation. Here's my take.\n\nThe behavioral core is genuinely new and useful. Three experiments with decent sample sizes, sensible controls—dialogue versus narrator framing to strip away politeness pressure, plus direct politeness ratings. They replicate the known 'quite' contrast and show it isn't just politeness norms: the narrator manipulation leaves 'quite' differences intact. That's a real empirical contribution. The RSA extension of Yoon et al. (2020) with culture-specific thresholds, weights, and costs is a reasonable first move, and they promise code and materials.\n\nThe soft spot is the model comparison, and it's load-bearing. Table 2 shows the paper's own BIC column puts M6 (only 'quite' threshold varies, semantic only) at 22590, while M9 (everything varies) is 22631—a 41-point BIC win for M6. AIC prefers M9, but they never justify choosing AIC over BIC or report cross-validated log loss. So the headline conclusion—cross-cultural differences come from both different literal meanings and different weights on utterance cost—is not actually supported by the evidence as presented. The alternative reading, that only 'quite' has a different literal threshold, fits the data equally well under BIC. The paper even acknowledges M6 gets the best BIC, then pivots to M9 anyway.\n\nThe other issues are smaller but related. The model fits are in-sample; there is no parameter recovery, no confidence intervals, and no held-out validation, so the semantic/pragmatic decomposition could be a curve-fit artifact. The robustness check shows dropping 'extremely' changes things notably, suggesting one modifier carries a lot of weight. And the model pools trials across participants without a hierarchical structure, though they flag that as a limitation.\n\nOn the plus side, the behavioral analyses use mixed-effects models with participant and scenario random effects, which is right, and they are honest about the social utility null result. The paper is clearly written and engages the literature fairly.\n\nBottom line: the experiments deserve to be published, and the modeling is a reasonable first attempt, but the central claim is not yet established. Send it to peer review; reviewers should demand cross-validation, identifiability checks, and a principled model-selection story. I'd accept with major revisions rather than reject.","headline":"New UK/US intensifier data worth having, but the model comparison claiming 'semantic plus pragmatic' is not robust once you look at the paper's own BIC column.","tokens_in":10863,"tokens_out":5897,"would_cite":true,"duration_ms":61169,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"British and American English speakers interpret 'quite' differently because of both different literal thresholds and different utterance-cost weights, not politeness norms alone.","keywords":["cross-cultural pragmatics","intensifiers","British English","American English","politeness","computational cognitive model","utterance cost","semantic variation"],"falsifier":"Fit the same model to simulated datasets generated from a known model with only threshold differences between two cultures and no cost differences; if the optimizer attributes the difference to cost or informativity weights instead, the semantic/pragmatic split is not identifiable from this design. A simpler version: compute profile likelihood or bootstrap confidence intervals for the threshold and cost parameters and check whether the culture differences survive the uncertainty.","tokens_in":9850,"feed_emoji":"🗣️","tokens_out":7864,"duration_ms":78500,"temperature":0.7,"pith_summary":"The paper tries to establish where cross-cultural differences in word interpretation come from: are British and American speakers literally assigning different meanings to words like 'quite' and 'very', or are they applying different social norms such as politeness and brevity? After three rating experiments with UK and US participants, the authors build a computational listener that reasons recursively about a speaker who balances informativity, politeness, and utterance cost, and they fit culture-specific parameters for literal thresholds and pragmatic weights. Their central claim is that both levels vary across cultures: the literal threshold for 'quite' differs, and the weight placed on utterance cost and informativity differs, while politeness ratings alone do not explain the interpretation gap. A sympathetic reader should care because this implies that successful cross-cultural communication requires recalibrating not just dictionary meanings but also expectations about why a speaker bothered to modify an utterance at all.","feed_headline":"UK-US 'quite' gap comes from meaning plus cost, not politeness","feed_subtitle":"Model shows literal thresholds and utterance-cost weights shift across cultures, changing how intensifiers land.","key_machinery":"The central object is a recursive listener model built on a speaker utility $U(w|s,\\varphi_i,\\varphi_s) = \\varphi_i U_i(w|s) + \\varphi_s U_s(w) - C(w)$, where the informativity term $U_i$ is the log-likelihood that a literal listener recovers the intended state, $U_s$ is the perceived social appropriateness of the utterance (taken from politeness ratings), and $C(w)$ penalizes adding a modifier. Culture enters through twelve double-threshold parameters fixing the literal denotation of each modifier and the unmodified baseline, plus pragmatic weights $\\varphi_i,\\varphi_s$ and cost $C(w)$. A pragmatic listener inverts the speaker's softmax production model via Bayes' rule, and the authors compare models allowing different subsets of these parameters to vary between UK and US to see which cultural variation is necessary and sufficient.","core_discovery":"In the authors' terms, cross-cultural variation in modifier interpretation arises through both semantic variation and pragmatic variation: British and American participants differ in the literal threshold values they assign to intensifiers, most clearly for 'quite', and they differ in the perceived cost of producing a modifier and the weight placed on being informative. These conclusions come from comparing models in which different subsets of parameters are allowed to vary across cultures: the best-fitting model allows both semantic thresholds and pragmatic weights to differ, and constraining either the utterance cost or the informativity weight to be shared across cultures significantly degrades fit. The model's social-utility term contributed little to fit, and perceived politeness ratings did not linearly explain the cross-cultural interpretation differences, even though politeness ratings did predict interpretation overall. A robustness check shows the model generalizes reasonably to narrator-framed data when the social term is removed, with a moderate loss increase.","pith_inferences":["If the identified cost difference is real, one testable prediction is that UK and US listeners should differ in how much extra strength they infer from the mere presence of a modifier even for novel or invented modifiers with no established semantic difference.","The joint estimation of twelve thresholds and three pragmatic weights from z-scored ratings leaves a possible trade-off between threshold shifts and cost shifts; a synthetic recoverability analysis would tell whether the semantic/pragmatic split is an artifact of the optimizer.","The paper's own note that a culture-wide social utility parameter may average over heterogeneous subcommunities suggests a hierarchical extension with regional or dialect-level politeness norms could revive the role of politeness in the model.","The model's clustering by modifier regardless of predicate, and the outliers for marked expressions like 'extremely exhausted', point toward predicate-specific social meaning rather than a single modifier-level utility as the next modeling step."],"forward_implications":["A purely semantic account is insufficient: models that let only literal thresholds vary miss the improvement gained from also letting cost and informativity weights vary.","A purely politeness-based account is insufficient: the 'quite' gap persists in narrator-framed utterances where politeness pressure is removed, and politeness ratings do not explain cross-cultural differences.","Culture-specific utterance cost means listeners infer different amounts of extra meaning from the same act of modification depending on the culture of the speaker.","Practical cross-cultural communication and machine translation should treat intensifiers as carrying both culture-specific denotations and culture-specific pragmatic weights."],"supporting_citations":[{"why":"Supplies the recursive speaker-listener utility model of politeness that this paper extends to cross-cultural variation.","marker":"Yoon et al. (2020)"},{"why":"Extends the politeness-based utility approach to social influences on indirectness, informing the model's utility structure.","marker":"Lumer and Buschmeier (2022)"},{"why":"Provides the CMA-ES optimizer used to fit the culture-specific model parameters.","marker":"Hansen (2023)"},{"why":"Provides the probabilistic programming language (memo) used to implement the recursive listener model.","marker":"Chandra, Chen, Tenenbaum, & Ragan-Kelley (2025)"},{"why":"Supplies the cooperative informativity maxim that motivates the informativity term in the speaker utility.","marker":"Grice (1975)"},{"why":"Provides the politeness-theoretic background for treating politeness as a competing speaker goal.","marker":"Brown & Levinson (1987)"},{"why":"Documents the known British/American difference in the meaning of 'quite' that motivates the semantic-variation hypothesis.","marker":"Fowler (2015)"}],"fun_headline_variants":["UK-US 'quite' gap: meaning plus cost, not politeness","Cross-cultural 'quite' shift traced to semantics and cost","Why Brits and Americans read 'quite' differently","Politeness not enough: 'quite' differs by culture and cost","Intensifier gap: literal meaning and cost, not just politeness"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The clean separation of semantic from pragmatic causes assumes that the twelve threshold parameters and the pragmatic weights are jointly identifiable from the pooled z-scored ratings, so a fitted threshold difference is not just absorbing what is really a cost difference (or vice versa).","fun_headline_variants_meta":{"raw":{"variants":["UK-US 'quite' gap: meaning plus cost, not politeness","Cross-cultural 'quite' shift traced to semantics and cost","Why Brits and Americans read 'quite' differently","Politeness not enough: 'quite' differs by culture and cost","Intensifier gap: literal meaning and cost, not just politeness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1315,"prompt_tokens":857,"completion_tokens":458,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":369}},"tokens_in":473,"tokens_out":458,"duration_ms":5195,"temperature":1.0,"reasoning_tokens":369,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:52:27.090030+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit the same model to simulated datasets generated from a known model with only threshold differences between two cultures and no cost differences; if the optimizer attributes the difference to cost or informativity weights instead, the semantic/pragmatic split is not identifiable from this design. A simpler version: compute profile likelihood or bootstrap confidence intervals for the threshold and cost parameters and check whether the culture differences survive the uncertainty.","supporting_citations":[{"cited_title":"J., Tessler, M","cited_arxiv_id":null,"evidence_quote":"Supplies the recursive speaker-listener utility model of politeness that this paper extends to cross-cultural variation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends the politeness-based utility approach to social influences on indirectness, informing the model's utility structure."},{"cited_title":"(2023).The CMA Evolution Strategy: A Tutorial","cited_arxiv_id":null,"evidence_quote":"Provides the CMA-ES optimizer used to fit the culture-specific model parameters."},{"cited_title":"B., & Ragan-Kelley, J","cited_arxiv_id":null,"evidence_quote":"Provides the probabilistic programming language (memo) used to implement the recursive listener model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the cooperative informativity maxim that motivates the informativity term in the speaker utility."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the politeness-theoretic background for treating politeness as a competing speaker goal."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the known British/American difference in the meaning of 'quite' that motivates the semantic-variation hypothesis."}],"review_version":1}