{"id":"e35e42bf-d73c-4c34-bda5-0d3d1ee06259","arxiv_id":"2607.05405","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Leading LLMs produce culturally appropriate health responses only 20-30% of the time on a new continuum-of-norm-adherence benchmark, with a strong bias toward Western defaults.","lead":"CCBENCH evaluates whether LLMs can infer and adapt to a user's cultural norms from conversation history rather than static demographics, using a health-query benchmark with 60 personas across six cultures. Leading models succeed only 20-30% of the time and adapt more readily when users avoid norms than when they follow them.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The Follow/Avoid asymmetry (and thus the Western-default claim) may be inflated by checklist construction that treats Western-default responses as automatic Avoid successes.","rationale":"The Reader correctly flags synthetic conversations + LLM checklists + moderate human agreement as the weakest assumption and therefore issues CONDITIONAL. That concern is real and already covers absolute CCS numbers. The more load-bearing issue for the paper's interpretive claim (Western-default bias evidenced by Follow/Avoid asymmetry) is a specific scoring asymmetry that the Reader does not isolate: Avoid items are satisfied by omission of cultural content, which is exactly what a Western-default model produces, while Follow items demand positive adaptation. This makes the headline asymmetry partly definitional rather than purely empirical. The concrete human re-scoring test would settle whether the gap survives a non-omission-based definition of competence. Because the paper already shows the gap even under the Norms upper-bound setting (where inference is removed), some genuine resistance remains; the concern therefore does not overturn CONDITIONAL but tightens what must be validated before the Western-default interpretation can be stated as strongly as it is in the abstract and conclusions. Data/code release plus the proposed human Follow/Avoid breakdown would address both the Reader's and this stress-test concern.","tokens_in":26268,"tokens_out":712,"duration_ms":6594,"concrete_test":"On a stratified sample of ≥100 responses (balanced Follow/Avoid, all six cultures), have two human annotators score each response twice: (1) against the existing LLM checklist items, and (2) with a forced-choice rubric that asks only \"Did the model actively accommodate the persona's stated preference?\" (yes/no, independent of omission). Recompute Follow Rate, Avoid Rate, and CCS under both rubrics. If the Follow–Avoid gap shrinks by >15 absolute points under the human forced-choice rubric, the asymmetry claim is partly artifactual.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's strongest claim is not merely low absolute CCS (20-30%) but the systematic Follow << Avoid asymmetry (e.g., GPT-5.2 Hist. Follow 6.5% vs Avoid 51.7%; Table 3), interpreted as evidence of a baked-in Western default that resists non-Western norms. This interpretation is load-bearing for the conclusions and abstract. It rests on the checklist pipeline (§3, F9): for every Ck(Ni) ≠ 0, GPT-5.2 generates a recommendation that the response must satisfy, then an LLM judge scores adherence. For Avoid personas the recommendation is effectively \"do not introduce or accommodate this cultural practice.\" A generic, culture-agnostic (Western-default) health answer will therefore satisfy most Avoid checklist items by simple omission, while the same answer fails nearly every Follow item that requires positive accommodation (e.g., Ramadan timing, family deferral, TCM balancing). The reported asymmetry is therefore partly mechanical: Avoid success is the default of non-adaptation; Follow success requires active cultural work. Human validation (50 responses, 62-74% agreement) does not break this out by Follow vs Avoid, so it cannot confirm that the gap is genuine cultural competence rather than an artifact of how \"correct Avoid\" is operationalized. If the gap shrinks substantially under a human-only or re-balanced scoring regime, the Western-default narrative weakens.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces CCBENCH, a domain-agnostic framework that evaluates LLM cultural competence by treating culture as a continuum of norm-adherence states (Follow/Avoid/Neutral) derived from binary value vectors, rather than binary demographic belonging. These states are revealed only implicitly via multi-turn conversation histories. As a case study it constructs CCBENCH-Health: 60 Mosaica-grounded personas across six cultures, each with 18 background dialogues, evaluated on 52 real-forum health queries (3,120 interactions). Five leading models are tested under four prompting regimes (no context, history, Culture-CoT, explicit norms). Headline results are that even the best models reach only 20–30 % culturally appropriate responses (CCS), Culture-CoT yields only 3–5 % gains, and a large Follow ≪ Avoid asymmetry appears (e.g., GPT-5.2 history: Follow 6.5 % vs Avoid 51.7 %), interpreted as a baked-in Western default that resists non-Western norms; Afghan performance is especially low (avg CCS 8.8 %).","tokens_in":26678,"tokens_out":893,"duration_ms":16998,"significance":"If the measurement is reliable, the work supplies a needed stress-test for cultural competence in high-stakes domains and shows that current LLMs remain far from equitable adaptation even when norms are made explicit. Strengths include the theoretically grounded continuum of adherence (Eqs. 1–3), the scale of the resource, the multi-model comparison, the attempt at human validation of both conversation filtering and checklist scoring, and the open pipeline that can be extended beyond health. The Follow/Avoid asymmetry and culture-specific disparities (esp. Afghan) would be actionable findings for alignment research if they survive tighter validation of the scoring regime.","major_comments":[{"comment":"Table 3 and §5: the central Western-default claim rests on the large Follow ≪ Avoid gap. Under the checklist pipeline (§3, prompt F9), an Avoid recommendation is essentially “do not introduce or accommodate this practice.” A generic, culture-agnostic (Western-default) answer therefore satisfies most Avoid items by simple omission, while the identical answer fails nearly every Follow item that requires positive accommodation (Ramadan timing, family deferral, TCM balancing, etc.). The reported asymmetry is therefore partly mechanical. Human validation of the judge (50 responses) is not stratified by Follow vs Avoid, so it cannot confirm that the gap is genuine competence rather than an artifact of how “correct Avoid” is operationalized. A re-analysis that (a) reports human agreement separately for Follow and Avoid items or (b) re-scores a stratified sample with human-only judgments is requ","section":null},{"comment":"§3 and A.3: checklist satisfaction is scored by GPT-5.2, the same model family used both to generate the background histories and as one of the evaluated systems. Human–LLM agreement on the 50-response sample is only 62–64 % (human–human 74 %). This level of agreement, combined with possible self-preference, is too low to underwrite the absolute CCS numbers (20–30 %) and the culture-wise rankings that appear in the abstract and conclusions. At minimum the paper must (i) report inter-annotator statistics broken down by culture and by Follow/Avoid and (ii) release the 50 annotated examples so that the community can assess judge reliability.","section":null},{"comment":"§3.2–3.3 and Table 2: the claim that the generated histories constitute valid “implicit” cultural signals rests on LLM-based filtering that removes only explicit identity statements while retaining greetings, dietary remarks, etc. The validation (A.3) shows high agreement that culture is not named, but does not establish that the retained cues are the same cues real users would produce or that models actually attend to them rather than to surface style. Because the entire benchmark is synthetic, a small human-authored or human-edited history subset (or a comparison against real multi-turn health dialogues) is needed to bound the ecological-validity risk.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that this is a genuine methodological step past binary culture benchmarks. They treat culture as a continuum of Follow/Avoid/Neutral states derived from value vectors, surface those states only through conversation history, and then score health answers with persona-specific checklists. That setup, plus the 60-persona / 3,120-interaction CCBENCH-Health resource grounded in Mosaica profiles, is new and usable. The multi-model results are consistent: absolute cultural competence stays low (roughly 20–30% even for the best models), Culture-CoT helps only a few points, and explicit norm injection still leaves Follow rates weak. The Afghan numbers are especially low. That pattern is reproducible enough to take seriously.\n\nWhat they do well is the pipeline discipline: value-to-norm derivation, filtering for explicit culture reveals, stratification of real forum queries, and the Follow vs Avoid decomposition. The asymmetry is real in their numbers (e.g., GPT-5.2 history setting: Follow ~6.5% vs Avoid ~52%). Human spot-checks on filtering and checklist scoring are only moderate (62–74%), but they are present and the multi-model consistency gives some external grounding.\n\nThe soft spot that matters is the stress-test point: Avoid success is largely automatic. A generic Western-default answer satisfies most “do not accommodate X” checklist items by omission, while Follow items require positive cultural work. Their human validation is not broken out by Follow vs Avoid, so we cannot yet tell how much of the Western-default narrative is measurement artifact versus genuine resistance. That does not kill the paper; it just means the strongest interpretive claim needs a cleaner human or re-balanced scoring check. Synthetic conversation fidelity and LLM-as-judge self-preference are secondary, acknowledged risks, not load-bearing collapses.\n\nThis is for people working on cultural alignment, medical NLP, and personalization audits. It is not trivia; the continuum framing and the resource are worth engaging. I would send it to peer review. Data/code release and a Follow/Avoid-stratified human validation would make the claims tighter, but the core contribution already clears the bar for serious referee time.","headline":"Useful continuum framing and a real Follow/Avoid gap on a carefully built health benchmark; the Western-default story is partly mechanical because Avoid success is mostly non-adaptation.","tokens_in":27221,"tokens_out":533,"would_cite":true,"duration_ms":5942,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Even the best LLMs give culturally appropriate health advice only 20-30% of the time when cultural norms appear only as implicit signals in prior conversation.","keywords":["cultural competence","large language models","health AI","implicit norms","norm adherence continuum","stereotype resistance","persona evaluation","chain-of-thought"],"falsifier":"Collect a held-out set of real multi-turn conversations from users of the six cultures who spontaneously reveal the same norm-adherence patterns, have the same five models answer the identical health queries, and have independent human cultural experts score the responses; if human competence rates substantially exceed the reported 20-30% band, the central claim fails.","tokens_in":27193,"feed_emoji":"🏥","tokens_out":867,"duration_ms":14515,"temperature":0.7,"pith_summary":"This paper argues that cultural competence for AI means inferring and adapting to a user's fluid, implicitly signaled norm adherence rather than treating culture as a binary demographic label. It introduces CCBENCH, a framework that builds personas as continua of Follow/Avoid/Neutral states on theoretically grounded norms, then tests models on whether they use conversational history to calibrate high-stakes answers. Instantiated as CCBENCH-Health (60 personas across six cultures, 52 real forum health queries, 3,120 interactions), it shows leading models top out at roughly 20-30% appropriate responses. Performance is asymmetrically higher when personas avoid cultural norms than when they follow them, and remains low even under explicit norm prompts or cultural chain-of-thought. A sympathetic reader cares because health advice that ignores or stereotypes cultural cues can erode trust and safety for the large non-Western user base already relying on these systems.","feed_headline":"LLMs get cultural health advice right only 20-30% of the time","feed_subtitle":"Models stick to Western defaults even when prior chat quietly signals a user follows different norms.","key_machinery":"CCBENCH: a domain-agnostic evaluation framework that represents culture as ternary norm-adherence states (Follow/Avoid/Neutral) derived from value averages, embeds those states as implicit behavioral cues inside multi-turn conversation histories, and scores model answers with persona-specific checklists generated for each non-neutral norm.","core_discovery":"Leading LLMs achieve culturally appropriate responses to health queries only 20-30% of the time when cultural position must be inferred from conversational history; they systematically succeed more often when personas avoid cultural norms than when they follow them, revealing a persistent Western-default bias that explicit norm lists and cultural chain-of-thought only modestly reduce.","pith_inferences":["The same continuum-of-adherence design could be ported to legal, educational, or financial advice domains where implicit value signaling is equally common.","The Follow/Avoid asymmetry predicts that models will look more 'competent' on progressive or secular personas than on traditional ones even when both are equally well-signaled.","Because checklist scoring itself relies on an LLM, residual Western bias may be double-counted; a pure human-scored subset would be a natural next measurement.","Cultures whose markers are rarer in pretraining (illustrated by the Afghan floor) will remain the hardest cases until data mixtures change."],"forward_implications":["Health chatbots that ignore implicit cultural cues will systematically under-serve users who follow non-Western norms, especially in Afghan and similar underrepresented contexts.","Stereotype-resistance metrics can look artificially high simply because models already omit non-Western content by default; true sensitivity requires active accommodation.","Explicit norm lists and cultural chain-of-thought are insufficient upper bounds; training or alignment data must encode continuum-style, multi-norm identity.","Topic effects are secondary to cultural distinctiveness: the same health category shows large competence gaps across cultures.","Communication-style cues are sometimes easier for models to match than practice-based norms, but only when those styles align with the Western default."],"fun_headline_variants":["LLMs get cultural health advice right only 20-30% of the time","Top models match health norms just 20-30% when cues stay implicit","LLMs succeed more when users avoid cultural norms than follow them","Cultural health replies from LLMs hit 20-30% even with chat history","Western defaults beat cultural signals for LLM health advice"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That LLM-simulated conversation histories plus LLM-generated checklists, after light filtering and only moderate human agreement, are faithful enough proxies for real human cultural signaling and for human judgments of cultural competence.","fun_headline_variants_meta":{"raw":{"variants":["LLMs get cultural health advice right only 20-30% of the time","Top models match health norms just 20-30% when cues stay implicit","LLMs succeed more when users avoid cultural norms than follow them","Cultural health replies from LLMs hit 20-30% even with chat history","Western defaults beat cultural signals for LLM health advice"]},"model":"grok-4.5","effort":"low","cost_usd":0.006924,"raw_usage":{"total_tokens":1700,"prompt_tokens":819,"num_sources_used":0,"completion_tokens":77,"cost_in_usd_ticks":69240000,"prompt_tokens_details":{"text_tokens":819,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":804,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":819,"tokens_out":77,"duration_ms":7325,"temperature":1.0,"reasoning_tokens":804,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T14:29:22.408979+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Collect a held-out set of real multi-turn conversations from users of the six cultures who spontaneously reveal the same norm-adherence patterns, have the same five models answer the identical health queries, and have independent human cultural experts score the responses; if human competence rates substantially exceed the reported 20-30% band, the central claim fails.","supporting_citations":[],"review_version":1}