{"id":"6252e614-dbaf-487d-b431-389eef4dc798","arxiv_id":"2501.07751","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Cultural alignment scores of GPT-4o vary with prompting style (classification, chain-of-thought, scenarios), supporting a context-dependent, bidirectional view of AI cultural alignment.","lead":"This paper argues that AI cultural alignment should be seen as bidirectional, with values depending on how users interact with the system, and uses a GPT-4o experiment to show that different prompting styles change measured alignment scores. A generalist reader might care because it challenges the common practice of baking static survey values into AI systems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The case study's cross-condition differences are statistically unsupported: all Table 1 confidence intervals overlap and no significance tests are reported, so the central empirical claim may be noise.","rationale":"The paper is best read as a position piece with an illustrative case study: its conceptual thesis—that cultural alignment is bidirectional and context-dependent—is plausible and supported by prior work (e.g., Shen et al. 2024; Röttger et al. 2024; Ge et al. 2024). The original empirical contribution is Table 1, which is supposed to demonstrate that interaction patterns change AI cultural alignment. That demonstration is the load-bearing evidence for the paper's central empirical statement. The most direct threat is statistical: the reported bootstrap confidence intervals overlap between conditions in every country, and no significance tests are provided. The reader's weakest-assumption diagnosis (filtering of unclassifiable outputs) is a real but secondary threat; it is a potential source of bias in the point estimates, whereas the significance concern attacks the reported estimates themselves even under the paper's own filtering rule. This is not a disagreement with the conceptual argument, and the paper's candid acknowledgment of limitations (single model, single task) is a credit. Rather, the empirical illustration is not yet strong enough to carry the weight assigned to it in Section 3. A pairwise significance test and a robustness check of the filtering rule would settle the matter. If those fail, the case study should be presented as illustrative only, not as evidence for the ordering. The verdict remains CONDITIONAL, as the reader proposed, because the conceptual contribution can survive the loss of the empirical illustration but the paper should then temper its claims accordingly.","tokens_in":6021,"tokens_out":6633,"duration_ms":62532,"concrete_test":"Using the original dataset, compute pairwise permutation or bootstrap tests of the Wasserstein-score difference between each pair of conditions (Classification vs CoT, Classification vs Scenarios, CoT vs Scenarios) for each country, with 10,000 resamples and a family-wise error correction. In parallel, recompute the scenario-condition scores without dropping unclassifiable outputs—for example by treating 'unclassifiable' as a third response category or by reweighting using the stance extractor's class probabilities. If the adjusted 95% confidence intervals for the differences include zero, or if the condition ordering changes, then Table 1 does not establish that interaction patterns shape cultural alignment.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3's central empirical claim—that 'interaction patterns fundamentally shape how cultural alignment manifests'—rests entirely on the ordering of Wasserstein similarity scores in Table 1. But in each of the four countries, the bootstrap 95% confidence intervals for the three conditions overlap substantially (e.g., US: Classification 0.66 [0.62,0.71], CoT 0.71 [0.67,0.75], Scenarios 0.70 [0.66,0.74]; China: 0.60 [0.55,0.66], 0.68 [0.63,0.73], 0.65 [0.60,0.70]). No pairwise significance tests or multiple-comparison corrections are reported, yet the text calls the variation 'significant.' Because the effective sample is only 72 binary-choice questions per country (Appendix A.1), the observed differences of 0.02–0.08 may easily be sampling noise. A separate measurement-pipeline threat compounds this: the scenario condition required GPT-4 stance extraction and dropped 28–40% of outputs as unclassifiable; if those outputs are non-random, the scenario scores are biased. Both issues attack the same inference: the data do not yet show that interaction patterns, rather than noise or filtering artifacts, drive the reported ordering.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that cultural alignment of AI systems should be reframed as a bidirectional process: rather than embedding static, survey-derived values into models, alignment should account for the context of specific AI systems and the ways humans structure their interactions. The central empirical support is a GPT-4o case study that measures Wasserstein similarity between model answer distributions and human survey responses across four countries, under three prompting conditions: direct classification, chain-of-thought, and open-ended scenarios. The authors report that chain-of-thought prompting yields the highest alignment scores in all four countries and conclude that interaction patterns fundamentally shape how cultural alignment manifests.","tokens_in":6260,"tokens_out":4777,"duration_ms":47268,"significance":"If the empirical claim were robust, the paper would make a useful conceptual contribution by drawing attention to the interaction-dependent nature of cultural alignment and by proposing a bidirectional framing that goes beyond static value repositories. The paper is clearly written, the literature review is relevant, and the experimental protocol is described in sufficient detail to be reproduced, including the exact prompts, sampling parameters (temperature 0.7, top-p 1), and the use of bootstrap confidence intervals. However, the empirical demonstration is currently fragile: the confidence intervals for the three conditions overlap in every country, no significance tests are reported, and the scenario condition discards a large fraction of outputs as unclassifiable. These issues directly undermine the claimed evidence for the paper's central empirical conclusion.","major_comments":[{"comment":"The claim that 'Table 1 highlights significant variation in alignment metrics across interaction types' is not supported by the reported statistics. For every country, the bootstrap 95% confidence intervals for the three conditions overlap substantially; for example, in the US, classification is 0.66 [0.62,0.71] and CoT is 0.71 [0.67,0.75], so the intervals overlap on [0.67,0.71]. The same pattern is visible for China, Japan, and India. No pairwise significance tests, p-values, or multiple-comparison corrections are reported, and the effective sample is only 72 binary-choice questions per country. The text should either report appropriate significance tests (e.g., bootstrap hypothesis tests with a defined null of no difference) or temper the wording to 'descriptive variation' rather than 'significant variation.'","section":"Section 3, Table 1"},{"comment":"The scenario condition drops 28–40% of outputs as unclassifiable before computing the Wasserstein score, and the paper does not analyze whether these dropped outputs differ systematically from the classified ones. If unclassifiable responses are more frequent for certain question types or for less decisive model outputs, then the scenario score is a conditional statistic, and the comparison to the classification and CoT conditions may reflect the filtering rule rather than genuine cultural alignment. To support the ordering in Table 1, the authors should provide a sensitivity analysis, such as reweighting the remaining responses, computing bounds that assume extreme outcomes for the dropped outputs, or reporting scores only for the subset of questions with low unclassifiable rates.","section":"Appendix A.1 and Section 3"},{"comment":"The stance extraction for CoT and scenario responses is performed by GPT-4, with manual validation on only 50 samples per setting. Given that the scenario condition has a very high unclassifiable rate (up to 40%), 50 samples are insufficient to rule out systematic misclassification that correlates with the outcome. The paper should report inter-annotator agreement (e.g., Cohen's kappa) between the two authors and between the authors and GPT-4, and it should describe the distribution of unclassifiable outputs across the 10 scenarios and across countries to help assess whether the high drop rate is concentrated in particular conditions.","section":"Appendix A.1 (stance extraction)"}],"minor_comments":[{"comment":"The word 'boostrapping' is misspelled in the description of the confidence intervals; it should be 'bootstrapping.'","section":"Section 3"},{"comment":"The sentence 'for each of the 10 scenarios outlined in Appendix' is missing the subsection reference; it should read 'outlined in Appendix B.5.'","section":"Appendix A.1"},{"comment":"The first sentence contains a stray space: 'be come' should be 'become.'","section":"Abstract"},{"comment":"In the caption, the downward arrow after 'Percentage of Unclassifiable Outputs' is nonstandard; a short note explaining that lower percentages are better would improve clarity.","section":"Table 1"},{"comment":"The paper does not state whether multiple runs with different random seeds were averaged; since temperature is set to 0.7, sampling is stochastic, and the reported scores presumably come from a single set of API calls. Reporting variance across independent runs would strengthen the reliability of the numbers.","section":"Appendix A.1"}],"recommendation":"major_revision","confidential_remarks":"This is a position paper with an exploratory empirical demonstration, and it was presented at a workshop. The central conceptual argument does not depend solely on the case study, so the paper is not beyond repair. However, as it stands, the empirical evidence advertised in the abstract—that AI alignment 'depends on how humans structure their interactions'—is not statistically supported by the reported analysis, and the filtering bias is a serious threat to internal validity. A revision that adds explicit significance tests, a sensitivity analysis for dropped outputs, and appropriate hedging of the claims would be within the manuscript's scope. The paper's self-citations (Kwok et al., Barez and Torr) appear as related work and are not load-bearing, so there is no circularity concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. The conceptual pitch—cultural alignment is bidirectional and interaction-relative—is reasonable but not new; Shen et al. (2024) already makes that argument, and prompt-sensitivity of LLM opinions is known from Röttger et al. The paper's own empirical support is shakier than the prose suggests: Table 1 shows overlapping bootstrap intervals for every country, yet the text calls the variation 'significant' and reports no significance tests. The scenario condition also discards 28–40% of outputs before scoring, and if that filtering is not random, the ordering of scores is suspect.\n\nThat said, the paper does some things well. It compiles a coherent set of references on why static value surveys are imperfect proxies, running from cultural psychology to LLM evaluation. It actually runs an experiment—a GPT-4o case study with three interaction conditions on GlobalOpinionQA—which is more than many position pieces do. The stance extraction step is checked by manual evaluation (98% agreement on a 50-item sample), which is honest. And it states its limitations plainly: one model, one task, one cultural theory.\n\nThe soft spots are central, not cosmetic. The core empirical claim about interaction patterns shaping cultural alignment rests on comparisons that are statistically indistinguishable from noise. The 72 binary questions per country is a small n, and the differences are 0.02–0.08 with wide intervals. The filtering issue compounds this: up to 40% of scenario outputs are unclassifiable, and the paper does not address how this could bias the Wasserstein scores. So the stress-test note is right: the data do not yet show what the abstract says they show.\n\nWho gains from this? People building cultural alignment benchmarks should read it as a reminder that the prompt format and user framing can move the numbers, even if this paper doesn't prove exactly how. It would be a decent workshop paper, which I think it is. For a top-venue submission, the experiment would need post-hoc tests, a treatment of missing/uncodable responses, and a more careful wording of 'significant'.\n\nMy take: I'd send it to review if it crossed my desk, because the question matters and the fixes are concrete. I wouldn't cite its empirical results yet, but the framing is worth keeping in mind.","headline":"Reasonable position piece but its own GPT-4o case study is statistically too weak to support the empirical claim; the bidirectional framing is already in Shen et al. (2024).","tokens_in":6791,"tokens_out":2622,"would_cite":false,"duration_ms":26947,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that cultural alignment between humans and AI is bidirectional and context-dependent, and demonstrates with a GPT-4o case study that the prompt structure a user chooses changes how closely the model's answers match human…","keywords":["cultural alignment","large language models","bidirectional human-AI alignment","prompting","chain-of-thought","Wasserstein similarity","GlobalOpinionQA"],"falsifier":"Recompute the Wasserstein similarity for the scenario condition without filtering unclassifiable outputs, either by counting them as a third outcome or assigning them to options at random; if the chain-of-thought advantage over classification disappears, the reported ordering is an artifact of the filtering rule rather than of interaction structure.","tokens_in":5805,"feed_emoji":"🌐","tokens_out":6685,"duration_ms":59964,"temperature":0.7,"pith_summary":"Cultural alignment is usually treated as a one-way process: take values from static surveys and embed them in an AI. This paper argues the opposite: alignment is bidirectional, because both the human values in play and the AI's expressed values shift with the context of the interaction. A GPT-4o case study across the US, China, Japan, and India shows that the same survey questions produce different levels of match between model answers and human answers depending on whether the model is asked to classify, reason step by step, or respond in an open-ended scenario. Chain-of-thought prompting achieves the highest similarity scores, while direct classification scores lowest, and scenario prompts produce many unclassifiable outputs. The authors conclude that how people interact with an AI is part of cultural alignment, not a neutral channel around it.","feed_headline":"Prompt style shifts GPT-4o's cultural alignment scores","feed_subtitle":"Chain-of-thought beats direct prompts in all four countries, suggesting alignment is shaped by interaction.","key_machinery":"Two tools carry the argument. The first is the comparison of three interaction structures: single-token classification, chain-of-thought reasoning, and open-ended scenario responses drawn from ten everyday formats such as phone surveys, editorials, and radio interviews, all prompted with 'respond as someone from [country] would.' The second is the Wasserstein similarity score, a distribution-distance measure between GPT-4o's answers and human answers from 72 binary-choice questions in the GlobalOpinionQA dataset, with bootstrap confidence intervals. Open-ended outputs are converted to stances by a separate GPT-4 classifier that the authors manually validated at 98% accuracy on 50 sampled responses per setting.","core_discovery":"The paper's central claim is that interaction patterns fundamentally shape how cultural alignment manifests. In the authors' experiments, GPT-4o's Wasserstein similarity to human survey responses varies across three prompting conditions in every country tested, with chain-of-thought reaching the highest average similarity and direct classification the lowest. The paper treats this variation as evidence that a model cannot be assigned a fixed cultural alignment score, and that alignment should be modeled as a bidirectional process in which human-imposed interaction structures co-determine the AI's cultural behavior.","pith_inferences":["If the pattern holds across other models, prompt-condition rankings could be used as a quick diagnostic: a model whose alignment ranking flips across conditions may be restating prompt stereotypes rather than stable cultural knowledge.","The high rate of unclassifiable scenario outputs suggests open-ended interaction exposes ambiguity that multiple-choice questions hide; a testable extension would analyze what kinds of answers get dropped and whether those answers cluster by country or by question topic.","A stronger test of bidirectionality would swap the direction of influence—showing that interacting with an AI changes human survey answers over time—which the paper's cross-sectional case study does not measure."],"forward_implications":["A single cultural alignment score for an AI model is not well defined; the score depends on the interaction format used to measure it.","Models compared for cultural alignment should be tested under matched interaction structures, or the comparison can reflect the prompt rather than the model.","Free-form scenario prompts produce 28–40% unclassifiable outputs, so multiple-choice evaluations may overstate how confidently models match human cultural opinions.","Design choices in AI products—how questions are posed and how users can respond—become part of the cultural alignment problem, not just implementation details.","Static surveys remain useful but are insufficient as a complete measure of cultural alignment."],"supporting_citations":[{"why":"It supplies the GlobalOpinionQA dataset and the Wasserstein similarity methodology used to compare model and human answer distributions.","marker":"Durmus et al. (2023)"},{"why":"It supplies the three interaction prompt designs and the stance extraction approach for open-ended outputs.","marker":"Röttger et al. (2024)"},{"why":"It frames alignment as bidirectional, the conceptual claim this paper argues for.","marker":"Shen et al. (2024)"},{"why":"It provides evidence that cultural preferences for AI vary across groups and concrete use cases.","marker":"Ge et al. (2024)"},{"why":"It is the GPT-4o system card, the model used in the case study.","marker":"Hurst et al. (2024)"},{"why":"It is the GPT-4 technical report; the paper uses GPT-4 for stance extraction to avoid contamination.","marker":"Achiam et al. (2023)"},{"why":"It shows different LLM personalities accentuate different human values, supporting context-dependent value expression.","marker":"Kirk et al. (2024)"}],"fun_headline_variants":["Different prompts, different cultural alignment for GPT-4o","GPT-4o's alignment scores change with prompt style","Bidirectional alignment: interaction shapes AI cultural scores","Cultural alignment is not fixed: it's interaction-dependent","GPT-4o's cultural alignment varies with prompt design"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison of alignment scores across prompt conditions is only valid if dropping unclassifiable outputs—28–40% in the scenario condition—does not remove systematically different kinds of answers from one condition than another.","fun_headline_variants_meta":{"raw":{"variants":["Different prompts, different cultural alignment for GPT-4o","GPT-4o's alignment scores change with prompt style","Bidirectional alignment: interaction shapes AI cultural scores","Cultural alignment is not fixed: it's interaction-dependent","GPT-4o's cultural alignment varies with prompt design"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001193,"raw_usage":{"total_tokens":4835,"prompt_tokens":769,"completion_tokens":4066,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":385,"completion_tokens_details":{"reasoning_tokens":3988}},"tokens_in":385,"tokens_out":4066,"duration_ms":31792,"temperature":1.0,"reasoning_tokens":3988,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:36:16.907745+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the Wasserstein similarity for the scenario condition without filtering unclassifiable outputs, either by counting them as a third outcome or assigning them to options at random; if the chain-of-thought advantage over classification disappears, the reported ordering is an artifact of the filtering rule rather than of interaction structure.","supporting_citations":[],"review_version":1}