{"id":"1528ac9e-005e-4943-ba66-642ac227528f","arxiv_id":"2502.02696","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Language models' judgments on social and moral rules align most closely with younger, higher-income human annotators, raising concerns about whose values AI reflects.","lead":"This paper asks whether AI language models share the same social and moral judgments as different groups of people, and finds they most often match the views of younger, higher-income adults. It measures this by comparing model answers to a dataset of 400 rules-of-thumb annotated by 100 human raters.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Demographic alignment ranking may be an artifact of uneven per-group RoT coverage and noisy modes from groups as small as n=9.","rationale":"The reader’s weakest_assumption identifies the core problem: the demographic alignment scores rely on modes from very small subgroups and lack uncertainty quantification. My stress-test confirms this and adds precision about the sampling design: because each RoT is labeled by only 50 of the 100 annotators, the effective per-RoT sample for the smallest groups is even smaller than the group size, and the set of RoTs contributing to each group’s average may differ. That is a direct threat to the validity of comparing ADA-MetDk across groups. The paper is otherwise transparent about limitations and provides code, and the central claim is plausible, but the quantitative demographic comparison is not yet supported. The reader’s CONDITIONAL verdict is appropriate; my concern does not change it, so I recommend UNCHANGED. A concrete computational check—coverage restriction plus bootstrapping—would settle whether the observed ordering is robust.","tokens_in":11729,"tokens_out":6223,"duration_ms":58419,"concrete_test":"Recompute the Table 2 demographic ranking using only RoTs for which every demographic group has at least 10 responses, and bootstrap annotator resampling within each group to attach 95% confidence intervals to each ADA-MetDk; if the youngest/affluent groups no longer consistently rank lowest, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in §4.2 is that LMs align most closely with younger (under 40) and higher-income groups, based on the ordering of ADA-MetDk from Eq. 3. That ordering is only interpretable if each demographic group’s average is computed over a comparable set of RoTs and if the per-group mode is a stable estimate of the group’s norm. Both conditions are questionable. Each RoT is annotated by only 50 of the 100 annotators, so for the smallest groups (Lower economic class n=9, age 50-69 n=11), a given RoT may have as few as 4–5 responses, and some RoTs likely have zero responses from those groups. The paper does not report nDk for each group, nor the distribution of per-RoT group sample sizes. If group Dk and group Dj are averaged over different RoT subsets, the difference in ADA-MetDk could reflect RoT composition rather than true alignment. Additionally, with overall human Krippendorff’s α = -0.032, the modes capture no strong consensus, so the ADA-Met differences across groups may lie within sampling noise. Without confidence intervals, significance tests, or coverage tables, the headline demographic alignment claim is statistically unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether language models (LMs) reflect the social and moral norms of particular demographic groups. Using 400 rules-of-thumb (RoTs) from the Social Chemistry 101 dataset, annotated by 100 US-based Mechanical Turk workers (50 per RoT), the authors prompt 11 LMs with three prompt variants and compare LM responses with human responses using a newly introduced Absolute Distance Alignment Metric (ADA-Met). The central empirical claim is that LMs align most closely with younger (under 40) and, to a lesser extent, higher-income demographic groups, and the authors interpret this as evidence that LMs restrict the representation of minority perspectives. The paper also reports that LMs disagree less among themselves than human annotators do.","tokens_in":12067,"tokens_out":6150,"duration_ms":52739,"significance":"If the demographic alignment claim were statistically robust, the paper would make a useful contribution to the growing literature on whose opinions LMs reflect, and ADA-Met could be reused by other researchers. The strengths of the paper include a direct measurement approach with no fitted parameters, a reproducible experimental design, released code and prompts, and transparent acknowledgment of the dataset's geographic and temporal limitations. The comparison of 11 LMs across three prompt conditions is a solid resource for the community. However, the central demographic finding is currently not supported by appropriate statistical inference, which substantially tempers the significance of the conclusions.","major_comments":[{"comment":"The demographic alignment ordering reported in Figure 3 and Table 2 is computed directly from per-group averages in Eq. (3), but the paper provides no confidence intervals, standard errors, or significance tests for these averages. Because several demographic groups are very small (e.g., 9 annotators in the lower economic class, 11 in the age 50–69 bin), and because overall human agreement is negative (Krippendorff’s α = −0.032), the per-group modes used in Eq. (3) are likely noisy. The paper also does not report n_Dk (the number of RoTs contributing to each group) or the per-RoT response counts per subgroup; with 50 annotators per RoT drawn from 100 total, a group of size 9 may have zero responses for many RoTs, so the group averages may be computed over different RoT subsets. Without these statistics, the claim that 'LMs tend to align most closely with a narrow demographic range' is not statistically substantiated. I recommend bootstrapped confidence intervals for ADA-Met_Dk, permutation tests between the best and second-best groups, and a table of per-group RoT coverage.","section":"§4.2, Eq. (3)"},{"comment":"The paper's abstract and conclusion state that LMs align more closely with 'affluent backgrounds,' but Table 2 shows that three of the eleven models (Gemini 1.0 Pro, Gemini 1.5 Pro, GPT-4o) align most closely with the 0–30k income group. The tendency is thus not uniform, and the paper does not report the magnitude of the advantage of the top income group over the others. The conclusion should either be softened to reflect the mixed pattern or supported by effect sizes and tests across income groups for each model.","section":"Table 2, §4.2"},{"comment":"The aggregation of human responses uses the mode, and in case of ties the arithmetic mean of the tied ordinal options (e.g., a tie between B=1 and C=2 yields s_Hi = 1.5). This assumes both that the ordinal scale is interval and that a non-integer consensus value is a meaningful reference point for the absolute distance. The percentage bins are not evenly spaced in underlying percentages (<1%, 5–25%, 50%, 75–90%, >90%), so the equal-interval assumption is questionable. Because the same aggregation is used for every demographic group, this issue affects all demographic comparisons and should be addressed, for example by using a distributional distance that respects ordinality or by reporting sensitivity to the mapping.","section":"Appendix B.2, Eq. (3)"},{"comment":"Refusals are assigned the maximum distance of 4, and the paper notes only two 'irrelevant responses' without clarifying how the much more frequent refusals (e.g., Llama-3.1-8B refuses 20 RoTs in the zero-shot condition, Llama-3.1-405B refuses 9) are counted. Assigning distance 4 to a refusal is a defensible conservative choice, but it may bias the model-level and demographic alignment scores if refused RoTs are not uniformly distributed across groups. The authors should provide a robustness check that excludes refused RoTs or treats them as missing, and should report refusal counts per condition.","section":"Appendix B.3, Table 5"}],"minor_comments":[{"comment":"The phrase 'how well do these models making judgements' should be 'how well do these models make judgements'.","section":"Abstract"},{"comment":"The statement that the 400 RoTs are 'labeled by 50 human annotators each' could be misinterpreted as each RoT having 50 unique annotators; clarify that there are 100 total annotators and each RoT is annotated by a subset of 50 of them.","section":"Section 2"},{"comment":"The statement 'only 2 instances where an LM provided an irrelevant response' should be reconciled with the much larger number of refusals in Table 5; the reader needs a definition that distinguishes 'refusal' from 'irrelevant response' and a clear statement of how each is treated.","section":"Appendix B.3"},{"comment":"The figure caption says circle positions correspond to demographic bins rather than specific values; adding horizontal jitter or separate panels per demographic attribute would make the plot easier to read.","section":"Figure 3"},{"comment":"The paper uses 'affluent' in the conclusion, but the demographic analysis in Table 2 is based on income categories; clarify whether 'affluent' refers to income, economic class, or both, and ensure the terminology is consistent.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of cs.CL and the measurement study is worthwhile, but the authors should be encouraged to strengthen the statistical analysis. I see no signs of circularity or overclaimed novelty; the main gap is that the central demographic claim lacks inferential support. A revision that adds confidence intervals, coverage tables, and sensitivity analyses for the metric choices would address the core concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a decent empirical paper that overstates its headline result. Read it as a useful first measurement of how LM judgments track different demographic groups on social-norm questions, with the understanding that the demographic comparisons are not statistically grounded as written.\n\nWhat's actually good: the authors used a well-established dataset (Social Chemistry 101), got the annotator demographics directly from the dataset creators, ran 11 models under three prompt variants, and shipped their prompts and code. ADA-Met is just mean absolute error on an ordinal scale, but it's appropriate for the task and clearly explained. The finding that LMs agree more among themselves (α ~ 0.1-0.16) than humans agree with each other (α = -0.032) is solid and worth reporting. They also list sensible limitations: US-only annotators, temporal gap, mode-based aggregation ignoring label variation.\n\nThe soft spots are mostly concentrated in §4.2. The central claim that LMs align most with younger, affluent groups has no confidence intervals or significance tests, and the per-group modes come from groups as small as n=9 (lower economic class) and n=11 (ages 50-69). The stress-test concern is legitimate: each RoT is answered by only 50 of the 100 annotators, so for the small groups some RoTs almost certainly have zero (or very few) responses, meaning the average per-group alignment may be computed over different RoT subsets. Uneven coverage alone could produce the observed ranking. The paper doesn't report how many RoTs contribute to each group's average. Also, assigning refusals a distance of 4 is defensible, but the paper only mentions two 'irrelevant' responses while the refusal table shows Llama models refusing plenty; the interaction between refusals and missing group coverage needs careful handling. Finally, Table 2 doesn't uniformly support 'higher-income': three models align most with the 0-30k bracket, so the abstract's phrasing is stronger than the data.\n\nNone of this is fatal. The overall picture—LMs do reflect certain demographics more than others on subjective moral questions—is plausible and consistent with prior work. But the specific younger-and-affluent conclusion needs a per-group coverage table, bootstrap CIs, and ideally a test that accounts for RoT-level variation before I'd use it as evidence.\n\nVerdict: send to peer review, yes. A serious referee can push on the statistics and the paper will come out better for it. I'd cite it with caveats, but I wouldn't bring it to reading group next week.","headline":"A useful first measurement of LM alignment with demographic groups, but the headline demographic claim lacks the statistical support to be trusted as stated.","tokens_in":12490,"tokens_out":5094,"would_cite":true,"duration_ms":42948,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Language models align most closely with younger, higher-income adults when judging social and moral norms.","keywords":["language model alignment","social norms","moral norms","demographic bias","ADA-Met","ordinal alignment metric","annotator disagreement","subjective opinion representation"],"falsifier":"Resample the annotators within each demographic group with replacement, recompute each group's modal answer and the ADA-Met gap to language models, and see whether the younger and wealthier alignment advantage survives bootstrap confidence intervals; if the gaps shrink to zero, the ranking is an artifact of small-group modes. A fresh annotation study recruiting a larger sample of older and lower-income U.S. participants would test the same thing directly.","tokens_in":11543,"feed_emoji":"⚖️","tokens_out":5299,"duration_ms":47534,"temperature":0.7,"pith_summary":"The paper asks whose judgments language models reproduce when they estimate how widely a social or moral norm is shared. Prompting 11 language models with 400 rules of thumb and comparing their ordinal answers to those of 100 human annotators, it finds the models align most closely with younger adults under 40 and with higher-income backgrounds. To make this comparison, the authors introduce ADA-Met, a metric that measures the absolute distance between a model's answer and a demographic group's modal answer on a five-point scale. The paper argues that this narrow alignment raises concerns about older, lower-income, and otherwise marginalized perspectives being under-represented in AI systems.","feed_headline":"LMs side with the young and affluent on moral norms","feed_subtitle":"An 11-model test on 400 rules of thumb finds models track younger, higher-income views most closely.","key_machinery":"The load-bearing object is ADA-Met, the Absolute Distance Alignment Metric: human answers to each rule of thumb are aggregated by taking the modal option, with the arithmetic mean used in ties, mapped to ordinal positions 0 through 4, and the metric is the absolute difference between the language model's choice and that aggregate. A lower value means closer alignment, and the demographic-alignment conclusion is defined operationally by which groups have the lowest mean ADA-Met across the 400 rules of thumb. The paper also uses Krippendorff's alpha as a separate tool to compare inter-annotator agreement among humans and among language models.","core_discovery":"Using a rules-of-thumb dataset with annotator demographics, the authors compute an absolute-distance alignment score between each of 11 language models and each demographic subgroup. Across gender, age, income, marital status, education, parental status, and geographic area, the lowest ADA-Met values consistently fall in younger age brackets, primarily 18-29 and 30-39, and in higher income brackets, such as 75-100k. The paper concludes that language models tend to align most closely with a narrow demographic range, primarily younger individuals under 40 and those from affluent backgrounds. It also reports that the models agree more with one another than the human annotators agree among themselves, interpreting this as a restriction of perceived norm diversity.","pith_inferences":["The authors leave open why the alignment skew exists; a natural extension is to test whether it tracks the demographics of training or preference-tuning data, which would make the bias a data property rather than an architectural necessity.","A concrete testable extension is to repeat the study with non-U.S. annotators; the paper's own limitation note predicts the alignment pattern would shift, since all current annotations come from U.S. residents.","One could also compute ADA-Met against the full distribution of human answers rather than the modal answer, which would quantify how much opinion diversity a model compresses away.","The refusal behavior of some models is treated as maximum misalignment; an alternative design would exclude refusals and compare only answerable items, which could change the alignment ranking."],"forward_implications":["Content moderation, advice chatbots, and preference-based assistants built on this class of model will systematically over-represent younger and higher-income judgments about what is socially acceptable.","The ADA-Met framing offers a routine audit: before deployment, run a model against demographic subgroups and look for low-distance clusters that reveal whose norms it reflects.","Adding written descriptions of the answer options improves alignment for most tested models, so prompt design can partly shift which human views a model appears to hold.","Because the models agree with one another more than the human annotators do, subjective norm judgments extracted from language models are likely to be narrower than the real distribution of human opinion."],"supporting_citations":[{"why":"Supplies the Social Chemistry 101 dataset: the 400 rules of thumb and the 100 human annotators whose responses every model comparison uses.","marker":"Forbes et al. (2020)"},{"why":"Defines Krippendorff's alpha, the agreement coefficient used to compare how much humans and language models disagree among themselves.","marker":"Krippendorff (2013)"},{"why":"Provides the route by which the paper obtained annotator demographic information for the dataset, enabling the demographic alignment analysis.","marker":"Wan et al. (2023a)"},{"why":"Establishes the 'whose opinions do language models reflect' question that this paper extends from general opinions to demographic-specific social and moral norms.","marker":"Santurkar et al. (2023)"},{"why":"Makes the case that label variation matters, which the paper relies on when acknowledging that mode aggregation overlooks annotator diversity.","marker":"Plank (2022)"}],"fun_headline_variants":["LMs skew toward young, affluent views on norms","Moral norms: LMs favor younger, higher-income groups","Language models narrow moral views to youth and wealth","Norms alignment: LMs side with the privileged young","LMs reflect only young, rich perspectives on norms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results assume that the most frequent answer within a demographic group stands for that group's true norm, even though some groups have only 9 to 11 annotators and overall human disagreement is high.","fun_headline_variants_meta":{"raw":{"variants":["LMs skew toward young, affluent views on norms","Moral norms: LMs favor younger, higher-income groups","Language models narrow moral views to youth and wealth","Norms alignment: LMs side with the privileged young","LMs reflect only young, rich perspectives on norms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000792,"raw_usage":{"total_tokens":3441,"prompt_tokens":851,"completion_tokens":2590,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":467,"completion_tokens_details":{"reasoning_tokens":2512}},"tokens_in":467,"tokens_out":2590,"duration_ms":18296,"temperature":1.0,"reasoning_tokens":2512,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T11:24:46.058480+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Resample the annotators within each demographic group with replacement, recompute each group's modal answer and the ADA-Met gap to language models, and see whether the younger and wealthier alignment advantage survives bootstrap confidence intervals; if the gaps shrink to zero, the ranking is an artifact of small-group modes. A fresh annotation study recruiting a larger sample of older and lower-income U.S. participants would test the same thing directly.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the 'whose opinions do language models reflect' question that this paper extends from general opinions to demographic-specific social and moral norms."}],"review_version":1}