Pith. sign in

REVIEW 4 major objections 5 minor 35 references

How Inclusively do LMs Perceive Social and Moral Norms?

T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Language models align most closely with younger, higher-income adults when judging social and moral norms.

desk verdict A useful first measurement of LM alignment with demographic groups, but the headline demographic claim lacks the statistical support to be trusted as stated. read the letter →

arxiv 2502.02696 v2 pith:ATCN55RE submitted 2025-02-04 cs.CL

classification cs.CL
keywords languagemodelalignmentsocialnormsmoraldemographicbiasADA-Metordinalmetricannotatordisagreementsubjectiveopinionrepresentation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whose judgments language models reproduce when they estimate how widely a social or moral norm is shared. Prompting 11 language models with 400 rules of thumb and comparing their ordinal answers to those of 100 human annotators, it finds the models align most closely with younger adults under 40 and with higher-income backgrounds. To make this comparison, the authors introduce ADA-Met, a metric that measures the absolute distance between a model's answer and a demographic group's modal answer on a five-point scale. The paper argues that this narrow alignment raises concerns about older, lower-income, and otherwise marginalized perspectives being under-represented in AI systems.

What carries the argument

The load-bearing object is ADA-Met, the Absolute Distance Alignment Metric: human answers to each rule of thumb are aggregated by taking the modal option, with the arithmetic mean used in ties, mapped to ordinal positions 0 through 4, and the metric is the absolute difference between the language model's choice and that aggregate. A lower value means closer alignment, and the demographic-alignment conclusion is defined operationally by which groups have the lowest mean ADA-Met across the 400 rules of thumb. The paper also uses Krippendorff's alpha as a separate tool to compare inter-annotator agreement among humans and among language models.

What would settle it

Resample the annotators within each demographic group with replacement, recompute each group's modal answer and the ADA-Met gap to language models, and see whether the younger and wealthier alignment advantage survives bootstrap confidence intervals; if the gaps shrink to zero, the ranking is an artifact of small-group modes. A fresh annotation study recruiting a larger sample of older and lower-income U.S. participants would test the same thing directly.

Watch

Extended reading notes

Core claim

Using a rules-of-thumb dataset with annotator demographics, the authors compute an absolute-distance alignment score between each of 11 language models and each demographic subgroup. Across gender, age, income, marital status, education, parental status, and geographic area, the lowest ADA-Met values consistently fall in younger age brackets, primarily 18-29 and 30-39, and in higher income brackets, such as 75-100k. The paper concludes that language models tend to align most closely with a narrow demographic range, primarily younger individuals under 40 and those from affluent backgrounds. It also reports that the models agree more with one another than the human annotators agree among themselves, interpreting this as a restriction of perceived norm diversity.

Load-bearing premise

The results assume that the most frequent answer within a demographic group stands for that group's true norm, even though some groups have only 9 to 11 annotators and overall human disagreement is high.

Editorial extensions

If this is right

  • Content moderation, advice chatbots, and preference-based assistants built on this class of model will systematically over-represent younger and higher-income judgments about what is socially acceptable.
  • The ADA-Met framing offers a routine audit: before deployment, run a model against demographic subgroups and look for low-distance clusters that reveal whose norms it reflects.
  • Adding written descriptions of the answer options improves alignment for most tested models, so prompt design can partly shift which human views a model appears to hold.
  • Because the models agree with one another more than the human annotators do, subjective norm judgments extracted from language models are likely to be narrower than the real distribution of human opinion.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors leave open why the alignment skew exists; a natural extension is to test whether it tracks the demographics of training or preference-tuning data, which would make the bias a data property rather than an architectural necessity.
  • A concrete testable extension is to repeat the study with non-U.S. annotators; the paper's own limitation note predicts the alignment pattern would shift, since all current annotations come from U.S. residents.
  • One could also compute ADA-Met against the full distribution of human answers rather than the modal answer, which would quantify how much opinion diversity a model compresses away.
  • The refusal behavior of some models is treated as maximum misalignment; an alternative design would exclude refusals and compare only answerable items, which could change the alignment ranking.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper investigates whether language models (LMs) reflect the social and moral norms of particular demographic groups. Using 400 rules-of-thumb (RoTs) from the Social Chemistry 101 dataset, annotated by 100 US-based Mechanical Turk workers (50 per RoT), the authors prompt 11 LMs with three prompt variants and compare LM responses with human responses using a newly introduced Absolute Distance Alignment Metric (ADA-Met). The central empirical claim is that LMs align most closely with younger (under 40) and, to a lesser extent, higher-income demographic groups, and the authors interpret this as evidence that LMs restrict the representation of minority perspectives. The paper also reports that LMs disagree less among themselves than human annotators do.

Significance. If the demographic alignment claim were statistically robust, the paper would make a useful contribution to the growing literature on whose opinions LMs reflect, and ADA-Met could be reused by other researchers. The strengths of the paper include a direct measurement approach with no fitted parameters, a reproducible experimental design, released code and prompts, and transparent acknowledgment of the dataset's geographic and temporal limitations. The comparison of 11 LMs across three prompt conditions is a solid resource for the community. However, the central demographic finding is currently not supported by appropriate statistical inference, which substantially tempers the significance of the conclusions.

major comments (4)
  1. [§4.2, Eq. (3)] The demographic alignment ordering reported in Figure 3 and Table 2 is computed directly from per-group averages in Eq. (3), but the paper provides no confidence intervals, standard errors, or significance tests for these averages. Because several demographic groups are very small (e.g., 9 annotators in the lower economic class, 11 in the age 50–69 bin), and because overall human agreement is negative (Krippendorff’s α = −0.032), the per-group modes used in Eq. (3) are likely noisy. The paper also does not report n_Dk (the number of RoTs contributing to each group) or the per-RoT response counts per subgroup; with 50 annotators per RoT drawn from 100 total, a group of size 9 may have zero responses for many RoTs, so the group averages may be computed over different RoT subsets. Without these statistics, the claim that 'LMs tend to align most closely with a narrow demographic range' is not statistically substantiated. I recommend bootstrapped confidence intervals for ADA-Met_Dk, permutation tests between the best and second-best groups, and a table of per-group RoT coverage.
  2. [Table 2, §4.2] The paper's abstract and conclusion state that LMs align more closely with 'affluent backgrounds,' but Table 2 shows that three of the eleven models (Gemini 1.0 Pro, Gemini 1.5 Pro, GPT-4o) align most closely with the 0–30k income group. The tendency is thus not uniform, and the paper does not report the magnitude of the advantage of the top income group over the others. The conclusion should either be softened to reflect the mixed pattern or supported by effect sizes and tests across income groups for each model.
  3. [Appendix B.2, Eq. (3)] The aggregation of human responses uses the mode, and in case of ties the arithmetic mean of the tied ordinal options (e.g., a tie between B=1 and C=2 yields s_Hi = 1.5). This assumes both that the ordinal scale is interval and that a non-integer consensus value is a meaningful reference point for the absolute distance. The percentage bins are not evenly spaced in underlying percentages (<1%, 5–25%, 50%, 75–90%, >90%), so the equal-interval assumption is questionable. Because the same aggregation is used for every demographic group, this issue affects all demographic comparisons and should be addressed, for example by using a distributional distance that respects ordinality or by reporting sensitivity to the mapping.
  4. [Appendix B.3, Table 5] Refusals are assigned the maximum distance of 4, and the paper notes only two 'irrelevant responses' without clarifying how the much more frequent refusals (e.g., Llama-3.1-8B refuses 20 RoTs in the zero-shot condition, Llama-3.1-405B refuses 9) are counted. Assigning distance 4 to a refusal is a defensible conservative choice, but it may bias the model-level and demographic alignment scores if refused RoTs are not uniformly distributed across groups. The authors should provide a robustness check that excludes refused RoTs or treats them as missing, and should report refusal counts per condition.
minor comments (5)
  1. [Abstract] The phrase 'how well do these models making judgements' should be 'how well do these models make judgements'.
  2. [Section 2] The statement that the 400 RoTs are 'labeled by 50 human annotators each' could be misinterpreted as each RoT having 50 unique annotators; clarify that there are 100 total annotators and each RoT is annotated by a subset of 50 of them.
  3. [Appendix B.3] The statement 'only 2 instances where an LM provided an irrelevant response' should be reconciled with the much larger number of refusals in Table 5; the reader needs a definition that distinguishes 'refusal' from 'irrelevant response' and a clear statement of how each is treated.
  4. [Figure 3] The figure caption says circle positions correspond to demographic bins rather than specific values; adding horizontal jitter or separate panels per demographic attribute would make the plot easier to read.
  5. [Table 3] The paper uses 'affluent' in the conclusion, but the demographic analysis in Table 2 is based on income categories; clarify whether 'affluent' refers to income, economic class, or both, and ensure the terminology is consistent.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the demographic-alignment claim is a direct empirical measurement with no fitted parameter, self-citation chain, or derivation that reduces to its inputs.

full rationale

The paper's derivation chain is entirely empirical. Human annotations and demographics come from Social Chemistry 101 / the dataset creators (via external prior work by Wan et al., not the present authors). LM outputs are obtained by prompting 11 models with fixed prompts and extracting ordinal answers. ADA-Met (Eq. 1) is defined as the absolute distance between an LM's ordinal response and the mode (or tie-mean) of a human group's responses; Eq. 3 averages these distances per demographic group. The headline result—that younger and higher-income groups show lower ADA-MetDk—is the direct numerical output of these measurement steps. There are no fitted parameters, no 'prediction' that was an input to a fit, no self-citation used as load-bearing justification, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in by citation. The choice to represent a group's norm by the mode/mean is an operationalization and could be debated, but it does not make the ranking equal to its own definition. ADA-Met resembles a mean absolute error on an ordinal scale, but naming or repackaging a metric does not make the empirical finding circular. Concerns about small subgroup sizes, uneven RoT coverage, and missing confidence intervals are statistical-validity issues, not circularity. Therefore no circular step can be quoted; the correct finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No numerical free parameters are fitted; the ordinal mapping A=0..E=4 and the mode/mean aggregation are modeling choices, not parameters fitted to data. No new physical or conceptual entities are postulated; ADA-Met is a metric, not an entity. The axioms listed are the load-bearing assumptions behind the demographic alignment measurements.

assumptions (4)
  • domain assumption The mode (or mean in ties) of ordinal human responses represents the collective norm of a demographic group.
    Used in Eq. (1)-(3) to define the human reference for ADA-Met. Humans have low agreement (alpha=-0.032), so the mode may be unstable for small groups.
  • domain assumption The five ordinal bins are equally spaced such that absolute differences in mapped values (A=0 to E=4) are meaningful.
    The bins (<1%, 5-25%, 50%, 75-90%, >90%) are not equal in probability width; ADA-Met treats a move from A to B as identical to D to E. Invoked in Eq. (1).
  • domain assumption A refusal or irrelevant LM response can be coded as distance 4, the maximum misalignment.
    Appendix B.3 assigns refusals the maximum distance; this is a strong assumption that affects models with many refusals (Llama-3.1-8B, 405B).
  • domain assumption Demographic information obtained privately from dataset creators is accurate and complete.
    Section 2; demographics are not in the public dataset and were provided via personal communication.

how reviews work

0 comments
Cite this review

Pith. "Pith review of How Inclusively do LMs Perceive Social and Moral Norms?." pith.science (2026). https://pith.science/paper/ATCN55RE

@misc{pith2026250202696,
  author       = {Pith},
  title        = {Pith review of: How Inclusively do LMs Perceive Social and Moral Norms?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ATCN55RE}},
  note         = {Machine review of arXiv:2502.02696}
}
read the original abstract

This paper discusses and contains offensive content. Language models (LMs) are used in decision-making systems and as interactive assistants. However, how well do these models making judgements align with the diversity of human values, particularly regarding social and moral norms? In this work, we investigate how inclusively LMs perceive norms across demographic groups (e.g., gender, age, and income). We prompt 11 LMs on rules-of-thumb (RoTs) and compare their outputs with the existing responses of 100 human annotators. We introduce the Absolute Distance Alignment Metric (ADA-Met) to quantify alignment on ordinal questions. We find notable disparities in LM responses, with younger, higher-income groups showing closer alignment, raising concerns about the representation of marginalized perspectives. Our findings highlight the importance of further efforts to make LMs more inclusive of diverse human values. The code and prompts are available on GitHub under the CC BY-NC 4.0 license.

Figures

Figures reproduced from arXiv: 2502.02696 by the authors.

Figure 1
Figure 1. Rule of thumb definition, anticipated agree [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Experimental pipeline of creating prompts, prompting LMs, extracting answers, and comparing LM [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. LM alignment with demographic groups based [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Absolute distance alignment matrices allow for the comparison between demographic groups and LMs. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: depicts the distribution of human-LM ADA-Met distances for different LMs. We observe that responses from Arctic and Llama-3.1-405B are mostly 0 or 1 option away from the human choice. This indicates a strong agreement between humans and LMs. F Refusal to Answer To use …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

35 extracted references · 11 canonical work pages

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Yang, Dylan Hadfield-Menell, Gillian K

    Aparna Balagopalan, David Madras, David H. Yang, Dylan Hadfield-Menell, Gillian K. Hadfield, and Marzyeh Ghassemi. 2023. https://doi.org/10.1126/sciadv.abq0701 Judging facts, judging norms: Training machine learning models to judge humans requires a modified approach to labeling data . Science Advances, 9(19):eabq0701

  4. [4]

    Su Lin Blodgett, Solon Barocas, Hal Daum \'e III, and Hanna Wallach. 2020. https://doi.org/10.18653/v1/2020.acl-main.485 Language (technology) is power: A critical survey of `` bias '' in NLP . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454--5476, Online. Association for Computational Linguistics

  5. [5]

    Esin Durmus, Karina Nguyen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. 2024. http://arxiv.org/abs/2306.16388 Towards measuring the representation of subj...

  6. [6]

    Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi

    Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.48 Social chemistry 101: Learning to reason about social and moral norms . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 653--670, Online. Association for Computational Linguistics

  7. [7]

    Frank, Manuel Cebrian, Galen Pickard, and Iyad Rahwan

    Morgan R. Frank, Manuel Cebrian, Galen Pickard, and Iyad Rahwan. 2017. https://doi.org/10.1371/journal.pone.0177385 Validating bayesian truth serum in large-scale online human experiments . PLOS ONE, 12(5):1--13

  8. [8]

    Gemini Team et al. 2024. http://arxiv.org/abs/2312.11805 Gemini: A family of highly capable multimodal models

Show all 35 references
  1. [9]

    Kivlichan, Rachel Rosen, and Lucy Vasserman

    Nitesh Goyal, Ian D. Kivlichan, Rachel Rosen, and Lucy Vasserman. 2022. https://doi.org/10.1145/3555088 Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation . Proceedings of the ACM on Human-Computer Interaction, 6(CSCW2)

  2. [10]

    Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023. http://arxiv.org/abs/2301.01768 The political ideology of conversational ai: Converging evidence on chatgpt's pro-environmental, left-libertarian orientation

  3. [11]

    Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...

  4. [12]

    Krippendorff

    K. Krippendorff. 2013. https://books.google.com/books?id=s_yqFXnGgjQC Content analysis: An introduction to its methodology . SAGE Publications

  5. [13]

    Yiwei Luo, Dallas Card, and Dan Jurafsky. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.296 Detecting stance in media on global warming . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3296--3315, Online. Association for Computational L...

  6. [14]

    Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016. https://doi.org/10.18653/v1/N16-1098 A corpus and cloze evaluation for deeper understanding of commonsense stories . In Proceedings of the 2...

  7. [15]

    Fabio Motoki, Valdemar Pinho Neto, and Victor Rodrigues. 2023. More human than human: Measuring chatgpt political bias. Public Choice, pages 1--21

  8. [16]

    Tarek Naous, Michael Ryan, Alan Ritter, and Wei Xu. 2024. https://doi.org/10.18653/v1/2024.acl-long.862 Having beer after prayer? measuring cultural bias in large language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volu...

  9. [17]

    OpenAI . 2023a. Gpt-4 technical report. Technical report, OpenAI. Available at https://doi.org/10.48550/arXiv.2303.08774

  10. [18]

    i'm fully who i am

    Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023. https://doi.org/10.1145/3593013.3594078 “i'm fully who i am”: Towards centering transgender and non-binary voices to measure biases in open languag...

  11. [19]

    Luger, Tharindu Ranasinghe, Ashiqur R

    Deepak Pandita, Tharindu Cyril Weerasooriya, Sujan Dutta, Sarah K. Luger, Tharindu Ranasinghe, Ashiqur R. KhudaBukhsh, Marcos Zampieri, and Christopher M. Homan. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.296 Rater cohesion and quality from a vicarious perspective ....

  12. [20]

    Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan

    Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Benjamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, D...

  13. [21]

    Barbara Plank. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.731 The `` problem '' of human label variation: On ground truth in data, modeling and evaluation . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671--10682, Ab...

  14. [22]

    Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. 2021. https://doi.org/10.18653/v1/2021.law-1.14 On releasing annotator-level labels and information in datasets . In Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Re...

  15. [23]

    Drazen Prelec. 2004. https://doi.org/10.1126/science.1102081 A bayesian truth serum for subjective data . Science, 306(5695):462--466

  16. [24]

    S. A. R. Team . 2024. https://www.snowflake.com/blog/arctic-open-efficient-foundation-language-models-snowflake/ Snowflake arctic: The best llm for enterprise ai — efficiently intelligent, truly open

  17. [25]

    Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning (ICML), page 1244. JMLR.org

  18. [26]

    Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022. https://doi.org/10.18653/v1/2022.naacl-main.431 Annotators with attitudes: How annotator beliefs and identities bias toxic language detection . In Proceedings of the 2022 Conference...

  19. [27]

    Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R

    Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna M. Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Pe...

  20. [28]

    Yuan Sui, Mengyu Zhou, Mingjie Zhou, Shi Han, and Dongmei Zhang. 2024. https://doi.org/10.1145/3616855.3635752 Table meets llm: Can large language models understand structured table data? a benchmark and empirical study . In Proceedings of the 17th ACM International Conference...

  21. [29]

    The Mosaic Research Team . 2024. https://www.databricks.com/blog/introducing-dbrx-new-state-art-open-llm Introducing dbrx: A new state-of-the-art open llm

  22. [30]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://arxiv.org/abs/2302.1397...

  23. [31]

    Ruyuan Wan, Jaehyung Kim, and Dongyeop Kang. 2023 a . Everyone’s voice matters: Quantifying annotation disagreement using demographic information. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 14523--14530. AAAI

  24. [32]

    Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023 b . https://doi.org/10.18653/v1/2023.findings-emnlp.243 `` kelly is a warm person, joseph is a role model '' : Gender biases in LLM -generated reference letters . In Findings of the Associat...

  25. [33]

    Su Wang, Greg Durrett, and Katrin Erk. 2018. https://doi.org/10.18653/v1/N18-2049 Modeling semantic plausibility by injecting world knowledge . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language...

  26. [34]

    Tharindu Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher Homan, and Ashiqur KhudaBukhsh. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.713 Vicarious offense and noise audit of offensive speech classifiers: Unifying human and machine disagreemen...

  27. [35]

    Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2017. https://doi.org/10.18653/v1/D17-1323 Men also like shopping: Reducing gender bias amplification using corpus-level constraints . In Proceedings of the 2017 Conference on Empirical Methods in Natur...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.