Pith. sign in

REVIEW 4 major objections 5 minor 56 references

Words of Warmth: Trust and Sociability Norms for over 26k English Words

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Words of Warmth provides manual trust and sociability scores for over 26,000 English words, with warmth taken from the stronger of the two ratings.

desk verdict Valuable T/S norms for 26k words, but the derived warmth score is presented as manually measured; the paper needs reframing before the stronger claims stand. read the letter →

arxiv 2506.03993 v1 pith:5462B7JU submitted 2025-06-04 cs.CL cs.CY

classification cs.CLcs.CY
keywords warmthnormstrustsociabilitystereotypecontentsocialcognitioncrowdsourcedlexiconageofacquisitionmediaanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces Words of Warmth, a manually built set of trust, sociability, and warmth association scores for over 26,000 English words, with warmth derived from whichever of the trust or sociability rating has the larger absolute value. It claims these crowdsourced scores are highly reliable, with split-half correlations above 0.94. The motivation is that warmth and competence are core dimensions of social perception, and word-level norms let researchers study stereotypes, child development, and public discourse at scale. The paper demonstrates the resource by tracking how children acquire warmth- and competence-related words at different ages and by measuring how social groups are described with warmer or colder words in social media.

What carries the argument

The carrying mechanism is a crowdsourced annotation pipeline with gold control questions and split-half reliability measurement. Annotators rate each word's trustworthiness and sociability on a seven-point scale from -3 to +3 after receiving detailed instructions, and a warmth score is assigned per word by taking whichever of trust or sociability has the larger absolute value. The age-of-acquisition analysis combines these scores with an external dataset of age-of-acquisition ratings; the stereotype case studies use direct lookup of target words and co-occurrence scores computed as the percentage of high-warmth words minus low-warmth words in tweets mentioning a target.

What would settle it

A concrete check is to compare Words of Warmth scores against an independent set of human warmth ratings obtained with a direct 'warmth' question: if the max-of-trust-and-sociability rule does not reproduce those rankings on a shared set of words, the operationalization of warmth would not match the psychological construct. A second check is to see whether the lexicon's warmth scores reproduce the rankings from the existing smaller manual warmth dictionary; a low correlation would indicate that the trust and sociability annotations do not capture warmth as social psychologists define it.

Watch

Extended reading notes

Core claim

The central discovery is a large-scale repository of word-trust and word-sociability association norms, each word rated on a -3 to +3 scale by roughly eight to eleven annotators, with a combined warmth lexicon formed by taking the trust or sociability score with the larger absolute value. The lexicon covers over 26,000 English words selected from a valence, arousal, and dominance resource, deliberately excluding words with neutral valence. The paper shows that repeating the annotation process yields nearly identical scores, with split-half reliability around 0.94 to 0.97, and demonstrates that the resource can reveal how children acquire warmth, competence, trust, and sociability words at different ages, and how social groups are described with warmer or colder words in social media.

Load-bearing premise

The load-bearing assumption is that trust and sociability are the two core components of warmth and that a short crowdsourced rating question on each dimension captures those psychological constructs; if the ratings measure something else, the warmth scores and the developmental and stereotype conclusions built on them do not follow.

Editorial extensions

If this is right

  • Researchers can now measure warmth, trust, and sociability in text at scale for over 26,000 English words, complementing existing competence norms.
  • The high split-half reliability implies that crowdsourced annotation at this scale can yield stable association scores, supporting the creation of similar lexicons for other languages.
  • Age-of-acquisition patterns suggest that warmth-polar words, especially sociability words, are acquired earlier than competence-polar words, consistent with the primacy of valence and with sociability being developmentally early.
  • Co-term warmth-competence plots enable stereotype research, showing gender stereotypes, ingroup-outgroup patterns, and differences in how social groups are discussed.
  • The resource supports longitudinal tracking of warmth and competence in discourse toward targets such as social groups, professions, and pronouns.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the lexicon's reliability extends to construct validity, it could serve as a scalable proxy for stereotype content in large text corpora, complementing implicit association measures, though the paper does not establish this link.
  • The choice to define warmth as the larger-magnitude trust or sociability score is one of several possible aggregation rules; because the paper releases the separate trust and sociability scores, researchers can test whether averaging or another combination better matches established stereotype content results.
  • The same annotation pipeline could be applied to other languages, and cross-language warmth norms would let researchers test whether the trust-sociability split is universal or culturally shaped, a direction the paper lists as future work.
  • The stability of co-term scores across years and morphological variants suggests the method could monitor stereotype change over time, such as shifts in discourse about immigrants or LGBTQ groups.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces Words of Warmth, a set of crowd-sourced English word-association norms for trust (T), sociability (S), and a derived warmth (W) score for about 26k words, together with competence scores taken from the NRC VAD lexicon. The authors report high split-half reliabilities for T, S, and W, analyze the age of acquisition of WCTS words using Kuperman et al.'s norms, and present case studies on social groups, genders, pronouns, and professions using direct lexicon lookup and co-occurrence analysis over a large tweet corpus. The central contribution is the resource itself, which is publicly released.

Significance. If the trust and sociability lexicons are valid, this is a valuable contribution: it is far larger than existing manually compiled warmth-related lexicons (compare the 341 words in Nicolas et al. 2021) and is released openly. The paper's clear strengths are the scale of annotation, the use of gold questions and popup feedback for quality control, the high split-half reliabilities reported in Table 1, and the transparent description of the annotation procedure. However, the warmth score is not independently measured: it is deterministically derived from T and S via a max-absolute rule, and no external validation of W is provided. The developmental and stereotype analyses are mostly descriptive and lack inferential statistics. The resource may still be useful as a T/S lexicon, but the paper's framing of W as a directly measured construct needs revision or additional validation.

major comments (4)
  1. [Section 3, bullet 6; Abstract; FAQ (Appendix A.1)] The abstract claims 'manually derived word-warmth associations', but the paper describes no direct warmth rating task. Warmth is defined as a deterministic function: W equals T when |T|>|S|, S when |S|>|T|, and the common score when they are equal. This is load-bearing because W is the headline construct and is used throughout the developmental and stereotype analyses. The FAQ even notes that politicians are often seen as untrustworthy yet sociable; under the max-absolute rule, whenever |S|>|T| such a target receives a high positive W and is coded as warm, collapsing the very divergence the FAQ says is important. Please either collect direct warmth ratings, validate W against direct warmth judgments or an external warmth norm (e.g., the 341-word Nicolas et al. 2021 lexicon), or revise the abstract and Section 7 so that W is explicitly described as an algorithmically derived score rather than a manually rated construct.
  2. [Table 1] The split-half reliability reported for warmth in Table 1 does not by itself validate W as a measure of warmth. Because W is computed deterministically from T and S, its split-half reliability is inherited from the reliable T and S ratings; the correlation between split-half-derived W scores would be high even if the max-absolute rule mis-specified the warmth construct. Reliability is necessary but not sufficient for construct validity. Please report the SHR of W as a derived quantity and provide construct validation evidence, such as correlations with independent warmth ratings or with warmth terms from existing dictionaries.
  3. [Section 5, Figures 2a-d] The developmental claims are presented without inferential statistics. The stream graphs are descriptive, yet the discussion states that the patterns are 'consistent with the primacy of valence hypothesis' and that sociability is 'more important than trust' in early years. No confidence intervals, permutation tests, or model-based estimates accompany these claims. In addition, the term selection step in Section 3, bullet 1 intentionally excluded neutrally valenced words (valence between -0.2 and 0.2), which may bias the proportions of polar versus neutral words at each age. Please add appropriate statistical tests and a discussion of the sampling frame, or explicitly reframe the developmental section as an illustrative case study rather than a substantive claim.
  4. [Section 6, Figures 3-5] The stereotype case studies report point estimates from co-term percentages without measures of uncertainty, confidence intervals, or multiple-comparison control. For example, comparisons in Figures 3(b), 4(b), and 5 are made by visual inspection of points whose sampling variability is not quantified. Since these case studies are used to demonstrate the utility of the resource, I ask that they either be accompanied by bootstrap confidence intervals or permutation tests, or be explicitly labeled as descriptive demonstrations rather than confirmatory findings.
minor comments (5)
  1. [Appendix A.3, Table 2 caption] The caption says 'anxiety-association score' for a random sample from Words of Warmth; this should read 'warmth-association score' or similar.
  2. [Section 3, bullet 6] In the categorical labeling description, 'neither sociable nor unsociability' should be 'neither sociable nor unsociable'.
  3. [Section 1] The acronyms WCST and WCTS are both used in the same section ('WCST assessment capabilities' and later 'WCTS capabilities'); please standardize to one form.
  4. [References] The reference to 'Hilton and V on Hippel, 1996' has erroneous spacing and capitalization; it should be 'Hilton and Von Hippel'.
  5. [Appendix A.5] The text says Figure 12 is 'described in Section 5', but Figure 12 is first referenced in Section 3, bullet 6; please correct the cross-reference.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular reasoning in the central derivation; trust and sociability norms come from independent crowd annotations, and the warmth score is a transparent deterministic combination rather than a hidden reuse of the output.

full rationale

The core resource is built from independent crowdsourced ratings for trust and sociability, with warmth explicitly defined as a deterministic function of those two scores: the paper states that if the absolute T score is greater than the absolute S score, W equals T, and if the absolute S score is greater, W equals S. This is a stated operationalization, not a circular derivation: W does not feed back into T or S, and no claim is made that T or S are predicted from W. The age-of-acquisition analysis uses an external dataset (Kuperman et al., 2012), and the stereotype case studies use independent tweets, so neither reduces to the lexicon's own inputs. The paper does cite the author's own NRC VAD lexicons for term selection and competence scores, and its own TUSC corpus and co-term formula for downstream analyses, but these are data sources or application tools, not load-bearing support for the trust and sociability norms themselves. The abstract's phrase 'manually derived word--warmth associations' overstates the fact that warmth scores are algorithmically derived from the manual T and S scores; however, this is a construct-validity and reporting concern, not circularity. No step in the derivation chain is equivalent by construction to its input.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central contribution is a dataset. It uses several hand-chosen thresholds and assumptions from psychology, but introduces no new theoretical entities. The main free parameters are cut-offs for term selection and for polar categories; none are fitted to predict an external outcome, so the circularity burden is low.

free parameters (5)
  • Valence exclusion threshold = -0.2 to 0.2
    Terms with NRC VAD valence scores between -0.2 and 0.2 are excluded; this choice shapes the word sample and may bias WTS distributions (Section 3, bullet 1).
  • Polar word threshold for WST = 1.5
    In the age-of-acquisition analysis, words with absolute score >= 1.5 are considered high or low; the threshold is arbitrary and not justified theoretically (Section 5).
  • Polar word threshold for competence = 0.33
    Competence scores from NRC VAD v2 are split at 0.33 and -0.33 into high, neutral, low; this is an ad hoc cut point (Section 5).
  • Gold question accuracy cutoff = 80%
    Annotators with accuracy below 80% on gold questions are excluded; a design choice with no sensitivity analysis (Section 3, bullet 5).
  • Number of annotations per word = 9 for sociability and warmth, 12 for trust
    Target of 9 responses per word, later 12 for trust; chosen for cost rather than derived from reliability analysis (Section 3, bullet 4).
assumptions (4)
  • domain assumption Trust and sociability are two distinct components of the warmth dimension of social cognition.
    Adopted from Abele et al. (2016) and Koch et al. (2024); the entire annotation task is built on this two-factor model (Section 1 and questionnaires).
  • domain assumption The NRC VAD lexicon's valence, arousal, and dominance scores are valid and the selected word list is appropriate for warm/trust/sociability annotation.
    Terms are sampled from NRC VAD v2 (Mohammad 2018, 2025); the paper does not compare with other term sources (Section 3).
  • domain assumption The Kuperman et al. (2012) age-of-acquisition ratings are accurate for the subset of words shared with the lexicon.
    Used without validation for the developmental analysis (Section 5).
  • domain assumption The co-occurrence formula (percentage of high-W minus low-W words in a target corpus) captures the affective tone of text.
    Taken from Teodorescu and Mohammad (2023) and Turney (2002); the paper notes other formulas are possible (Section 6, bullet 2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Words of Warmth: Trust and Sociability Norms for over 26k English Words." pith.science (2026). https://pith.science/paper/5462B7JU

@misc{pith2026250603993,
  author       = {Pith},
  title        = {Pith review of: Words of Warmth: Trust and Sociability Norms for over 26k English Words},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5462B7JU}},
  note         = {Machine review of arXiv:2506.03993}
}
read the original abstract

Social psychologists have shown that Warmth (W) and Competence (C) are the primary dimensions along which we assess other people and groups. These dimensions impact various aspects of our lives from social competence and emotion regulation to success in the work place and how we view the world. More recent work has started to explore how these dimensions develop, why they have developed, and what they constitute. Of particular note, is the finding that warmth has two distinct components: Trust (T) and Sociability (S). In this work, we introduce Words of Warmth, the first large-scale repository of manually derived word--warmth (as well as word--trust and word--sociability) associations for over 26k English words. We show that the associations are highly reliable. We use the lexicons to study the rate at which children acquire WCTS words with age. Finally, we show that the lexicon enables a wide variety of bias and stereotype research through case studies on various target entities. Words of Warmth is freely available at: http://saifmohammad.com/warmth.html

Figures

Figures reproduced from arXiv: 2506.03993 by the authors.

Figure 1
Figure 1. Distribution of terms in Words of Warmth: percentage and number of terms associated with each class. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Stream charts of the percentages of high-, low-, and neutral WCTS words acquired in ages 3 to 17. (The [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. W–C plots for social groups [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: W–C plots for various gender terms. grandmother) compared to son, wife, and hus￾band. The co-term plot also shows certain ad￾ditional related terms for comparison such as person and human. • Importantly, we see clear gender stereotypes re￾flected in these scores with m…
Figure 5
Figure 5. Figure 5: Co-terms W–C plot for tweets by Canadians [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Trust Questionnaire: Detailed instructions. [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]
Figure 7
Figure 7. Figure 7: Trust Questionnaire: Sample question [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: Sociability Questionnaire: Detailed instructions. [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]
Figure 9
Figure 9. Figure 9: Sociability Questionnaire: Sample question. [PITH_FULL_IMAGE:figures/full_fig_p018_9.png]
Figure 10
Figure 10. Figure 10: Trust Questionnaire: Examples [PITH_FULL_IMAGE:figures/full_fig_p019_10.png]
Figure 11
Figure 11. Figure 11: Sociability Questionnaire: Examples [PITH_FULL_IMAGE:figures/full_fig_p019_11.png]
Figure 13
Figure 13. Figure 13: Stacked bar charts showing a break down of [PITH_FULL_IMAGE:figures/full_fig_p020_13.png]
Figure 14
Figure 14. Figure 14: Direct and co-term W–C plots for various pronouns. [PITH_FULL_IMAGE:figures/full_fig_p021_14.png]
Figure 15
Figure 15. Figure 15: Direct and co-term W–C plots for various professions. [PITH_FULL_IMAGE:figures/full_fig_p021_15.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

56 extracted references · 41 canonical work pages

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Andrea E Abele, Nicole Hauke, Kim Peters, Eva Louvet, Aleksandra Szymkow, and Yanping Duan. 2016. Facets of the fundamental content dimensions: Agency with competence and assertiveness—communion with warmth and morality. Frontiers in psychology, 7:1810

  4. [4]

    Inna Altschul, Shawna J Lee, and Elizabeth T Gershoff. 2016. Hugs, not hits: Warmth and spanking as predictors of child social competence. Journal of Marriage and Family, 78(3):695--714

  5. [5]

    Alejandro Ariza-Casabona, Wolfgang S Schmeisser-Nieto, Montserrat Nofre, Mariona Taul \'e , Enrique Amig \'o , Berta Chulvi, and Paolo Rosso. 2022. Overview of detests at iberlef 2022: Detection and classification of racial stereotypes in spanish. Procesamiento del lenguaje natural, 69:217--228

  6. [6]

    Alexander Baines, Lidia Gruia, Gail Collyer-Hoar, and Elisa Rubegni. 2024. Playgrounds and prejudices: Exploring biases in generative ai for children. In Proceedings of the 23rd Annual ACM Interaction Design and Children Conference, pages 839--843

  7. [7]

    Su Lin Blodgett, Solon Barocas, Hal Daum \'e III, and Hanna Wallach. 2020. Language (technology) is power: A critical survey of" bias" in nlp. arXiv preprint arXiv:2005.14050

  8. [8]

    Galen V Bodenhausen, Sonia K Kang, and Destiny Peery. 2012. Social categorization and the perception of social groups. The Sage handbook of social cognition, pages 311--329

Show all 56 references
  1. [9]

    Cristina Bosco, Viviana Patti, Simona Frenda, Alessandra Teresa Cignarella, Marinella Paciello, and Francesca D’Errico. 2023. Detecting racial stereotypes: An italian social media corpus where psychology meets nlp. Information Processing & Management, 60(1):103118

  2. [10]

    Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183--186

  3. [11]

    Amy JC Cuddy, Susan T Fiske, and Peter Glick. 2007. The bias map: behaviors from intergroup affect and stereotypes. Journal of personality and social psychology, 92(4):631

  4. [12]

    Amy JC Cuddy, Peter Glick, and Anna Beninger. 2011. The dynamics of warmth and competence judgments, and their outcomes in organizations. Research in organizational behavior, 31:73--98

  5. [13]

    Federica Durante and Susan T Fiske. 2017. How social-class stereotypes maintain inequality. Current opinion in psychology, 18:43--48

  6. [14]

    Adar B Eisenbruch and Max M Krasnow. 2022. Why warmth matters more than competence: A new evolutionary approach. Perspectives on Psychological Science, 17(6):1604--1623

  7. [15]

    Susan Fiske, Amy Cuddy, Peter Glick, and Jun Xu. 2002. https://doi.org/10.1037/0022-3514.82.6.878 A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition . Journal of Personality and Social Psychology, 82:878--902

  8. [16]

    Susan T Fiske. 2018. Stereotype content: Warmth and competence endure. Current directions in psychological science, 27(2):67--73

  9. [17]

    Susan T Fiske and Federica Durante. 2016. Stereotype content across cultures. Handbook of advances in culture and psychology, 6:209--258

  10. [18]

    Susan T Fiske, Federica Durante, et al. 2014. Never trust a politician? collective distrust, relational accountability, and voter response. Power, politics, and paranoia: Why people are suspicious of their leaders, pages 91--105

  11. [19]

    Kathleen Fraser, Svetlana Kiritchenko, and Isar Nejadgholi. 2024. https://doi.org/10.18653/v1/2024.starsem-1.2 How does stereotype content differ across data sources? In Proceedings of the 13th Joint Conference on Lexical and Computational Semantics (*SEM 2024), pages 18--34, ...

  12. [20]

    Dmitry Grigoryev, Susan T Fiske, and Anastasia Batkhina. 2019. Mapping ethnic stereotypes and their antecedents in russia: The stereotype content model. Frontiers in psychology, 10:1643

  13. [21]

    James L Hilton and William Von Hippel. 1996. Stereotypes. Annual review of psychology, 47(1):237--271

  14. [22]

    Ewa Kacewicz, James W Pennebaker, Matthew Davis, Moongee Jeon, and Arthur C Graesser. 2014. Pronoun use reflects standings in social hierarchies. Journal of Language and Social Psychology, 33(2):125--143

  15. [23]

    Svetlana Kiritchenko and Saif Mohammad. 2018. https://doi.org/10.18653/v1/S18-2005 Examining gender and race bias in two hundred sentiment analysis systems . In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, pages 43--53, New Orleans, Louis...

  16. [24]

    Alex Koch, Austin Smith, Susan T Fiske, Andrea E Abele, Naomi Ellemers, and Vincent Yzerbyt. 2024. Validating a brief measure of four facets of social evaluation. Behavior Research Methods, 56(8):8521--8539

  17. [25]

    Melissa A Koenig and Catharine H Echols. 2003. Infants' understanding of false labeling events: The referential roles of words and the speakers who use them. Cognition, 87(3):179--208

  18. [26]

    Hadas Kotek, Rikker Dockum, and David Sun. 2023. Gender bias and stereotypes in large language models. In Proceedings of the ACM collective intelligence conference, pages 12--24

  19. [27]

    Victor Kuperman, Hans Stadthagen-Gonzalez, and Marc Brysbaert. 2012. Age-of-acquisition ratings for 30,000 english words. Behavior research methods, 44:978--990

  20. [28]

    Kevin MacDonald. 1992. Warmth as a developmental construct: An evolutionary analysis. Child development, 63(4):753--773

  21. [29]

    Ivy W Maina, Tanisha D Belton, Sara Ginzberg, Ajit Singh, and Tiffani J Johnson. 2018. A decade of studying implicit racial/ethnic bias in healthcare providers using the implicit association test. Social science & medicine, 199:219--229

  22. [30]

    Saif Mohammad. 2018. https://doi.org/10.18653/v1/P18-1017 Obtaining reliable human ratings of valence, arousal, and dominance for 20,000 E nglish words . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1...

  23. [31]

    Mohammad

    Saif M. Mohammad. 2022. Ethics sheet for automatic emotion recognition and sentiment analysis. Computational Linguistics, 48(2):239--278

  24. [32]

    Mohammad

    Saif M. Mohammad. 2023. Best practices in the creation and use of emotion lexicons. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, Dubrovnik, Croatia. Association for Computational Linguistics

  25. [33]

    Mohammad

    Saif M. Mohammad. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.910 W orry W ords: Norms of anxiety association for over 44k E nglish words . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 16261--16278, Miami, Florida, USA....

  26. [34]

    Mohammad

    Saif M. Mohammad. 2025. https://arxiv.org/abs/2503.23547 NRC VAD Lexicon v2: Norms for Valence, Arousal, and Dominance for over 55k English Terms . arXiv preprint arXiv:2503.23547

  27. [35]

    Agnes Moors, Jan De Houwer, Dirk Hermans, Sabine Wanmaker, Kevin Van Schie, Anne-Laura Van Harmelen, Maarten De Schryver, Jeffrey De Winne, and Marc Brysbaert. 2013. Norms of valence, arousal, dominance, and age of acquisition for 4,300 dutch words. Behavior research methods, ...

  28. [36]

    Robert Morabito, Sangmitra Madhusudan, Tyler McDonald, and Ali Emami. 2024. Stop! benchmarking large language models with sensitivity testing on offensive progressions. arXiv preprint arXiv:2409.13843

  29. [37]

    Gandalf Nicolas, Xuechunzi Bai, and Susan T Fiske. 2021. Comprehensive stereotype content dictionaries using a semi-automated method. European Journal of Social Psychology, 51(1):178--196

  30. [38]

    Brian A Nosek, Anthony G Greenwald, and Mahzarin R Banaji. 2005. Understanding and using the implicit association test: Ii. method variables and construct validity. Personality and Social Psychology Bulletin, 31(2):166--180

  31. [39]

    James W Pennebaker. 2011. The secret life of pronouns. New Scientist, 211(2828):42--45

  32. [40]

    Gina Roussos and Yarrow Dunham. 2016. https://doi.org/https://doi.org/10.1016/j.jecp.2015.08.009 The development of stereotype content: The use of warmth and competence in assessing social groups . Journal of Experimental Child Psychology, 141:133--144

  33. [41]

    Javier S \'a nchez-Junquera, Berta Chulvi, Paolo Rosso, and Simone Paolo Ponzetto. 2021. How do you speak about immigrants? taxonomy and stereoimmigrants dataset for identifying stereotypes about immigrants. Applied Sciences, 11(8):3610

  34. [42]

    Wolfgang S Schmeisser-Nieto, Alessandra Teresa Cignarella, Tom Bourgeade, Simona Frenda, Alejandro Ariza-Casabona, Mario Laurent, Paolo Giovanni Cicirelli, Andrea Marra, Giuseppe Corbelli, Farah Benamara, et al. 2024. Stereohoax: a multilingual corpus of racial hoaxes and soci...

  35. [43]

    Marie Gustafsson Send \'e n, Torun Lindholm, and Sverker Sikstr \"o m. 2014. Biases in news media as reflected by personal pronouns in evaluative contexts. Social Psychology

  36. [44]

    Jillian K Swencionis, Cydney H Dupree, and Susan T Fiske. 2017. Warmth-competence tradeoffs in impression management across race and social-class divides. Journal of Social Issues, 73(1):175--191

  37. [45]

    Yi Chern Tan and L Elisa Celis. 2019. Assessing social and intersectional biases in contextualized word representations. Advances in neural information processing systems, 32

  38. [46]

    Daniela Teodorescu and Saif Mohammad. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.271 Evaluating emotion arcs across languages: Bridging the global divide in sentiment analysis . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 4124--41...

  39. [47]

    Mike Thelwall. 2018. Gender bias in sentiment analysis. Online Information Review, 42(1):45--57

  40. [48]

    Kristen Swan Tummeltshammer, Rachel Wu, David M Sobel, and Natasha Z Kirkham. 2014. Infants track the reliability of potential informants. Psychological science, 25(9):1730--1738

  41. [49]

    Peter Turney. 2002. https://doi.org/10.3115/1073083.1073153 Thumbs up or thumbs down? semantic orientation applied to unsupervised classification of reviews . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 417--424, Philadelph...

  42. [50]

    Mohammad

    Krishnapriya Vishnubhotla and Saif M. Mohammad. 2022. https://aclanthology.org/2022.lrec-1.442/ Tweet Emotion Dynamics : Emotion word usage in tweets from US and C anada . In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 4162--4176, Marseill...

  43. [51]

    Melissa LH V \ o , Markus Conrad, Lars Kuchinke, Karolina Urton, Markus J Hofmann, and Arthur M Jacobs. 2009. The berlin affective word list reloaded (bawl-r). Behavior research methods, 41(2):534--538

  44. [52]

    Mohammad

    Jan Philip Wahle, Krishnapriya Vishnubhotla, Bela Gipp, and Saif M. Mohammad. 2025. Affect, body, cognition, demographics, and emotion: The abcde of text features for computational affective science. arXiv

  45. [53]

    Amy Beth Warriner, Victor Kuperman, and Marc Brysbaert. 2013. Norms of valence, arousal, and dominance for 13,915 E nglish lemmas. Behavior Research Methods, 45(4):1191--1207

  46. [54]

    Joseph P Weir. 2005. Quantifying test-retest reliability using the intraclass correlation coefficient and the sem. The Journal of Strength & Conditioning Research, 19(1):231--240

  47. [55]

    Bogdan Wojciszke, Andrea E Abele, and Wies aw Baryla. 2009. Two dimensions of interpersonal attitudes: Liking depends on communion, respect depends on agency. European Journal of Social Psychology, 39(6):973--990

  48. [56]

    Mi Zhou, Vibhanshu Abhishek, Timothy Derdenger, Jaymo Kim, and Kannan Srinivasan. 2024. Bias in G enerative AI . arXiv preprint arXiv:2403.02726

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.