Pith. sign in

REVIEW 3 major objections 4 minor 121 references

STEREODISCO: Discovering Stereotypicality in LLMs

T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read The paper argues that stereotypes in large language models are encoded as geometric axes in the models' internal activation space, and that two different models agree on those axes more than either agrees with humans.

desk verdict A genuinely useful probing framework whose headline stereotype-discovery results are undermined by an unmatched reference set: C′ is inanimate nouns, so the KS test may be measuring animacy rather than stereotypicality. read the letter →

arxiv 2607.27824 v1 pith:FFJKY3PP submitted 2026-07-30 cs.AI cs.LG

classification cs.AIcs.LG
keywords stereotypediscoverysemanticdifferentialprobinggeometricaxesLLMinternalrepresentationsactivationspaceKolmogorov-Smirnovtestsocialgroupstereotypes
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's central claim is that a language model's stereotype content is not scattered across its parameters but is laid out as directions in its internal activation space: each pair of opposites, like cowardly versus brave, becomes a line, and a social group like 'CEOs' is located by projecting its neural representation onto that line. To test this, the authors build a discovery pipeline that starts from about 2,000 antonym pairs, recovers each pair as a geometric axis by probing attention-head activations, and then flags axes as stereotypical when social-group projections spread out significantly more than projections of random noun phrases. Applied to two 7–8 billion parameter instruction-tuned models, the pipeline finds that the models rate social groups more like each other (73–75% position agreement) than like human survey ratings (55–63%), and it surfaces axes such as cowardly/brave and narrow-minded/broad-minded that prior stereotype dictionaries omit. If the claim holds, audits of LLM bias can look for stereotypical axes beyond the few studied in social psychology, and the recovered directions can be used for targeted steering without extra prompting or fine-tuning.

What carries the argument

The load-bearing identity is the geometric axis: the direction θ = μ+ − μ− formed by subtracting the mean activation of one pole's sentences from the mean activation of the other pole's sentences at a given attention head. Because swapping poles only negates the direction, the axis itself is label-invariant. The projection score π(c) = x(c)ᵀθ/‖θ‖² places each concept along the axis, and the stereotypicality decision is a two-sample Kolmogorov–Smirnov test on the distribution of projections for the concept set versus a random reference set. What the machinery does is convert an intangible semantic opposition into a linear direction that can be located in specific attention heads and then stee

What would settle it

Re-run STEREODISCO with a reference set matched on animacy and concreteness (e.g., other human-relevant words that are not social groups) and check whether axes such as small/large, green/ripe, and rough/smooth — which the current reference set flags as stereotypical while humans call them not applicable — still pass the p<0.05 test. If they do, the stereotypicality claim survives with a stronger reference; if they drop out, the current headline list partly depends on the reference-set artifact.

Watch

Extended reading notes

Core claim

STEREODISCO treats a semantic differential axis — a pair of opposite adjectives — as a candidate geometric axis in the model's activation space. For each of roughly 2,000 antonym pairs, it builds a probing dataset of sentences instantiating both poles, reads activations from every attention head, and takes the difference of the two pole means as the axis direction. Social-group mentions are then projected onto the top-scoring heads' axes and z-scored, and a Kolmogorov–Smirnov test compares the projections of 50 social groups with those of 50 frequency-matched random noun phrases. An axis is called stereotypical when the two distributions differ at p<0.05. The case study's empirical finding i

Load-bearing premise

The load-bearing premise is that a significant difference between how 50 social-group words and 50 random noun phrases project onto an axis measures stereotypicality; since the random nouns are not matched for human-relevance, the flagged set may partly reflect the general difference between people and objects rather than stereotypes about specific groups.

Editorial extensions

If this is right

  • Stereotype audits can move beyond the small set of warmth/competence dimensions: STEREODISCO surfaces axes like cowardly/brave that previous dictionaries miss, so bias evaluations that only use predefined axes will undercount the stereotypes an LLM actually encodes.
  • Because the stereotypical axis is a linear direction in specific middle-to-late attention heads, interventions can steer outputs along it without prompting or fine-tuning.
  • Cross-model agreement is not evidence of human alignment: the two models agree with each other more than with human ratings, so an LLM's 'consensus' stereotype content can diverge from documented human stereotypes.
  • The line between stereotypical and non-stereotypical axes is visible in the projection distributions: groups spread toward both poles on stereotypical axes and overlap with random phrases on non-stereotypical ones, giving a quantitative criterion for what counts as a stereotype axis.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the reference set were matched on human-relevance — replacing random nouns like 'capacitor' and 'ginger nut' with phrases that can describe people — several well-powered axes flagged as stereotypical (small/large, green/ripe, rough/smooth) would likely stop being flagged, because the distribution shift on those axes may just be the difference between social groups and objects.
  • The paper's H3 test treats any distribution shift as stereotypical, which means the method measures a relative property (axes along which social groups differ from random nouns), not an absolute property of the model. The same method applied with a different reference set (e.g., occupations vs. social groups) would discover different axes.
  • If the geometric-axis claim scales, the discovery pipeline could be applied to food, brands, or individuals, producing stereotype maps for any concept family that has a meaningful opposite set — but the paper only demonstrates social groups, so this is an extension, not a result.
  • A testable extension: use the discovered axes as steering directions and measure whether downstream generation shifts accordingly; that would connect the representational discovery to behavior.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces STEREODISCO, a framework for discovering stereotypical semantic axes in the internal representations of LLMs. It constructs ~2,000 candidate antonym axes from WordNet, recovers each as a geometric direction in attention-head activation space via mass-mean probing, projects concept mentions onto those directions, and uses a two-sample Kolmogorov–Smirnov test to identify axes along which projections of social groups differ from those of random nouns. In a case study with Llama-3-8B-Instruct and Mistral-7B-Instruct, the authors find that the two LLMs agree more with each other on social-group ratings (0.73–0.75 position accuracy) than either agrees with human survey ratings (0.55–0.63), and they report discovery of stereotype axes such as humble/proud, narrow-minded/broad-minded, and cowardly/brave, which are confirmed by human annotators.

Significance. If the method is validated, the contribution is significant: it moves stereotype analysis from predefined, theory-driven axes to systematic discovery in LLM internal states, localizes the axes to specific attention heads, and is generalizable to other concept families. The paper also raises an important empirical observation about LLM–human stereotype divergence. Strengths include a well-structured framework with explicit design choices, a full ablation study, power analysis, human annotation with quality controls, and clear reporting of data and compute resources. The central concern is whether the statistical test actually measures stereotypicality rather than merely animacy/applicability differences between social groups and inanimate reference nouns.

major comments (3)
  1. [§4.4 and §5.2] The H3 stereotypicality test compares projections of 50 social-group mentions (C, Table 8) against 50 frequency-matched random nouns/noun phrases (C′, Table 9) that are predominantly inanimate objects (e.g., country house, capacitor, atomic bomb). The two-sample KS test therefore rejects H0 whenever social-group projections differ from object projections, which can be driven by animacy, concreteness, or applicability rather than stereotype content. This is confirmed in Tables 10–11: both models flag small/large (Llama D=0.44, Mistral D=0.36), green/ripe (D=0.50, 0.46), and rough/smooth (D=0.40, 0.36) as well-powered stereotypical axes, while human annotators mark them 'not applicable.' The test's construct validity as a measure of stereotypicality is compromised, and the counts of discovered stereotypical axes are inflated. Please re-run with a reference set matched on human-applicabilit
  2. [Table 1 and §6] The headline discovery results (humble/proud, narrow-minded/broad-minded, cowardly/brave) are all axes that apply to humans but not to objects. Because C is human and C′ is not, these axes may pass the KS test for exactly the same reason as small/large and green/ripe, rather than because they encode stereotype content. The human confirmation for the specific axes in Table 1 is encouraging, but it does not establish that the framework's statistical test itself selects stereotypical axes; it only shows that some outputs happen to be confirmed. The paper should either (a) reframe the discovery claim as 'axes whose projections separate social groups from objects,' or (b) validate the test on an applicability-matched reference set and show that the Table 1 axes are still flagged.
  3. [Appendix C, Table 6] The classification 'genuine nulls' is assigned to not-flagged axes with power <0.30 and justified by the statement that 'the test would reliably detect even moderate effects but finds none.' This is incorrect: power <0.30 means the test is unlikely to detect a moderate effect, so non-significance is inconclusive, not evidence of a null. This category should be labeled 'inconclusive low power' or the power threshold/definition should be revised. The interpretation of not-flagged axes in Tables 10–11 is therefore affected.
minor comments (4)
  1. [§4.1] Line 1: 'Kondovs.messy' should read 'Kondo vs. messy'.
  2. [Figure 2] Caption contains missing spaces: 'STEREODISCOinstantiated' and 'STEREODISCOprediction'.
  3. [§5.2] Clarify how axes marked 'not applicable' by human annotators are treated: they appear in Tables 10–11 with Hum. '–' but are excluded from the 76 axes in Fig. 2. Please report how many of the 100 axes were excluded and whether any model-flagged axes fall in this group.
  4. [Tables 10–11] Add a column indicating whether each axis is included in the 76 applicable axes, to make the exclusion transparent and to help readers separate the discovery test from the human-comparison analysis.

Circularity Check

1 steps flagged · score 2.0 of 10

No core circularity: axes and KS test are self-contained; only the hyperparameter selection for the human-agreement evaluation is circularity-adjacent.

  1. fitted input called prediction [Appendix A (Selected configuration) and §5.1 (Position Prediction, Fig. 1)]
    "Selected configuration. The Listing template with n=30 templates, mean over all tokens, and k∈[128,512] is the best configuration in all four (model, dimension) cells. We therefore use this configuration in the main paper, fixing k=128 as a single operating point that balances Warmth (which peaks at k=128) and Competence (which keeps improving up to k=512)."

    The agreement rates reported as evidence for H4 (§5.1, Fig. 1) are computed on the same human Warmth/Competence labels (Fraser et al. 2021 compilation) that Appendix A used to select the method's configuration (template type, token position, ensemble size k). The configuration was chosen as 'the best configuration in all four (model, dimension) cells' on those exact labels, and the main paper then reports agreement on the same labels. The reported human-agreement numbers are therefore in-sample, partly determined by the selection step, rather than independent predictions. The geometric axes (Eqs. 1-3) and the KS stereotypicality test (§4.4) are computed from separate pole-sentence activations and concept/reference projections, so the core discovery pipeline is not circular; this is a local

full rationale

The derivation chain is largely self-contained. Semantic axes are recovered with mass-mean probing from WordNet antonym-pole sentences (Eqs. 1-3); concept projections come from separate prompts (Eq. 6); the KS test compares the 50 social-group projections against 50 random-noun projections (§4.4). No axis or projection is defined in terms of the stereotypicality labels, and the novel-axis claims (Table 1) are empirical outputs, not inputs. There are no load-bearing self-citations: the references to Marks & Tegmark, Li et al., Nicolas et al., and Fraser et al. are external prior work, not the authors' own chain. The H3 reference-set confound — C is human social groups while C′ is inanimate nouns/phrases (Tables 8-9), so axes like small/large and green/ripe are flagged while humans mark them Not applicable (Tables 10-11) — is a construct-validity threat to the stereotypicality operationalization, but it is not circularity: the KS result is not built into the input, and the paper's Limitations and Appendix A explicitly disclaim ground truth for LLM stereotypes. The only circularity-adjacent issue is the Appendix A selection of templates/token/k on the same human SCM labels later used for the §5.1 agreement headline; that is a selection-bias concern, not a reduction of the framework's core result. Hence score 2.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The framework rests on standard ML probing assumptions and on the semantic-differential mapping from WordNet opposites to geometric axes. No new physical or conceptual entities are postulated. The most consequential free parameter is the ensemble size k, selected on the human-agreement metric that later becomes the main evaluation. The reference-set construction (random noun phrases) is the key domain assumption that can confound the stereotypicality test.

free parameters (4)
  • ensemble size k = 128
    Head-aggregation size in Eq. 7; selected in Appendix A as the best operating point on the human-agreement evaluation, which is also the main metric in §5.1.
  • template count n = 30
    Number of sentences per pole in probing dataset; ablation shows negligible difference vs. n=15, so a minor free choice.
  • Zipf frequency threshold = 3.0
    Filtering criterion for eligible antonym pairs in stratified sampling (§B.2.1); affects which axes enter the candidate pool.
  • significance level α = 0.05
    Standard threshold for the KS stereotypicality test; not data-fitted.
assumptions (5)
  • domain assumption WordNet antonym synset pairs operationalize meaningful semantic axes.
    The entire candidate pool comes from WordNet antonymy at sense level (§4.1).
  • domain assumption Attention-head outputs of decoder-only LLMs linearly encode semantic axes, recoverable by mass-mean probing.
    Assumed in §4.2 as H1, supported by citations to prior probing work, but not independently verified in this paper.
  • domain assumption The prompt 'Provide the best description of {c} using adjectives' elicits stereotype-relevant activations equally for social groups and random phrases.
    §4.3 concept activation; no validation that this prompt is equally apt for object phrases like 'country house' and group mentions like 'CEOs'.
  • domain assumption A two-sample KS test on projections of 50 vs. 50 items is an appropriate and sufficiently powered stereotypicality test.
    §4.4; the paper's own power analysis (Appendix C, Tables 10–11) shows many non-significant axes are underpowered, so this assumption holds only partially.
  • domain assumption Majority-vote aggregation of 5 annotators (Fleiss κ=0.375) yields a reliable ground truth for whether an axis is stereotypical.
    Appendix B.2.2; inter-annotator agreement is only 'fair', and the pool is unpaid volunteers with diverse backgrounds.

how reviews work

0 comments
Cite this review

Pith. "Pith review of STEREODISCO: Discovering Stereotypicality in LLMs." pith.science (2026). https://pith.science/paper/FFJKY3PP

@misc{pith2026260727824,
  author       = {Pith},
  title        = {Pith review of: STEREODISCO: Discovering Stereotypicality in LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FFJKY3PP}},
  note         = {Machine review of arXiv:2607.27824}
}
read the original abstract

LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psychology, and operates on word embeddings produced by language models, leaving open which other semantic axes carry stereotypical associations in LLMs and how LLMs internally represent such axes. We introduce STEREODISCO, a framework that adapts the semantic differential method (Osgood et al., 1957) to the systematic study of stereotypes in LLM internal representations. STEREODISCO constructs approx. 2,000 candidate semantic axes from WordNet antonym synsets, recovers each as a geometric axis in the LLM's activation space via probing, and identifies stereotypical axes via a statistical test over concept projections. As a case study, we apply STEREODISCO to social group stereotypes with LLAMA-3-8B-INSTRUCT and MISTRAL-7B-INSTRUCT. We find that the two LLMs agree with each other on social group ratings more than with humans, suggesting that LLM-encoded stereotype content diverges from that documented in social psychology. We also discover stereotypical axes not investigated in prior work -- including humble vs. proud, narrow-minded vs. broad-minded, and cowardly vs. brave, which human annotators independently confirm.

Figures

Figures reproduced from arXiv: 2607.27824 by the authors.

Figure 1
Figure 1. Pairwise agreement (position prediction accu [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. STEREODISCO instantiated with LLAMA-3- 8B-INSTRUCT compared against human annotations on a stratified sample of semantic axes, judging whether each axis is stereotypical. Each point is one of the 76 semantic axes, positioned by human judgment (x-axis) and STEREODISCO prediction (y-axis); diagonal points indicate agreement (43 axes: 31 stereotypical, 12 non￾stereotypical). Off-diagonal points show disagreements (19 h… view at source ↗
Figure 3
Figure 3. Projection distributions for STEREODISCO with LLAMA-3-8B-INSTRUCT. Each point is a social group mention (red) or random phrase (blue) projected onto the geometric axis (Eq. 7); vertical lines mark group medians. Semantic axis (Neg./Pos.) Llama Mistral unreasonable.a.01/reasonable.a.01 0.62 0.44 cowardly.a.01/brave.a.01 0.56 0.32 narrow-minded.a.02/broad￾minded.a.02 0.34 0.60 retarded.a.01/precocious.a.01 0.32 0.40 h… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Variance-ratio heatmaps (Eq. 4) on LLAMA-3-8B-INSTRUCT (left two columns) and MISTRAL-7B￾INSTRUCT (right two columns) for representative semantic axes. For each LLM, the left panel uses the Last token and the right panel uses the Mean over all tokens. Last-token heatma…
Figure 5
Figure 5. Figure 5: Screenshot of our annotation interface. C Power Analysis We estimate the reliability of the Kolmogorov– Smirnov test conducted in §4.4 by Monte Carlo simulation (Mooney, 1997). We draw 1,000 boot￾strap resamples of size nC = nC ′ = 50 from the ob￾served projections and…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

121 extracted references · 20 canonical work pages

  1. [1]

    , title =

    Stephan, Walter G. , title =. Stereotyping and Prejudice , series =. 1989 , publisher =

  2. [2]

    , title =

    Stangor, Charles and Lange, James E. , title =. Advances in Experimental Social Psychology , volume =. 1994 , publisher =

  3. [3]

    Journal of Personality and Social Psychology , volume =

    Stereotypes and prejudice: Their automatic and controlled components , author =. Journal of Personality and Social Psychology , volume =. 1989 , doi =

  4. [4]

    and Abele, Andrea E

    Koch, Alex and Smith, Austin and Fiske, Susan T. and Abele, Andrea E. and Ellemers, Naomi and Yzerbyt, Vincent , title =. Behavior Research Methods , year =. doi:10.3758/s13428-024-02489-y , url =

  5. [5]

    Jenkins and Pierre Karashchuk and Lusha Zhu and Ming Hsu , title =

    Adrianna C. Jenkins and Pierre Karashchuk and Lusha Zhu and Ming Hsu , title =. Proceedings of the National Academy of Sciences , volume =. 2018 , doi =

  6. [6]

    Fiske and Gandalf Nicolas , title =

    Vincent Yzerbyt and Alex Koch and Marco Brambilla and Naomi Ellemers and Susan T. Fiske and Gandalf Nicolas , title =. Current Directions in Psychological Science , volume =. 2025 , doi =

  7. [7]

    White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLM s

    Wan, Yixin and Chang, Kai-Wei. White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLM s. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. doi:10.18653/v1/2025.acl-long.445

  8. [8]

    Proceedings of the AAAI Conference on Artificial Intelligence , author=

    SCoUT: A Framework for Structured Stereotype Analysis in Language Models , volume=. Proceedings of the AAAI Conference on Artificial Intelligence , author=. 2026 , month=. doi:10.1609/aaai.v40i32.39905 , number=

Show all 121 references
  1. [9]

    LL a MA s Have Feelings Too: Unveiling Sentiment and Emotion Representations in LL a MA Models Through Probing

    Di Palma, Dario and De Bellis, Alessandro and Servedio, Giovanni and Anelli, Vito Walter and Narducci, Fedelucio and Di Noia, Tommaso. LL a MA s Have Feelings Too: Unveiling Sentiment and Emotion Representations in LL a MA Models Through Probing. Proceedings of the 63rd Annual...

  2. [10]

    ``Kelly is a Warm Person, Joseph is a Role Model'': Gender Biases in LLM -Generated Reference Letters

    Wan, Yixin and Pu, George and Sun, Jiao and Garimella, Aparna and Chang, Kai-Wei and Peng, Nanyun. ``Kelly is a Warm Person, Joseph is a Role Model'': Gender Biases in LLM -Generated Reference Letters. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023...

  3. [11]

    Fiske and Chris Malone , keywords =

    Nicolas Kervyn and Susan T. Fiske and Chris Malone , keywords =. Brands as intentional agents framework: How perceived intentions and ability can map brand perception , journal =. 2012 , issn =. doi:https://doi.org/10.1016/j.jcps.2011.09.006 , url =

  4. [12]

    1988 , publisher =

    Cohen, Jacob , title =. 1988 , publisher =. doi:10.4324/9780203771587 , url =

  5. [13]

    , title =

    Mooney, Christopher Z. , title =. 1997 , series =

  6. [14]

    Journal of Statistical and Econometric Methods , volume=

    Comparison of the powers of the Kolmogorov-Smirnov Two-Sample Test and the Mann-Whitney Test for different kurtosis and Skewness coefficients using the Monte Carlo simulation method , author=. Journal of Statistical and Econometric Methods , volume=. 2013 , publisher=

  7. [15]

    2022 , publisher =

    Robyn Speer , title =. 2022 , publisher =. doi:10.5281/zenodo.7199437 , url =

  8. [16]

    American Sociological Review , volume =

    András Tilcsik , title =. American Sociological Review , volume =. 2021 , doi =

  9. [17]

    and Hauke, Nicole and Peters, Kim and Louvet, Eva and Szymkow, Aleksandra and Duan, Yanping , journal =

    Abele, Andrea E. and Hauke, Nicole and Peters, Kim and Louvet, Eva and Szymkow, Aleksandra and Duan, Yanping , journal =. Facets of the fundamental content dimensions:. 2016 , publisher =

  10. [18]

    , journal =

    Nicolas, Gandalf and Bai, Xuechunzi and Fiske, Susan T. , journal =. A spontaneous stereotype content model:. 2022 , month = dec, publisher =. doi:10.1037/pspa0000312 , pmid =

  11. [19]

    , title =

    Nicolas, Gandalf and Bai, Xuechunzi and Fiske, Susan T. , title =. European Journal of Social Psychology , volume =. doi:https://doi.org/10.1002/ejsp.2724 , url =. https://onlinelibrary.wiley.com/doi/pdf/10.1002/ejsp.2724 , year =

  12. [20]

    and Rossi, Ryan A

    Gallegos, Isabel O. and Rossi, Ryan A. and Barrow, Joe and Tanjim, Md Mehrab and Kim, Sungchul and Dernoncourt, Franck and Yu, Tong and Zhang, Ruiyi and Ahmed, Nesreen K. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics. 2024. doi:10.1162/coli_a_00524

  13. [21]

    Layers at Similar Depths Generate Similar Activations Across

    Christopher Wolfram and Aaron Schein , booktitle=. Layers at Similar Depths Generate Similar Activations Across. 2025 , url=

  14. [22]

    2026 , eprint=

    Testing the Limits of Truth Directions in LLMs , author=. 2026 , eprint=

  15. [23]

    Detecting (Un)answerability in Large Language Models with Linear Directions

    Lavi, Maor Juliet and Milo, Tova and Geva, Mor. Detecting (Un)answerability in Large Language Models with Linear Directions. Proceedings of the 19th Conference of the E uropean Chapter of the A ssociation for C omputational L inguistics (Volume 1: Long Papers). 2026. doi:10.18...

  16. [24]

    Nature , volume =

    Hofmann, Valentin and Kalluri, Pratyusha Ria and Jurafsky, Dan and King, Sharese , title =. Nature , volume =. 2024 , doi =

  17. [25]

    2023 , eprint=

    LLaMA: Open and Efficient Foundation Language Models , author=. 2023 , eprint=

  18. [26]

    arXiv preprint arXiv:2307.09288 , year=

    Llama 2: Open Foundation and Fine-Tuned Chat Models , author=. arXiv preprint arXiv:2307.09288 , year=

  19. [27]

    G lo V e: Global Vectors for Word Representation

    Pennington, Jeffrey and Socher, Richard and Manning, Christopher. G lo V e: Global Vectors for Word Representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing ( EMNLP ). 2014. doi:10.3115/v1/D14-1162

  20. [28]

    Advances in Pre-Training Distributed Word Representations

    Mikolov, Tomas and Grave, Edouard and Bojanowski, Piotr and Puhrsch, Christian and Joulin, Armand. Advances in Pre-Training Distributed Word Representations. Proceedings of the Eleventh International Conference on Language Resources and Evaluation ( LREC 2018). 2018

  21. [29]

    Proceedings of the International Conference on Learning Representations (ICLR) , year =

    Tomas Mikolov and Kai Chen and Greg Corrado and Jeffrey Dean , title =. Proceedings of the International Conference on Learning Representations (ICLR) , year =

  22. [30]

    S ense POLAR : Word sense aware interpretability for pre-trained contextual word embeddings

    Engler, Jan and Sikdar, Sandipan and Lutz, Marlene and Strohmaier, Markus. S ense POLAR : Word sense aware interpretability for pre-trained contextual word embeddings. Findings of the Association for Computational Linguistics: EMNLP 2022. 2022. doi:10.18653/v1/2022.findings-emnlp.338

  23. [31]

    Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words

    Zhou, Kaitlyn and Ethayarajh, Kawin and Card, Dallas and Jurafsky, Dan. Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2022. d...

  24. [32]

    Understanding Undesirable Word Embedding Associations

    Ethayarajh, Kawin and Duvenaud, David and Hirst, Graeme. Understanding Undesirable Word Embedding Associations. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. doi:10.18653/v1/P19-1166

  25. [33]

    Social-Group-Agnostic Bias Mitigation via the Stereotype Content Model

    Omrani, Ali and Salkhordeh Ziabari, Alireza and Yu, Charles and Golazizian, Preni and Kennedy, Brendan and Atari, Mohammad and Ji, Heng and Dehghani, Morteza. Social-Group-Agnostic Bias Mitigation via the Stereotype Content Model. Proceedings of the 61st Annual Meeting of the ...

  26. [34]

    and Rudinger, Rachel

    May, Chandler and Wang, Alex and Bordia, Shikha and Bowman, Samuel R. and Rudinger, Rachel. On Measuring Social Biases in Sentence Encoders. Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Techn...

  27. [35]

    2021 , isbn =

    Guo, Wei and Caliskan, Aylin , title =. 2021 , isbn =. doi:10.1145/3461702.3462536 , booktitle =

  28. [36]

    Science , volume=

    Semantics derived automatically from language corpora contain human-like biases , author=. Science , volume=. 2017 , url=

  29. [37]

    S tereo S et: Measuring stereotypical bias in pretrained language models

    Nadeem, Moin and Bethke, Anna and Reddy, Siva. S tereo S et: Measuring stereotypical bias in pretrained language models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Proc...

  30. [38]

    Frontiers in Artificial Intelligence , volume=

    Computational Modeling of Stereotype Content in Text , author=. Frontiers in Artificial Intelligence , volume=. 2022 , publisher=

  31. [39]

    A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI Evaluations

    Mostafazadeh Davani, Aida and Dev, Sunipa and P \'e rez-Urbina, H \'e ctor and Prabhakaran, Vinodkumar. A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI Evaluations. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Pr...

  32. [40]

    RoBERTa: A Robustly Optimized

    Liu, Yinhan and Ott, Myle and Goyal, Naman and Du, Jingfei and Joshi, Mandar and Chen, Danqi and Levy, Omer and Lewis, Mike and Zettlemoyer, Luke and Stoyanov, Veselin , journal =. RoBERTa: A Robustly Optimized

  33. [41]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding , author =. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , year =

  34. [42]

    Stereotype Activation and Inhibition , isbn =

    Bodenhausen, Galen and Macrae, C , year =. Stereotype Activation and Inhibition , isbn =

  35. [43]

    2024 , eprint=

    The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets , author=. 2024 , eprint=

  36. [44]

    Truth is Universal: Robust Detection of Lies in

    Lennart B. Truth is Universal: Robust Detection of Lies in. The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

  37. [45]

    Attention is All you Need , url =

    Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser, ukasz and Polosukhin, Illia , booktitle =. Attention is All you Need , url =

  38. [46]

    Transformer Circuits Thread , year =

    A Mathematical Framework for Transformer Circuits , author =. Transformer Circuits Thread , year =

  39. [47]

    Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume =

    Nicolas, Gandalf and Caliskan, Aylin , title =. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume =. 2025 , doi =

  40. [48]

    How Does Stereotype Content Differ across Data Sources?

    Fraser, Kathleen and Kiritchenko, Svetlana and Nejadgholi, Isar. How Does Stereotype Content Differ across Data Sources?. Proceedings of the 13th Joint Conference on Lexical and Computational Semantics (*SEM 2024). 2024

  41. [49]

    The Lancet Digital Health , year =

    Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study , author =. The Lancet Digital Health , year =

  42. [50]

    OpenAI Blog , volume=

    Language Models are Unsupervised Multitask Learners , author=. OpenAI Blog , volume=

  43. [51]

    The Impact of Name Age Perception on Job Recommendations in LLM s

    Kamruzzaman, Mahammed and Kim, Gene Louis. The Impact of Name Age Perception on Job Recommendations in LLM s. Findings of the Association for Computational Linguistics: ACL 2025. 2025. doi:10.18653/v1/2025.findings-acl.778

  44. [52]

    Tessa E. S. Charlesworth and Aylin Caliskan and Mahzarin R. Banaji , title =. Proceedings of the National Academy of Sciences , volume =

  45. [53]

    What ' s in the Box? An Analysis of Undesirable Content in the C ommon C rawl Corpus

    Luccioni, Alexandra and Viviano, Joseph. What ' s in the Box? An Analysis of Undesirable Content in the C ommon C rawl Corpus. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Languag...

  46. [54]

    Kozlowski and Matt Taddy and James A

    Austin C. Kozlowski and Matt Taddy and James A. Evans , title =. American Sociological Review , volume =. 2019 , doi =

  47. [55]

    2025 , eprint=

    Semantic Structure in Large Language Model Embeddings , author=. 2025 , eprint=

  48. [56]

    B abel D omains: Large-Scale Domain Labeling of Lexical Resources

    Camacho-Collados, Jose and Navigli, Roberto. B abel D omains: Large-Scale Domain Labeling of Lexical Resources. Proceedings of the 15th Conference of the E uropean Chapter of the Association for Computational Linguistics: Volume 2, Short Papers. 2017

  49. [57]

    B abel N et: Building a Very Large Multilingual Semantic Network

    Navigli, Roberto and Ponzetto, Simone Paolo. B abel N et: Building a Very Large Multilingual Semantic Network. Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics. 2010

  50. [58]

    , title =

    Phipson, Belinda and Smyth, Gordon K. , title =. Statistical Applications in Genetics and Molecular Biology , year =

  51. [59]

    Ford and George R

    Thomas E. Ford and George R. Tonander , journal =. The Role of Differentiation Between Groups and Social Identity in Stereotype Formation , urldate =

  52. [60]

    Charlesworth, Tessa E. S. and Sanjeev, Nikhil and Hatzenbuehler, Mark L. and Banaji, Mahzarin R. , title =. Journal of Personality and Social Psychology , year =. doi:10.1037/pspa0000354 , pmid =

  53. [61]

    2020 , issn =

    Groups' warmth is a personal matter: Understanding consensus on stereotype dimensions reconciles adversarial models of social evaluation , journal =. 2020 , issn =. doi:https://doi.org/10.1016/j.jesp.2020.103995 , author =

  54. [62]

    Thirty-seventh Conference on Neural Information Processing Systems , year=

    Inference-Time Intervention: Eliciting Truthful Answers from a Language Model , author=. Thirty-seventh Conference on Neural Information Processing Systems , year=

  55. [63]

    , title =

    Hebb, Donald O. , title =. 1949 , publisher =

  56. [64]

    Trends in Cognitive Sciences , volume =

    A Network Neuroscience of Human Learning: Potential to Inform Quantitative Theories of Brain and Behavior , author =. Trends in Cognitive Sciences , volume =. 2017 , month =. doi:10.1016/j.tics.2017.01.010 , pmid =

  57. [65]

    ``You Gotta be a Doctor, Lin'' : An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations

    Nghiem, Huy and Prindle, John and Zhao, Jieyu and Daum \'e Iii, Hal. ``You Gotta be a Doctor, Lin'' : An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Process...

  58. [66]

    Employed persons by detailed occupation, sex, race, and Hispanic or Latino ethnicity , year =

  59. [67]

    Language (Technology) is Power: A Critical Survey of ``Bias'' in NLP

    Blodgett, Su Lin and Barocas, Solon and Daum \'e III, Hal and Wallach, Hanna. Language (Technology) is Power: A Critical Survey of ``Bias'' in NLP. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.485

  60. [68]

    Journal of Personality and Social Psychology , volume =

    Koch, Alex and Imhoff, Roland and Dotsch, Ron and Unkelbach, Christian and Alves, Hans , title =. Journal of Personality and Social Psychology , volume =. 2016 , month = may, publisher =. doi:10.1037/pspa0000046 , issn =

  61. [69]

    Theory-Grounded Measurement of U

    Cao, Yang Trista and Sotnikova, Anna and Daum \'e III, Hal and Rudinger, Rachel and Zou, Linda. Theory-Grounded Measurement of U . S . Social Stereotypes in E nglish Language Models. Proceedings of the 2022 Conference of the North American Chapter of the Association for Comput...

  62. [70]

    2006 , edition =

    Evans, Vyvyan and Green, Melanie , title =. 2006 , edition =

  63. [71]

    Sex Roles , year =

    Eckes, Thomas , title =. Sex Roles , year =. doi:10.1023/A:1021020920715 , issn =

  64. [72]

    S tereo D etect: Detecting Stereotypes and Anti-stereotypes the Correct Way Using Social Psychological Underpinnings

    Shejole, Kaustubh Shivshankar and Bhattacharyya, Pushpak. S tereo D etect: Detecting Stereotypes and Anti-stereotypes the Correct Way Using Social Psychological Underpinnings. Findings of the Association for Computational Linguistics: EMNLP 2025

  65. [73]

    Proceedings of the AAAI Conference on Artificial Intelligence , author=

    ConceptNet 5.5: An Open Multilingual Graph of General Knowledge , volume=. Proceedings of the AAAI Conference on Artificial Intelligence , author=. 2017 , month=. doi:10.1609/aaai.v31i1.11164 , number=

  66. [74]

    Charlesworth, Tessa E. S. and Ghate, Kshitish and Caliskan, Aylin and Banaji, Mahzarin R. , title =. PNAS Nexus , year =

  67. [75]

    Social Psychology , volume =

    The effects of status on perceived warmth and competence: Malleability of the relationship between status and stereotype content , author =. Social Psychology , volume =. 2010 , doi =

  68. [76]

    2006 , note =

    Not an outgroup, not yet an ingroup: Immigrants in the Stereotype Content Model , journal =. 2006 , note =. doi:https://doi.org/10.1016/j.ijintrel.2006.06.005 , url =

  69. [77]

    2007 , issn =

    Universal dimensions of social cognition: warmth and competence , journal =. 2007 , issn =. doi:https://doi.org/10.1016/j.tics.2006.11.005 , url =

  70. [78]

    2011 , issn =

    The dynamics of warmth and competence judgments, and their outcomes in organizations , journal =. 2011 , issn =. doi:https://doi.org/10.1016/j.riob.2011.10.004 , url =

  71. [79]

    2024 , doi =

    Dubey, Abhimanyu and Jauhri, Abhinav and Pandey, Abhinav and Kadian, Abhishek and Al-Dahle, Ahmad and Leszczynski, Aiesha and Asaadi, Akula and others , journal =. 2024 , doi =

  72. [80]

    Jiang, Albert Q. and Sablayrolles, Alexandre and Mensch, Arthur and Bamford, Chris and Chaplot, Devendra Singh and de las Casas, Diego and Bressand, Florian and Lengyel, Gianna and Lample, Guillaume and Saulnier, Lucile and others , journal =. 2023 , doi =

  73. [81]

    Journal of Personality and Social Psychology , volume =

    Bad but Bold: Ambivalent Attitudes Toward Men Predict Gender Inequality in 16 Nations , author =. Journal of Personality and Social Psychology , volume =. 2004 , doi =

  74. [82]

    2008 , issn =

    Warmth and Competence as Universal Dimensions of Social Perception: The Stereotype Content Model and the BIAS Map , series =. 2008 , issn =. doi:https://doi.org/10.1016/S0065-2601(07)00002-0 , url =

  75. [83]

    Losh and Ryan Wilke and Margareta Pop , title =

    Susan C. Losh and Ryan Wilke and Margareta Pop , title =. International Journal of Science Education , volume =. 2008 , publisher =. doi:10.1080/09500690701250452 , URL =

  76. [84]

    , journal =

    Fiske, Susan T. , journal =. Venus and. 2010 , doi =

  77. [85]

    and Nejadgholi, Isar and Kiritchenko, Svetlana

    Fraser, Kathleen C. and Nejadgholi, Isar and Kiritchenko, Svetlana. Understanding and Countering Stereotypes: A Computational Approach to the Stereotype Content Model. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internat...

  78. [86]

    Discovering Differences in the Representation of People using Contextualized Semantic Axes

    Lucy, Li and Tadimeti, Divya and Bamman, David. Discovering Differences in the Representation of People using Contextualized Semantic Axes. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. doi:10.18653/v1/2022.emnlp-main.228

  79. [87]

    Probing Classifiers: Promises, Shortcomings, and Advances

    Belinkov, Yonatan. Probing Classifiers: Promises, Shortcomings, and Advances. Computational Linguistics. 2022. doi:10.1162/coli_a_00422

  80. [88]

    Daniel Jurafsky and James H. Martin. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, with Language Models. 2026

  81. [89]

    Current Directions in Psychological Science , volume =

    Stereotype Content: Warmth and Competence Endure , author =. Current Directions in Psychological Science , volume =. 2018 , doi =

  82. [90]

    The Journal of Abnormal and Social Psychology , volume =

    Racial Prejudice and Racial Stereotypes , author =. The Journal of Abnormal and Social Psychology , volume =. 1935 , doi =

  83. [91]

    Journal of Abnormal and Social Psychology , volume =

    Stereotype Persistence and Change Among College Students , author =. Journal of Abnormal and Social Psychology , volume =

  84. [92]

    Journal of Personality and Social Psychology , volume =

    On the Fading of Social Stereotypes: Studies in Three Generations of College Students , author =. Journal of Personality and Social Psychology , volume =. 1969 , doi =

  85. [93]

    and Bhiwandiwalla, Anahita and Kiritchenko, Svetlana

    Howard, Phillip and Fraser, Kathleen C. and Bhiwandiwalla, Anahita and Kiritchenko, Svetlana. Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computa...

  86. [94]

    Journal of Personality and Social Psychology , volume=

    A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition , author=. Journal of Personality and Social Psychology , volume=. 2002 , doi=

  87. [95]

    W ord N et: A Lexical Database for E nglish

    Miller, George A. W ord N et: A Lexical Database for E nglish. Speech and Natural Language: Proceedings of a Workshop Held at Harriman, New York, F ebruary 23-26, 1992. 1992

  88. [96]

    S em A xis: A Lightweight Framework to Characterize Domain-Specific Word Semantics Beyond Sentiment

    An, Jisun and Kwak, Haewoon and Ahn, Yong-Yeol. S em A xis: A Lightweight Framework to Characterize Domain-Specific Word Semantics Beyond Sentiment. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. doi:10.18...

  89. [97]

    Russell , title =

    Lisa Feldman Barrett and James A. Russell , title =. Current Directions in Psychological Science , volume =

  90. [98]

    Fiske and Amy J

    Susan T. Fiske and Amy J. C. Cuddy and Peter Glick and Jun Xu , title =. Social Cognition , pages =

  91. [99]

    Nature , year =

    A foundation model to predict and capture human cognition , author =. Nature , year =

  92. [100]

    F air S teer: Inference Time Debiasing for LLM s with Dynamic Activation Steering

    Li, Yichen and Fan, Zhiting and Chen, Ruizhe and Gai, Xiaotang and Gong, Luqi and Zhang, Yan and Liu, Zuozhu. F air S teer: Inference Time Debiasing for LLM s with Dynamic Activation Steering. Findings of the Association for Computational Linguistics: ACL 2025. 2025. doi:10.18...

  93. [101]

    2024 , isbn =

    Wang, Haoran and Shu, Kai , title =. 2024 , isbn =. doi:10.1145/3627673.3679821 , booktitle =

  94. [102]

    Proceedings of the First Workshop on Gender Bias in Natural Language Processing , pages=

    Measuring bias in contextualized word representations , author=. Proceedings of the First Workshop on Gender Bias in Natural Language Processing , pages=. 2019 , organization=

  95. [103]

    Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society , pages=

    Persistent anti-Muslim bias in large language models , author=. Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society , pages=. 2021 , organization=

  96. [104]

    arXiv preprint arXiv:2106.13219 , year=

    Towards understanding and mitigating social biases in language models , author=. arXiv preprint arXiv:2106.13219 , year=

  97. [105]

    The Thirteenth International Conference on Learning Representations , year=

    The Geometry of Categorical and Hierarchical Concepts in Large Language Models , author=. The Thirteenth International Conference on Learning Representations , year=

  98. [106]

    2020 , isbn =

    Mathew, Binny and Sikdar, Sandipan and Lemmerich, Florian and Strohmaier, Markus , title =. 2020 , isbn =. doi:10.1145/3366423.3380227 , booktitle =

  99. [107]

    The Thirteenth International Conference on Learning Representations , year=

    Linear Representations of Political Perspective Emerge in Large Language Models , author=. The Thirteenth International Conference on Learning Representations , year=

  100. [108]

    1957 , publisher=

    The measurement of meaning , author=. 1957 , publisher=

  101. [109]

    NAACL-HLT , year=

    Linguistic regularities in continuous space word representations , author=. NAACL-HLT , year=

  102. [110]

    NeurIPS , year=

    Man is to computer programmer as woman is to homemaker? Debiasing word embeddings , author=. NeurIPS , year=

  103. [111]

    Transformer Circuits Thread , year=

    Toy models of superposition , author=. Transformer Circuits Thread , year=

  104. [112]

    The Twelfth International Conference on Learning Representations , year=

    Linearity of Relation Decoding in Transformer Language Models , author=. The Twelfth International Conference on Learning Representations , year=

  105. [113]

    2014 , month =

    Measuring Occupational Prestige on the 2012 General Social Survey , author =. 2014 , month =

  106. [114]

    arXiv preprint arXiv:2404.02650 , year=

    Towards detecting unanticipated bias in Large Language Models , author=. arXiv preprint arXiv:2404.02650 , year=

  107. [115]

    and Roman, Maria-Alexandra and Ghatiwala, Shashwat and Groh, Georg

    Schuster, Carolin M. and Roman, Maria-Alexandra and Ghatiwala, Shashwat and Groh, Georg. Profiling Bias in LLM s: Stereotype Dimensions in Contextual Word Embeddings. Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Hum...

  108. [116]

    and Hauke, Nicole and Peters, Kim and Louvet, Eva and Szymkow, Aleksandra and Duan, Yanping , TITLE=

    Abele, Andrea E. and Hauke, Nicole and Peters, Kim and Louvet, Eva and Szymkow, Aleksandra and Duan, Yanping , TITLE=. Frontiers in Psychology , VOLUME=. 2016 , URL=. doi:10.3389/fpsyg.2016.01810 , ISSN=

  109. [117]

    Kolmogorov, A. N. , title =. Giornale dell'Istituto Italiano degli Attuari , volume =

  110. [118]

    Smirnov, N. V. , title =. Bulletin Moscow University , volume =

  111. [119]

    Proceedings of the National Academy of Sciences , volume =

    Nikhil Garg and Londa Schiebinger and Dan Jurafsky and James Zou , title =. Proceedings of the National Academy of Sciences , volume =. 2018 , doi =

  112. [120]

    Adaptive Axes: A Pipeline for In-domain Social Stereotype Analysis

    Zeng, Qingcheng and Jin, Mingyu and Voigt, Rob. Adaptive Axes: A Pipeline for In-domain Social Stereotype Analysis. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.872

  113. [121]

    Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLM s Across Logical Transformations and Question Answering Tasks

    Bao, Yuntai and Zhang, Xuhong and Du, Tianyu and Zhao, Xinkui and Feng, Zhengwen and Peng, Hao and Yin, Jianwei. Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLM s Across Logical Transformations and Question Answering Tasks. Findings of ...

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.