{"id":"7e350e35-ca11-4c4c-8298-d742a121eba9","arxiv_id":"2607.27824","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"LLMs encode stereotypes along recoverable geometric axes in attention heads, and two tested LLMs share more stereotype content with each other than with documented human stereotypes.","lead":"This paper presents STEREODISCO, a method that looks inside large language models to find the directions (axes) along which the models associate social groups with personality attributes. It finds that two different LLMs agree more with each other than with human survey responses, and unearths stereotype axes (e.g., cowardly vs. brave) that earlier research had not examined.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"H3's stereotypicality test compares social groups to inanimate random nouns, so distributional shift may reflect animacy/applicability rather than stereotype content; Tables 10–11 show physical axes flagged while humans mark them not applicable.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the reference set is not matched on animacy/concreteness/applicability, so the H3 distributional-shift test is confounded. This concern is central because H3 is the component that decides which axes are 'stereotypical'; it feeds directly into the discovery claim (Table 1), the agreement with human annotations (Fig. 2), and the paper's headline finding that the two LLMs share stereotype content divergent from human content. The paper's own Tables 10–11 provide direct evidence: physical axes such as small/large and green/ripe are flagged with high power while human annotators mark them not applicable. A matched-reference rerun would settle whether the discovered axes survive. The config-selection issue noted by the reader is secondary: it affects the magnitude of reported human-agreement numbers but does not determine the existence of discovered axes. Since the framework is novel, the internal-geometry evidence (H1/H2) is plausible, and the confound is fixable, the reader's CONDITIONAL verdict remains appropriate; there is no basis to reject outright or to accept as-is.","tokens_in":27705,"tokens_out":4678,"duration_ms":50036,"concrete_test":"Rerun §4.4 with a reference set C′ matched on animacy and human-applicability—e.g., 50 human-denoting noun phrases that are not social-group stereotype targets ('a person', 'an adult', 'someone', 'an individual', 'a pedestrian', 'a customer', 'an inhabitant'), frequency-matched to C. Recompute the stereotypicality classification for all 76 axes and the Table 1 discoveries. If most well-powered axes (especially small/large, green/ripe, rough/smooth) become non-significant and Table 1 shrinks or reverses, the original C′ is the source of the findings. As a negative control, also run the same test with C′ as the 50 social groups and C as random objects; H3 predicts the asymmetry should disappear if it measures stereotype content rather than animacy/applicability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is the construct validity of the H3 stereotypicality test (§4.4). The reference set C′ is 50 frequency-matched random nouns/noun phrases from WordNet (Table 9: country house, record, church, tank ship, atomic bomb, capacitor, etc.), while C is 50 social-group mentions (Table 8: women, CEOs, refugees, etc.). The two-sample KS test therefore rejects H0 whenever social-group projections differ from object/artifact projections. This difference can be driven by animacy, concreteness, or simple applicability of the axis to people, not by stereotypicality. Tables 10–11 expose the confound directly: both models flag small/large (Llama D=0.44, p<.001; Mistral D=0.36, p=.003), green/ripe (D=0.50 and 0.46), and rough/smooth (D=0.40 and 0.36) as well-powered positive axes, while human annotators mark them 'not applicable'. These are not stereotypes; they are artifacts of comparing animate human groups with inanimate objects. The same mechanism inflates the count of discovered axes and the off-diagonal entries in Fig. 2: any axis that applies to people but not to objects will tend to produce a distributional shift. Thus the central empirical claim—that STEREODISCO discovers stereotypical axes such as cowardly/brave and narrow-minded/broad-minded—is at risk, because these human-applicable axes may pass simply because C is human and C′ is not. The framework's internal hypotheses H1/H2 (linear geometric encoding and projection-as-rating) are not directly threatened; the problem is the operationalization of H3.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces STEREODISCO, a framework for discovering stereotypical semantic axes in the internal representations of LLMs. It constructs ~2,000 candidate antonym axes from WordNet, recovers each as a geometric direction in attention-head activation space via mass-mean probing, projects concept mentions onto those directions, and uses a two-sample Kolmogorov–Smirnov test to identify axes along which projections of social groups differ from those of random nouns. In a case study with Llama-3-8B-Instruct and Mistral-7B-Instruct, the authors find that the two LLMs agree more with each other on social-group ratings (0.73–0.75 position accuracy) than either agrees with human survey ratings (0.55–0.63), and they report discovery of stereotype axes such as humble/proud, narrow-minded/broad-minded, and cowardly/brave, which are confirmed by human annotators.","tokens_in":28034,"tokens_out":4876,"duration_ms":45286,"significance":"If the method is validated, the contribution is significant: it moves stereotype analysis from predefined, theory-driven axes to systematic discovery in LLM internal states, localizes the axes to specific attention heads, and is generalizable to other concept families. The paper also raises an important empirical observation about LLM–human stereotype divergence. Strengths include a well-structured framework with explicit design choices, a full ablation study, power analysis, human annotation with quality controls, and clear reporting of data and compute resources. The central concern is whether the statistical test actually measures stereotypicality rather than merely animacy/applicability differences between social groups and inanimate reference nouns.","major_comments":[{"comment":"The H3 stereotypicality test compares projections of 50 social-group mentions (C, Table 8) against 50 frequency-matched random nouns/noun phrases (C′, Table 9) that are predominantly inanimate objects (e.g., country house, capacitor, atomic bomb). The two-sample KS test therefore rejects H0 whenever social-group projections differ from object projections, which can be driven by animacy, concreteness, or applicability rather than stereotype content. This is confirmed in Tables 10–11: both models flag small/large (Llama D=0.44, Mistral D=0.36), green/ripe (D=0.50, 0.46), and rough/smooth (D=0.40, 0.36) as well-powered stereotypical axes, while human annotators mark them 'not applicable.' The test's construct validity as a measure of stereotypicality is compromised, and the counts of discovered stereotypical axes are inflated. Please re-run with a reference set matched on human-applicabilit","section":"§4.4 and §5.2"},{"comment":"The headline discovery results (humble/proud, narrow-minded/broad-minded, cowardly/brave) are all axes that apply to humans but not to objects. Because C is human and C′ is not, these axes may pass the KS test for exactly the same reason as small/large and green/ripe, rather than because they encode stereotype content. The human confirmation for the specific axes in Table 1 is encouraging, but it does not establish that the framework's statistical test itself selects stereotypical axes; it only shows that some outputs happen to be confirmed. The paper should either (a) reframe the discovery claim as 'axes whose projections separate social groups from objects,' or (b) validate the test on an applicability-matched reference set and show that the Table 1 axes are still flagged.","section":"Table 1 and §6"},{"comment":"The classification 'genuine nulls' is assigned to not-flagged axes with power <0.30 and justified by the statement that 'the test would reliably detect even moderate effects but finds none.' This is incorrect: power <0.30 means the test is unlikely to detect a moderate effect, so non-significance is inconclusive, not evidence of a null. This category should be labeled 'inconclusive low power' or the power threshold/definition should be revised. The interpretation of not-flagged axes in Tables 10–11 is therefore affected.","section":"Appendix C, Table 6"}],"minor_comments":[{"comment":"Line 1: 'Kondovs.messy' should read 'Kondo vs. messy'.","section":"§4.1"},{"comment":"Caption contains missing spaces: 'STEREODISCOinstantiated' and 'STEREODISCOprediction'.","section":"Figure 2"},{"comment":"Clarify how axes marked 'not applicable' by human annotators are treated: they appear in Tables 10–11 with Hum. '–' but are excluded from the 76 axes in Fig. 2. Please report how many of the 100 axes were excluded and whether any model-flagged axes fall in this group.","section":"§5.2"},{"comment":"Add a column indicating whether each axis is included in the 76 applicable axes, to make the exclusion transparent and to help readers separate the discovery test from the human-comparison analysis.","section":"Tables 10–11"}],"recommendation":"major_revision","confidential_remarks":"The core issue is the construct validity of the H3 test. The limitation section acknowledges that 'the selection of applicable semantic axes remains the responsibility of the analyst,' but the empirical counts and the discovery claim are presented as systematic results. A matched reference set is a feasible fix and would materially strengthen the paper. Note also that the human annotation pool is small (5 after quality filtering) with fair inter-annotator agreement (κ=0.375); the strength of the human confirmation should be described more cautiously."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The framework itself is the real contribution here. Adapting the semantic differential to attention-head activations, probing ~2,000 WordNet antonym axes, and localizing them to specific heads is a solid, reproducible engineering effort. The variance-ratio candidate selection, mass-mean probing, and the per-axis power analysis are all good practice. The position-prediction experiment is also a sane sanity check: LLMs recover Warmth/Competence labels above chance, and the Llama/Mistral agreement exceeding human agreement is an interesting finding even if not entirely new.\n\nBut the central empirical claim about discovering new stereotypical axes does not hold up as stated. The stereotypicality test in §4.4 compares projections of social groups (C) against random nouns and noun phrases (C′), and Table 9 shows those reference items are overwhelmingly inanimate: country house, tank ship, atomic bomb, capacitor. The KS test then flags any axis that applies to people but not to objects. Tables 10–11 prove the point directly: small/large, green/ripe, and rough/smooth all come out as well-powered stereotypical while human annotators mark them “not applicable.” That is not a stereotype artifact; it is an animacy artifact. The paper’s own Limitations section states that axes must be applicable to humans to be meaningful, but the reference set is not filtered for human-applicability. This confound also threatens the supposedly discovered axes like cowardly/brave and humble/proud: they may pass simply because C is animate and C′ is not. The fact that some human-applicable axes are not flagged (slow/fast) shows the test is not vacuous, but the numerical strength of the discovery is uncertain.\n\nTwo softer issues. First, the ablation in Appendix A selects the configuration that maximizes agreement with human ratings, and then the same human-agreement metric is reported as the main result. That is selection on the evaluation metric, which inflates the headline accuracy numbers. Not fatal, but it should be disclosed more prominently or validated on a held-out set. Second, the “discovery” is actually only over the 100-axis stratified sample, not the full ~2,000 candidate pool, so the title slightly overstates the scope.\n\nNone of this breaks H1/H2, which are about geometric encoding and projection-as-rating. H3 is the load-bearing weak point, and H4 outcomes are therefore conditional. The paper deserves a serious referee, but the referee should demand a matched reference set (e.g., animate or human-applicable noun phrases, or a control set of non-stereotypical human traits), plus word-embedding baselines and per-model significance for each discovered axis. With that revision, the framework could be genuinely useful for bias auditing and interpretability. I’d send it to review, not desk-reject it.","headline":"A genuinely useful probing framework whose headline stereotype-discovery results are undermined by an unmatched reference set: C′ is inanimate nouns, so the KS test may be measuring animacy rather than stereotypicality.","tokens_in":28611,"tokens_out":2914,"would_cite":true,"duration_ms":29816,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that stereotypes in large language models are encoded as geometric axes in the models' internal activation space, and that two different models agree on those axes more than either agrees with humans.","keywords":["stereotype discovery","semantic differential","probing","geometric axes","LLM internal representations","activation space","Kolmogorov-Smirnov test","social group stereotypes"],"falsifier":"Re-run STEREODISCO with a reference set matched on animacy and concreteness (e.g., other human-relevant words that are not social groups) and check whether axes such as small/large, green/ripe, and rough/smooth — which the current reference set flags as stereotypical while humans call them not applicable — still pass the p<0.05 test. If they do, the stereotypicality claim survives with a stronger reference; if they drop out, the current headline list partly depends on the reference-set artifact.","tokens_in":27546,"feed_emoji":"🧭","tokens_out":5175,"duration_ms":49057,"temperature":0.7,"pith_summary":"The paper's central claim is that a language model's stereotype content is not scattered across its parameters but is laid out as directions in its internal activation space: each pair of opposites, like cowardly versus brave, becomes a line, and a social group like 'CEOs' is located by projecting its neural representation onto that line. To test this, the authors build a discovery pipeline that starts from about 2,000 antonym pairs, recovers each pair as a geometric axis by probing attention-head activations, and then flags axes as stereotypical when social-group projections spread out significantly more than projections of random noun phrases. Applied to two 7–8 billion parameter instruction-tuned models, the pipeline finds that the models rate social groups more like each other (73–75% position agreement) than like human survey ratings (55–63%), and it surfaces axes such as cowardly/brave and narrow-minded/broad-minded that prior stereotype dictionaries omit. If the claim holds, audits of LLM bias can look for stereotypical axes beyond the few studied in social psychology, and the recovered directions can be used for targeted steering without extra prompting or fine-tuning.","feed_headline":"LLM stereotypes are geometric axes — and models share them more than people","feed_subtitle":"Two models rate social groups more like each other than like humans, and the method spots axes dictionaries miss.","key_machinery":"The load-bearing identity is the geometric axis: the direction θ = μ+ − μ− formed by subtracting the mean activation of one pole's sentences from the mean activation of the other pole's sentences at a given attention head. Because swapping poles only negates the direction, the axis itself is label-invariant. The projection score π(c) = x(c)ᵀθ/‖θ‖² places each concept along the axis, and the stereotypicality decision is a two-sample Kolmogorov–Smirnov test on the distribution of projections for the concept set versus a random reference set. What the machinery does is convert an intangible semantic opposition into a linear direction that can be located in specific attention heads and then stee","core_discovery":"STEREODISCO treats a semantic differential axis — a pair of opposite adjectives — as a candidate geometric axis in the model's activation space. For each of roughly 2,000 antonym pairs, it builds a probing dataset of sentences instantiating both poles, reads activations from every attention head, and takes the difference of the two pole means as the axis direction. Social-group mentions are then projected onto the top-scoring heads' axes and z-scored, and a Kolmogorov–Smirnov test compares the projections of 50 social groups with those of 50 frequency-matched random noun phrases. An axis is called stereotypical when the two distributions differ at p<0.05. The case study's empirical finding i","pith_inferences":["If the reference set were matched on human-relevance — replacing random nouns like 'capacitor' and 'ginger nut' with phrases that can describe people — several well-powered axes flagged as stereotypical (small/large, green/ripe, rough/smooth) would likely stop being flagged, because the distribution shift on those axes may just be the difference between social groups and objects.","The paper's H3 test treats any distribution shift as stereotypical, which means the method measures a relative property (axes along which social groups differ from random nouns), not an absolute property of the model. The same method applied with a different reference set (e.g., occupations vs. social groups) would discover different axes.","If the geometric-axis claim scales, the discovery pipeline could be applied to food, brands, or individuals, producing stereotype maps for any concept family that has a meaningful opposite set — but the paper only demonstrates social groups, so this is an extension, not a result.","A testable extension: use the discovered axes as steering directions and measure whether downstream generation shifts accordingly; that would connect the representational discovery to behavior."],"forward_implications":["Stereotype audits can move beyond the small set of warmth/competence dimensions: STEREODISCO surfaces axes like cowardly/brave that previous dictionaries miss, so bias evaluations that only use predefined axes will undercount the stereotypes an LLM actually encodes.","Because the stereotypical axis is a linear direction in specific middle-to-late attention heads, interventions can steer outputs along it without prompting or fine-tuning.","Cross-model agreement is not evidence of human alignment: the two models agree with each other more than with human ratings, so an LLM's 'consensus' stereotype content can diverge from documented human stereotypes.","The line between stereotypical and non-stereotypical axes is visible in the projection distributions: groups spread toward both poles on stereotypical axes and overlap with random phrases on non-stereotypical ones, giving a quantitative criterion for what counts as a stereotype axis."],"fun_headline_variants":["LLM stereotype axes are geometric — and models share them more than humans","Two LLMs rate social groups more like each other than like humans","Stereotypes in LLMs are geometric axes — models agree more than humans","LLMs share stereotypes more with each other than with humans","New stereotype axes found in LLMs: humble, narrow-minded, cowardly"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that a significant difference between how 50 social-group words and 50 random noun phrases project onto an axis measures stereotypicality; since the random nouns are not matched for human-relevance, the flagged set may partly reflect the general difference between people and objects rather than stereotypes about specific groups.","fun_headline_variants_meta":{"raw":{"variants":["LLM stereotype axes are geometric — and models share them more than humans","Two LLMs rate social groups more like each other than like humans","Stereotypes in LLMs are geometric axes — models agree more than humans","LLMs share stereotypes more with each other than with humans","New stereotype axes found in LLMs: humble, narrow-minded, cowardly"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001236,"raw_usage":{"total_tokens":4921,"prompt_tokens":766,"completion_tokens":4155,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":4061}},"tokens_in":510,"tokens_out":4155,"duration_ms":25135,"temperature":1.0,"reasoning_tokens":4061,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T00:43:50.598235+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run STEREODISCO with a reference set matched on animacy and concreteness (e.g., other human-relevant words that are not social groups) and check whether axes such as small/large, green/ripe, and rough/smooth — which the current reference set flags as stereotypical while humans call them not applicable — still pass the p<0.05 test. If they do, the stereotypicality claim survives with a stronger reference; if they drop out, the current headline list partly depends on the reference-set artifact.","supporting_citations":[],"review_version":1}