{"id":"4d60a3f4-dd80-4d4c-b389-5738d62b9270","arxiv_id":"2412.04936","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Behavior-based word embeddings, especially from free associations, capture unique affective, agentic, and socio-moral information compared to text-based embeddings.","lead":"This paper compares word meanings derived from internet text, human behavior (like word association games), and brain scans, using 292 psychological rating scales. It finds that behavior-based word maps capture emotional and moral information that text-based maps miss, suggesting they could make AI language models more human-aligned.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Selection of the top text and behavior representations on the same 292 norms used for evaluation likely inflates the reported unique-variance gains in Section 4.3; the claim needs a hold-out-norm or split-half selection check.","rationale":"The paper's strongest and most field-relevant claim is that behavior-derived representations, specifically PPMI SVD SWOW, add unique variance over text on affective, agentic, and socio-moral norms. The reader already identified representation selection on the evaluation norms as a key unaddressed condition, and my stress-test narrows in on that condition as the single most load-bearing threat. If the top representations were selected on the full set of 292 norms before the same norms were used to compare ensembles, then every category-level advantage and significance test in Section 4.3 is subject to winner's-curse inflation. This is not merely a call for more conservative p-values; it changes the interpretation of the result from 'behavior generally complements text' to 'one behavior representation, chosen for high performance on these exact targets, shows localized advantages here.' The concern is concrete and testable with a split-half or hold-out-norm design. I do not see an internal inconsistency or a need to reject the paper: the metabase, RCA framework, vocabulary-matched probes, and nested cross-validation are genuine contributions, and the RSA differences in Section 4.1 are less affected by this selection issue. The brain-data result is appropriately cautious in the authors' own limitations discussion. For these reasons, the reader's conditional verdict is preserved, pending the selection-control analysis.","tokens_in":13312,"tokens_out":4942,"duration_ms":56263,"concrete_test":"Re-run the §4.3 ensemble analysis with representation selection isolated from evaluation: split the 292 norms by category or randomly into two halves; select the top-2 text and top behavior representations using only norms in the selection half (or a separate development set of norms), then compute the Text & Text vs Text & Behavior differences and Wilcoxon p-values on the held-out half. If the significant advantages on affect, agency, and Social/Moral vanish or shrink below the reported |d~| values, the §4.3 claim is inflated by test-set selection. As a robustness variant, repeat using all 10 behavior representations individually in ensembles (not just PPMI SVD SWOW) and report the distribution of gain sizes; only if the gain is consistently positive across behavior models and held-out norms does the complementarity claim generalize.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3's central comparison is vulnerable to selection on the evaluation set. The 'top-2 Text (CBOW GoogleNews and fastText CommonCrawl)' and 'top Behavior (PPMI SVD SWOW)' representations are chosen as the best-performing representations on the very same 292 psychNorms targets used to compute every reported ensemble difference. Because there are 26 representations and 292 norms, random per-norm fluctuations guarantee that the maximum-performing representation on the full norm set will show category-level advantages that do not generalize to a fresh set of norms or a different selection split. The Wilcoxon signed-rank p-values in Figure 5 compare within selected representations only; they do not account for the fact that the representations were selected to perform well on these targets. Thus the claimed 'unique variance' of behavior on Dominance, Arousal, Valence, Emotion, Goals/Needs, Motor, and Social/Moral may be an artifact of selection noise, not a general property of behavior-derived representations. The paper does not report any split-half, hold-out-norm, or pre-registered selection procedure, and its own limitations section addresses probe-training-set size but not representation selection.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a systematic comparison of semantic representations derived from text, behavior, and brain data, using representational similarity analysis (RSA) and a new interpretability framework called representational content analysis (RCA). The authors introduce psychNorms, a metabase of 292 word norms grouped into 27 categories, and probe 10 text, 10 behavior, and 6 brain representations against these norms. They report that behavior representations, particularly a PPMI-SVD representation trained on free associations, capture unique variance on affective, agentic, and socio-moral dimensions beyond text representations when combined in an ensemble. The paper also finds that brain representations contain little probe-able psychological content at the word level. The central claim is that behavior data complement text data for human-aligned semantic modeling.","tokens_in":13420,"tokens_out":5607,"duration_ms":53130,"significance":"If the central claim holds, the paper would provide the first large-scale evidence that behavior-derived word representations (e.g., from free associations) can complement text embeddings on psychologically meaningful dimensions, with implications for sentiment analysis, cognitive modeling, and LLM alignment. The psychNorms metabase and the RCA framework are valuable contributions, and the authors make code and data available. The nested cross-validation for probing is a technical strength. However, the selection-on-evaluation issue in Section 4.3 currently undermines the central claim of unique behavior variance.","major_comments":[{"comment":"The top-2 Text representations (CBOW GoogleNews, fastText CommonCrawl) and the top Behavior representation (PPMI SVD SWOW) are selected based on overall RCA performance on the same 292 norms that are then used to compute the ensemble differences. This selection-on-evaluation inflates the reported unique-variance gains on Dominance, Arousal, Valence, Emotion, Goals/Needs, Motor, and Social/Moral, because the representations were chosen to perform well on these very targets. The Wilcoxon signed-rank tests reported in Figure 5 compare within the selected representations only and do not account for the selection step. Please re-run the analysis with a hold-out-norm or split-half selection procedure, or otherwise provide evidence that the category-level advantages are not driven by selection noise.","section":"§4.3, Figure 5"},{"comment":"The 'norms sensorimotor' behavior representation is the same rating data (Lynott et al., 2020) that constitutes the Sensory and Motor norm categories in the psychNorms metabase. Probing this representation against those norms is circular and can drive the apparent behavior advantages on Motor and related categories. The paper should exclude this representation from RCA when its own norms serve as targets, or explicitly report results without it. This is especially important because Section 4.2's claim that 'the best-performing behavior representations perform comparatively strongly on ... Motor' may be entirely due to this circular case.","section":"§4.2, Table 1"},{"comment":"The Text & Text ensemble concatenates two text representations, while the Text & Behavior ensemble concatenates one text representation and one behavior representation. The feature dimensionality of these concatenated spaces is not matched, and linear probe performance can depend on the number of input features even with L2 regularization (the optimal regularization path and the effective model capacity change). The reported marginal R2 gains may therefore reflect the dimensionality difference rather than the unique semantic content of behavior features. Please match the dimensionality across ensembles (e.g., by subsampling features or using a fixed total dimension) or show that the results are robust to the feature-count difference.","section":"§4.3"},{"comment":"The behavior representations (e.g., free associations) and the norm ratings are both elicited from human participants. Common-method variance (e.g., shared response biases, word-frequency effects, or social desirability) could masquerade as 'unique psychological variance' attributed to behavior. The paper should at least discuss this alternative explanation and, ideally, include a control analysis using a behavior representation that is not based on human elicitation (e.g., eye-tracking) or partial out a general word-property factor.","section":"§4.2"}],"minor_comments":[{"comment":"The text '1/apha' should read '1/alpha'.","section":"§3.3"},{"comment":"The typo 'Wilxocon' should be 'Wilcoxon'.","section":"§4.3"},{"comment":"The proportion of top-3 nearest neighbors for text is reported as '.97%' but the intended value is likely '97%'; please correct the formatting.","section":"§4.1"},{"comment":"The exact size of the collective vocabulary intersection used for the ensemble analysis is not reported; please provide it, as the probe sensitivity depends on this.","section":"§4.3"},{"comment":"The caption mentions 'all Text & Text and Text & Behavior ensemble combinations' but the main text does not specify how many combinations were evaluated or how they were aggregated; please clarify.","section":"§4.3, Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is clearly written and the resource contributions (psychNorms, RCA code) are real. The main issue is the selection-on-evaluation in Section 4.3, which is a standard error that the authors should be able to fix with a split-half procedure. I recommend major revision rather than rejection because the central question is interesting and the data and code are provided. Please ensure the authors explicitly address the circularity of the 'norms sensorimotor' representation and the dimensionality mismatch in the ensemble comparison."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The psychNorms metabase and the systematic comparison of 26 text, behavior, and brain representations are the real contributions here. If you work on semantic representation evaluation, this is a handy resource and a sensible map of the terrain. The finding that free-association embeddings (PPMI SVD SWOW) hold their own against text on affective, agentic, and socio-moral norms is genuinely interesting and worth following up.\n\nThe methodology is mostly careful: nested cross-validation, vocabulary subsetting for the ensemble comparisons, and a cautious interpretation of the weak brain results. I'm not going to fault the brain section; they explicitly note the limitations of word-level extraction from sentence-level data.\n\nThe soft spot is load-bearing. The top text and top behavior representations are selected by performance on the same 292 norms that are then used for the ensemble comparison in Section 4.3. With 26 representations, the best performer on a given norm category will appear to add unique variance simply because it was chosen to do well on those targets. The Wilcoxon tests only compare the selected ensembles; they don't account for selection. The paper's own limitations section mentions probe-training-set size but not representation selection. A split-half or hold-out-norm selection check is needed before I'd trust the unique-variance numbers.\n\nA secondary issue: the Text & Text and Text & Behavior ensembles likely have different input dimensionalities, which can affect linear probe performance independent of content. That's a smaller concern but worth controlling. Shared method variance between behavior elicitation and human norms is a minor issue; most norms are independent of the free-association task, so I wouldn't overweight it.\n\nBottom line: this paper deserves serious peer review, but the central claim needs a selection-robustness analysis. If the unique variance survives hold-out-norm selection, the paper becomes quite strong. As is, treat the descriptive patterns as plausible but unproven.","headline":"A useful norm metabase and the broadest representation comparison I've seen, but the unique-variance claim is likely inflated by selecting the top representations on the same norms used for evaluation.","tokens_in":14062,"tokens_out":1782,"would_cite":true,"duration_ms":21005,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Behavior-derived word vectors capture psychological content that text embeddings miss, complementing them on affective, agentic, and socio-moral dimensions.","keywords":["semantic representations","representational similarity analysis","representational content analysis","word norms","free associations","embedding probing","human-aligned language models","psychNorms"],"falsifier":"A concrete falsifying observation: if the Text & Behavior advantage over Text & Text disappears when the top behavior representation is selected using only a training subset of the 292 norms (or when both ensembles are matched for feature dimensionality, or when norms are collected from a non-verbal/implicit task), then the claim that behavior captures unique psychological variance would be undermined. Alternatively, training PPMI SVD on a random reshuffling of the SWOW association matrix while preserving marginal frequencies, and showing it yields a similar R2 gain in the ensemble, would indicate the gain is not due to semantic content.","tokens_in":12978,"feed_emoji":"🧠","tokens_out":5765,"duration_ms":52236,"temperature":0.7,"pith_summary":"This paper asks whether word representations built from human behavior, such as free associations, and from brain recordings carry meaning that ordinary text-based embeddings miss. To answer it, the authors assembled a metabase of 292 human-rated word properties (psychNorms), introduced a probing framework called representational content analysis, and compared ten text, ten behavior, and six brain representations. They find that behavior-derived vectors, particularly PPMI SVD SWOW trained on the Small World of Words free-association data, rival text on many psychological dimensions and explain additional variance on affective, agency, and social/moral norms when added to text in an ensemble. The paper concludes that behavior is a practical complement to text for applications that aim to model human semantic representations, from sentiment analysis to evaluating large language models.","feed_headline":"Free-association vectors carry psychological content text misses","feed_subtitle":"Behavioral word vectors add unique signal to text embeddings on affective, agentic, and moral judgments.","key_machinery":"The machinery that carries the argument is representational content analysis (RCA): L2-regularized linear probes are fit to predict each of 292 norms from each word-vector representation, yielding a psychological content profile per representation. The decisive test is the ensemble RCA, which concatenates the top text and top behavior representations and compares the marginal increase in cross-validated R2 against a text-plus-text ensemble, with all vocabularies subset to a common intersection so training-set size is matched. The behavior representation of interest, PPMI SVD SWOW, is built by applying positive pointwise mutual information and singular value decomposition to the Small World of Words cue–response matrix.","core_discovery":"The central claim is that free-association-derived word vectors encode psychological information that text embeddings do not, and that this information is not redundant: ensembling the top text representation (CBOW GoogleNews or fastText CommonCrawl) with the top behavior representation (PPMI SVD SWOW) beats ensembling two text representations on affective norms (Dominance, Arousal, Valence, Emotion), agency norms (Goals/Needs, Motor), and social/moral norms, with median differences in R2 between 0.03 and 0.08, all significant at p < .05. The paper also reports representational similarity analysis showing clear clustering by data type rather than by learning algorithm, with behavior the most distinct of the three types. On this basis, the authors maintain that behavior representations, trained on orders of magnitude less data than text, are an important complement for measuring and modeling human representations.","pith_inferences":["The unique-variance finding could be tested more stringently by selecting the top representations on a training split of the norms and evaluating on a held-out split; the current analysis selects representations using the same 292 norms it then evaluates on, which may inflate the reported gains.","Because behavior vectors and norm ratings both come from the same kind of human subjects, some of the shared variance may reflect common response styles rather than semantic content; collecting norms through implicit or task-based measures would clarify whether behavior truly adds semantic information.","The brain representations' poor performance in this study may be a byproduct of small vocabularies and crude word-level extraction from sentence recordings rather than a property of brain data generally; better word-level brain representations might change the comparisons.","RCA applied to non-English norms could reveal whether the text–behavior complementarity extends across languages or is specific to English datasets."],"forward_implications":["Behavior-based semantic representations can supplement text embeddings in sentiment analysis, cognitive modeling, and other applications that depend on affective, agentic, or social-moral content.","The psychNorms metabase of 292 human-rated norms is a reusable resource for probing any word-level representation along psychologically meaningful dimensions.","Large language models trained or fine-tuned on structured behavioral data, such as free associations, could achieve better alignment with humans on psychological dimensions that text alone underrepresents.","Representational content analysis gives a general recipe for turning opaque word vectors into interpretable profiles, which can clarify what different models capture and where they diverge.","Expanding behavior-data collection efforts (e.g., more free associations) could bring behavior representations closer to text in coverage and performance, potentially improving their practical usefulness."],"supporting_citations":[{"why":"Supplies the Small World of Words free-association dataset from which the PPMI SVD SWOW behavior representation is trained.","marker":"De Deyne et al. (2019)"},{"why":"Provides the PPMI-and-SVD method for turning free-association cue–response counts into word vectors.","marker":"Richie and Bhatia (2021)"},{"why":"Applies the same PPMI SVD SWOW construction in the context of risk-perception prediction, a precursor to this paper's behavior representations.","marker":"Hussain et al. (2024b)"},{"why":"Source of the SCOPE norm metabase from which 97 norms were contributed to psychNorms.","marker":"Gao et al. (2023)"},{"why":"Contributes the 65 experiential attribute ratings that form part of the psychNorms metabase and a behavior representation.","marker":"Binder et al. (2016)"},{"why":"Provides the probing-classifier methodology that representational content analysis extends.","marker":"Belinkov (2022)"},{"why":"Establishes representational similarity analysis, the method used to compare information across data types.","marker":"Kriegeskorte et al. (2008)"}],"fun_headline_variants":["Free-association words carry psychological signal text misses","Behavior vectors add unique mind dimensions beyond text","Sparse behavior data yields word vectors with human norms","Free-association embeddings beat text on affective and moral norms","Text embeddings miss what free associations encode"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the extra variance explained by adding the free-association vectors to text vectors reflects genuinely unique semantic content, rather than artifacts of choosing the best representations on the evaluation norms, having different feature dimensionality between the compared ensembles, or shared method variance between human ratings and human-elicited associations.","fun_headline_variants_meta":{"raw":{"variants":["Free-association words carry psychological signal text misses","Behavior vectors add unique mind dimensions beyond text","Sparse behavior data yields word vectors with human norms","Free-association embeddings beat text on affective and moral norms","Text embeddings miss what free associations encode"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00039,"raw_usage":{"total_tokens":2039,"prompt_tokens":914,"completion_tokens":1125,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":1054}},"tokens_in":530,"tokens_out":1125,"duration_ms":12013,"temperature":1.0,"reasoning_tokens":1054,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:07:27.617890+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete falsifying observation: if the Text & Behavior advantage over Text & Text disappears when the top behavior representation is selected using only a training subset of the 292 norms (or when both ensembles are matched for feature dimensionality, or when norms are collected from a non-verbal/implicit task), then the claim that behavior captures unique psychological variance would be undermined. Alternatively, training PPMI SVD on a random reshuffling of the SWOW association matrix while preserving marginal frequencies, and showing it yields a similar R2 gain in the ensemble, would indicate the gain is not due to semantic content.","supporting_citations":[{"cited_title":"Small World of Words","cited_arxiv_id":null,"evidence_quote":"Supplies the Small World of Words free-association dataset from which the PPMI SVD SWOW behavior representation is trained."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the PPMI-and-SVD method for turning free-association cue–response counts into word vectors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the SCOPE norm metabase from which 97 norms were contributed to psychNorms."}],"review_version":1}