{"id":"a026feab-d3c8-44fd-9535-129f627a1a59","arxiv_id":"2508.13729","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Predicting semantic feature norms from word embeddings is not evidence that the embeddings encode those features, because the same methods can predict random information just as well.","lead":"This paper challenges the common assumption that an embedding model accurately predicting human-labeled semantic features means the model truly encodes those meanings. It argues that prediction success can instead be explained by an algorithmic upper bound and geometric similarity, not by genuine semantic knowledge.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim hinges on the random-information baseline being statistically matched to real feature norms; the abstract does not establish this, so the upper-bound argument is not yet supported.","rationale":"The reader's weakest-assumption analysis, based on the abstract, identifies exactly the same condition: the random-information baseline must be statistically matched to real feature norms in all ways that affect prediction difficulty. This is the most load-bearing premise because the entire argument is a negative claim about semantic encoding supported by a control condition. If the control is mismatched, the central conclusion does not follow; if it is matched, the conclusion is strong. The abstract alone cannot establish which case holds, so the appropriate verdict remains UNVERDICTED. My stress-test pass does not identify a new concern beyond this, but it sharpens the required check: the matching must cover not only label statistics but also per-feature noise and annotator agreement, and the algorithmic upper bound must be derived independently of the evaluated accuracies. The proposed concrete test, a permutation-based matched control, would settle whether the concern lands. I agree with the reader's assessment and see no reason to change the verdict given the abstract-only evidence.","tokens_in":715,"tokens_out":1693,"duration_ms":19364,"concrete_test":"Inspect the paper's experimental protocol and check whether the random-information baseline is constructed by permuting the real feature vectors across words, thereby preserving marginal distributions, sparsity, dimensionality, and frequency, while only breaking the word-feature association. If such a permutation-based matched control is present and the reported high accuracy persists, the concern is resolved. If instead the random targets were generated independently (e.g., uniform random labels or random sparse vectors with different statistics), recompute the main result with matched random targets; if accuracy on matched random targets drops substantially below accuracy on real semantic features, the upper-bound claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that prediction accuracy on semantic feature norms does not indicate that embeddings encode those features, because the same mapping methods can 'successfully predict even random information.' For this conclusion to follow, the random targets must be matched to the real feature-norm targets in every aspect that affects prediction difficulty: feature dimensionality, per-feature sparsity, value distribution, word-frequency confounds, and annotator agreement (i.e., irreducible noise). The abstract does not state that such matching was performed. If the random targets are denser, lower-dimensional, or less noisy than real semantic features, then high accuracy on random targets is expected and says nothing about whether real features are encoded. Conversely, if the random targets are matched and accuracy remains at the same high level, the conclusion is substantially strengthened. A second aspect of the same concern is the 'algorithmic upper bound': it must be derived independently of the data used to evaluate it, not fitted to the observed accuracies. If the bound is fit post hoc, then the statement that results are 'predominantly determined by an algorithmic upper bound' becomes circular. Since the full text is unavailable, the existence of a matched control and an independently derived bound cannot be verified; these are the load-bearing premises on which the central claim rests.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper challenges the widely used assumption that high prediction accuracy when mapping word embeddings onto human-interpretable semantic feature norms demonstrates that those features are genuinely encoded in the embeddings. It reports that the same mapping methods can \"successfully predict even random information\" and argues that observed accuracy is mostly determined by an \"algorithmic upper bound\" rather than by semantic content. It concludes that comparisons between datasets based solely on prediction performance are unreliable and that such mappings mainly reflect geometric similarity in vector space.","tokens_in":947,"tokens_out":1595,"duration_ms":18579,"significance":"If the central claim is fully substantiated, the paper would have a substantial impact on interpretability evaluations of word embeddings and LLMs: it would invalidate a common evidential inference and require the community to redesign feature-norm probing experiments. The proposed random-information control is a sensible and potentially powerful experimental design, and the paper's explicit focus on a widely used but rarely challenged assumption is a strength. However, the abstract alone provides no data, no error bars, no comparison protocol, and no derivation of the algorithmic upper bound, so the significance is conditional on the full manuscript delivering on these points.","major_comments":[{"comment":"The central inference that \"successful prediction of random information\" implies that results are dominated by an algorithmic upper bound depends on the random targets being statistically matched to the real feature norms in every respect that affects prediction difficulty, including dimensionality, per-feature sparsity, value distribution, word-frequency confounds, and annotator agreement noise. The abstract does not state that such matching was performed. If the random targets are denser, lower-dimensional, or less noisy than the real semantic features, then high accuracy on random targets would be expected and would not show that the real feature accuracies are semantically uninformative. Please specify the exact construction of the random baseline and report the matching statistics in the manuscript; otherwise the main conclusion does not follow from the reported result.","section":"Abstract"},{"comment":"The claim that results are \"predominantly determined by an algorithmic upper bound\" appears to be offered as an explanation of the observed accuracies. For this explanation to be non-circular, the upper bound must be derived independently of the particular datasets and fitted accuracies it is used to explain, for example from the mathematical properties of the mapping algorithm and the target distribution. If the bound is estimated or fitted from the same experimental results, then the statement becomes a reformulation of the data rather than an explanation. The abstract provides no indication of how the bound was obtained, so the load-bearing derivation must be clearly presented in the full text.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase \"random information\" is ambiguous: it should be clarified what exactly is randomized (the feature labels, the feature values, the word-feature associations, or some combination) and at what level (features, items, or splits).","section":"Abstract"},{"comment":"The abstract does not state whether the random baseline is evaluated under the same protocol as the semantic feature experiments, including identical train/test splits, mapping dimensionality, and regularization; this should be stated to reassure readers that the comparison is apples-to-apples.","section":"Abstract"},{"comment":"The sentence about mappings \"primarily reflecting geometric similarity within vector spaces\" is a key interpretive claim but is not elaborated in the abstract; a brief explanation of the geometric mechanism would help readers understand the proposed alternative explanation.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This review is necessarily based on the abstract alone because the full text was not made available. The central hypothesis is plausible and the random-control design is a good idea, but the abstract does not provide enough information to verify the matched-control premise or the independent derivation of the algorithmic upper bound. I would need the full manuscript, particularly the methods and the construction of the random baseline, to render a confident recommendation. There is no sign of a fatal flaw in the stated approach, but the load-bearing details are exactly the ones missing from the available material."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper from the abstract alone: it argues that predicting semantic feature norms from word embeddings doesn't show the embeddings encode those features, because the same mapping methods can predict random targets just as well. That's a pointed and useful challenge to a common interpretability practice. The random-information control is the right kind of idea, and the authors deserve credit for proposing it. The conclusion that dataset comparisons based purely on prediction accuracy are unreliable also follows naturally if the control holds.\n\nWhat I cannot judge from the abstract is whether the control actually holds. The whole argument depends on the random targets being matched to the real feature norms in every way that affects prediction difficulty: dimensionality, sparsity, value distribution, word frequency, annotator noise. If the random targets are easier to predict, then matching accuracy proves nothing. The abstract does not say matching was done. That is not a flaw in the paper per se—it just means the load-bearing premise is invisible at this stage.\n\nThe second soft spot is the 'algorithmic upper bound.' The abstract says results are predominantly determined by this bound, but it doesn't say how the bound is derived. If it is computed from the same datasets and methods it is used to explain, the argument becomes circular. If it is derived independently, that is a strong result. Again, unverifiable from the abstract.\n\nSo this is a paper with a clear, coherent thesis and a sensible experimental design at the conceptual level, but with two load-bearing details I would want to see before believing the conclusion. It is not a refutation of anything yet; it is a promising line of attack.\n\nFor whom: anyone working on probing or interpretability of embeddings. The question matters, and the proposed control would be a useful addition to the toolkit even if the current experiment turns out to have flaws. I would bring it to a reading group if the full text were available, but not on the abstract alone.\n\nRecommendation: yes, send it to peer review. The editors should not desk-reject it. The referees should focus on the random-target matching and the derivation of the upper bound, and they need the full experimental protocol to judge.","headline":"A useful challenge to feature-norm probing, but the abstract's central claim rests on a random-target control and an algorithmic upper bound that cannot be assessed without the full text.","tokens_in":617,"tokens_out":705,"would_cite":false,"duration_ms":19114,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Word embeddings can predict random targets, so accuracy alone does not prove they encode meaning.","keywords":["word embeddings","feature norms","interpretability","probing","prediction accuracy","algorithmic upper bound","semantic representation","evaluation methodology"],"falsifier":"A study could construct random target sets that exactly match the statistical profile of a real feature norm (same number of features, same per-feature positive rate, same word frequency distribution, and same inter-annotator agreement) and then compare prediction accuracy. If real features are predicted substantially better than these matched random targets across many embeddings, the paper's central claim would be refuted.","tokens_in":545,"feed_emoji":"🎲","tokens_out":1728,"duration_ms":20371,"temperature":0.7,"pith_summary":"This paper challenges the widespread assumption that high prediction accuracy on human-annotated semantic features proves those features are encoded in word embeddings. The authors show that standard mapping methods can 'successfully' predict even random information, implying that observed accuracy is largely set by an algorithmic upper bound rather than by meaningful semantic content. Consequently, comparisons between feature datasets based solely on prediction scores cannot tell which dataset is genuinely better captured by the embeddings. The paper's point matters because it questions whether current interpretability tests for embeddings, and by extension LLMs, measure real knowledge or merely geometric regularities.","feed_headline":"Word embeddings 'predict' random data—so scores don't prove meaning","feed_subtitle":"Mapping methods hit high accuracy even on meaningless targets, suggesting an algorithmic ceiling rather than real semantic encoding.","key_machinery":"The central object is the mapping procedure that learns a function from word embedding vectors to collections of human-interpretable semantic features (feature norms). The paper's key mechanism is the 'algorithmic upper bound': a limit on prediction accuracy that arises from the mapping method itself and from statistical properties of the target set, independent of semantic content. This bound makes random targets predictable, so the method cannot distinguish meaningful semantic encoding from noise.","core_discovery":"The central claim is that prediction accuracy of mapping embeddings onto semantic feature norms is not reliable evidence of interpretability. The authors demonstrate that these methods achieve high accuracy even when the target features are random, concluding that results are predominantly determined by an algorithmic upper bound rather than by semantic representation in the embeddings. They further argue that such mappings primarily reflect geometric similarity within vector spaces, not the genuine emergence of semantic properties. If correct, this invalidates prior conclusions drawn from feature-prediction benchmarks about what embeddings know.","pith_inferences":["The same algorithmic-upper-bound critique likely applies to more complex nonlinear probes, where overfitting to random labels is an established risk, so the paper's argument strengthens the case for stricter null models throughout interpretability research.","A testable extension would be to check whether the upper bound scales with the dimension of the embedding space: if higher-dimensional embeddings predict random targets even better, the geometric-similarity explanation gains support.","The authors' reasoning implies that feature norms with low inter-annotator agreement may be particularly prone to producing spurious high accuracy, since their signal is closer to noise to begin with.","One could operationalize 'semantic encoding' as the excess accuracy over a well-matched random baseline; this paper provides a motivation for adopting that metric in future work."],"forward_implications":["Accuracy on semantic feature norms should no longer be treated as proof that embeddings encode those features.","Benchmark comparisons between feature datasets based on prediction performance alone are unreliable indicators of which dataset is better captured by the embeddings.","Future interpretability evaluations should include randomized control targets matched to real features to separate algorithmic upper bounds from genuine semantic signal.","The finding extends caution to any probing or mapping method that reports accuracy without a null baseline.","Claims about emergent semantic knowledge in LLMs that rely on feature-prediction probes may need to be revisited."],"supporting_citations":[],"fun_headline_variants":["High prediction scores on random targets: no real semantic signal","Embedding mappings hit ceiling on random data, not meaning","Predicting random features: why accuracy doesn't equal knowledge","Mapping embeddings: geometric similarity, not semantic truth","Can embeddings predict random info? Yes—so scores mislead"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the random information used as a control is statistically comparable to real semantic features in all ways that affect prediction difficulty, such as sparsity, frequency, dimensionality, and annotator agreement.","fun_headline_variants_meta":{"raw":{"variants":["High prediction scores on random targets: no real semantic signal","Embedding mappings hit ceiling on random data, not meaning","Predicting random features: why accuracy doesn't equal knowledge","Mapping embeddings: geometric similarity, not semantic truth","Can embeddings predict random info? Yes—so scores mislead"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000706,"raw_usage":{"total_tokens":3112,"prompt_tokens":808,"completion_tokens":2304,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":424,"completion_tokens_details":{"reasoning_tokens":2238}},"tokens_in":424,"tokens_out":2304,"duration_ms":19249,"temperature":1.0,"reasoning_tokens":2238,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:10:47.800115+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A study could construct random target sets that exactly match the statistical profile of a real feature norm (same number of features, same per-feature positive rate, same word frequency distribution, and same inter-annotator agreement) and then compare prediction accuracy. If real features are predicted substantially better than these matched random targets across many embeddings, the paper's central claim would be refuted.","supporting_citations":[],"review_version":2}