Pith. sign in

REVIEW 2 cited by

Prediction is not Explanation: Revisiting the Explanatory Capacity of Mapping Embeddings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.13729 v1 pith:QJPPFP22 submitted 2025-08-19 cs.CL cs.AIcs.LG

Prediction is not Explanation: Revisiting the Explanatory Capacity of Mapping Embeddings

classification cs.CL cs.AIcs.LG
keywords embeddingssemanticwordknowledgemethodspredictionencodedfeatures
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Understanding what knowledge is implicitly encoded in deep learning models is essential for improving the interpretability of AI systems. This paper examines common methods to explain the knowledge encoded in word embeddings, which are core elements of large language models (LLMs). These methods typically involve mapping embeddings onto collections of human-interpretable semantic features, known as feature norms. Prior work assumes that accurately predicting these semantic features from the word embeddings implies that the embeddings contain the corresponding knowledge. We challenge this assumption by demonstrating that prediction accuracy alone does not reliably indicate genuine feature-based interpretability. We show that these methods can successfully predict even random information, concluding that the results are predominantly determined by an algorithmic upper bound rather than meaningful semantic representation in the word embeddings. Consequently, comparisons between datasets based solely on prediction performance do not reliably indicate which dataset is better captured by the word embeddings. Our analysis illustrates that such mappings primarily reflect geometric similarity within vector spaces rather than indicating the genuine emergence of semantic properties.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery

    cs.AI 2026-07 unverdicted novelty 5.0

    Mechanistic World Models reframe AI scientific discovery as knowledge organisation around reusable explanatory mechanisms rather than predictive input–output mappings.

  2. From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery

    cs.AI 2026-07 accept novelty 5.0

    The paper proposes Mechanistic World Models — models organized as typed latent variables, a reusable mechanism library, and binding structures — as the route from AI forecasting to autonomous discovery.