REVIEW 2 major objections 4 minor 10 references
The Concept of Representation in ML: Beyond Plato and Aristotle
T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Alignment evidence alone cannot prove that AI models converge on a shared reality.
desk verdict A sharp philosophical critique of PRH that misses the structural-realist escape route; worth peer review with a request to engage that gap. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the disjunction problem, the argument that mere causal covariance cannot determine which of many possible causes fixes a state's content, together with the teleosemantic response that content is fixed by a state's proper function and its role for downstream consumers. The paper applies this pair to kernel-alignment evidence: alignment measures relational geometry, but geometry alone cannot resolve which content, if any, the aligned states carry. The paper also uses Shea's functional account as a constructive route for attributing content to models without evolutionary history.
What would settle it
A controlled experiment showing strong kernel alignment between models trained on datasets with no shared task statistics, where alignment cannot be explained by shared statistical structure or spurious correlations, would undermine the paper's claim that convergence only evidences shared useful structure rather than reality.
Extended reading notes
Core claim
The paper claims that the Platonic Representation Hypothesis, as currently formulated, does not provide an account of what fixes representational content, and therefore its move from observed representational alignment to a metaphysical conclusion about a single underlying reality is unsupported. The authors argue that representational content requires more than covariance: it requires an account of how a state has determinate content in the face of error and noise, which is exactly Fodor's disjunction problem, and how internal states serve functions within the system, as teleosemantic theories emphasize. They suggest that a defensible version of the hypothesis would need to individuate content-bearing vehicles, explain their causal role and task function, and show how they distinguish veridical representation from spurious correlation.
Load-bearing premise
The critique depends on the premise that Fodor's disjunction problem, developed for biological minds, applies to the internal activations of large machine learning models; if one holds a purely instrumental view of activations, the argument loses its target.
Editorial extensions
If this is right
- Alignment results should be interpreted as evidence of shared useful structure, not convergence to reality.
- Any stronger claim about models representing reality will need a criterion for content-fixing that goes beyond covariance.
- Future work on the Platonic Representation Hypothesis should adopt task-based and functional analyses of representation.
- Mechanistic interpretability, by identifying downstream consumers and causal roles of activations, could supply the needed content-fixing evidence.
- Philosophical constraints can guide concrete ML research, such as controlling for spurious correlations in alignment measurements.
Reading between the lines
- If the argument is right, benchmark-based claims that larger models are 'more aligned with reality' should be reframed as claims about task-generalizability and shared inductive biases.
- The same disjunction-problem critique applies to other 'representation' claims in interpretability, such as linear probes and concept classifiers, not just alignment.
- A testable extension: alignment between models with different architectures, controlled for dimensionality, should be re-examined; residual alignment may vary by task family, predicting an Aristotelian pattern.
- The paper's use of Shea suggests a research program: define content for a model relative to a set of downstream tasks, making content-fixing empirically tractable.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper examines the Platonic Representation Hypothesis (PRH) of Huh et al. (2024), which claims that as AI models scale, their internal representations converge because they approximate a unified structure of reality. The paper argues that this claim moves from a legitimate engineering notion of representation to a philosophically loaded notion of mental representation, and that this move requires addressing the problem of content-fixing. It introduces Fodor's disjunction problem, teleosemantic theories (Millikan, Dretske), Davidson's Swampman, and Shea's learning-based teleosemantics, and then claims that kernel alignment methods, which measure covariance-based similarity between model activations, cannot resolve the disjunction problem. The paper also cites a follow-up result by Gr\"oger et al. (2026) suggesting that apparent alignment effects may partly reflect spurious correlations. The conclusion is that alignment evidence alone is insufficient for strong metaphysical claims, and that future work should adopt task-based and mechanistic analyses of representation.
Significance. If the argument is accepted, the paper provides a useful conceptual corrective to a prominent high-profile claim in machine learning. Its strengths include a generally accurate presentation of the philosophical sources, a clear and falsifiable target (PRH's inference from alignment to reality), and constructive suggestions for future research in task-based evaluation and mechanistic interpretability. The paper also draws on an independent empirical result by Gr\"oger et al. (2026), which strengthens the claim that the philosophical worry has practical consequences. However, the paper's significance is moderate: it is largely a conceptual critique, and its central argument relies on an unargued bridge from philosophy of mind to ML and on an unexamined assumption about the available philosophical routes from alignment to realism.
major comments (2)
- [§3 (Revisiting the Platonic Representation Hypothesis)] The central inference assumes that any warrant for 'models represent a unified reality' must be mediated by a solution to Fodor's disjunction problem. The paper does not engage with structural-realist accounts of representation, under which representation is understood as structure preservation (e.g., partial isomorphism) and for which kernel alignment would be directly relevant evidence for convergence to a common structure. Since Huh et al. (2024) explicitly situate PRH in relation to 'Convergent Realism,' a structural-realist reading is a live interpretation. The paper should either show that a structural-realist warrant is also undermined by the disjunction problem, or weaken its conclusion to apply only to content-based readings of PRH. As written, the argument defeats a content-based reading but not the strongest metaphysical claim it targets.
- [§3 (covariance-based methods and the disjunction problem)] The paper asserts that 'covariance-based methods ... cannot resolve the disjunction problem,' but this presupposes that model activations are the kind of state for which representational content is a meaningful question. On an instrumental or deflationary view of model activations, the disjunction problem has no target. The paper should provide an explicit argument, or at least a stated assumption, that the internal states of large models are candidate content-bearers. Without this bridge, the central step from philosophy of mind to ML is asserted rather than demonstrated.
minor comments (4)
- [§3 (Groger et al. discussion)] The discussion of Gr\"oger et al. (2026) is too terse: the paper says the authors 'prove mathematically' a growth in possible spurious correlations and that 'many previously reported alignment effects disappeared,' but gives no theorem statement, method, or quantitative result. Since this empirical result is used to show that the philosophical worry 'is not merely abstract,' please add a concise specification of the result and its relation to the conceptual argument.
- [Author line and references] There are several typographical errors: 'A viv Keren' should be 'Aviv Keren'; 'inThe Platonic Representation Hypothesis' is missing a space; 'Gr\"oger' should be typeset as 'Gr\"oger' (with proper umlaut); 'V oss' in the Brown et al. reference should be 'Voss'.
- [§1 (kernel alignment notation)] The notation for kernel alignment is inconsistent: the text uses 'Kf' and 'K g' where it should use K_f and K_g consistently. This makes the definition harder to follow.
- [§1 and elsewhere] The paper contrasts 'engineering' and 'philosophical' notions of representation but never gives a formal or precise working definition of the engineering notion. Adding an explicit characterization, even a stipulative one, would sharpen the contrast and the subsequent argument.
Circularity Check
No circularity: the paper's critique applies external philosophical standards to PRH and draws on independent empirical results; no fit, prediction, or self-citation chain reduces the argument to its own inputs.
full rationale
The paper makes no empirical predictions and fits no parameters. Its central argument is that kernel-alignment evidence, as used by Huh et al. (2024), cannot by itself fix representational content in the philosophically robust sense, citing Fodor's disjunction problem, teleosemantic theories, and Shea's functional account. These are external philosophical sources, not definitions imported from the present authors' prior work. The paper also invokes an independent technical result by Gröger et al. (2026) that controlled for activation-space growth and found that alignment effects largely disappeared; this is an external, falsifiable empirical result rather than a restatement of the paper's own conclusion. The critique is conditional: if PRH is read as a strong metaphysical claim, then it must address content-fixing; the paper does not define PRH's conclusion into existence. The only possible concern is that the argument assumes a content-based reading of representation and does not engage structural realism, but that is a philosophical gap or correctness risk, not a circularity. No equation is equated to another by construction, and no fitted quantity is renamed as a prediction. Accordingly, the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Representational content requires more than covariance or causal correlation.
- domain assumption Fodor's disjunction problem applies to the internal activations of trained neural networks.
- domain assumption Functional-role and teleosemantic theories provide the appropriate standard for representational content.
- domain assumption The empirical result of Groger et al. (2026) is correct.
Cite this review
Pith. "Pith review of The Concept of Representation in ML: Beyond Plato and Aristotle." pith.science (2026). https://pith.science/paper/IBTHZS7F
@misc{pith2026260717800,
author = {Pith},
title = {Pith review of: The Concept of Representation in ML: Beyond Plato and Aristotle},
year = {2026},
howpublished = {\url{https://pith.science/paper/IBTHZS7F}},
note = {Machine review of arXiv:2607.17800}
}
read the original abstract
Representation is a central concept in modern machine learning, where it usually refers to internal encodings that support learning and generalization. As models scale and their capabilities become increasingly human-level, this representational language sometimes shifts from an engineering context into the more philosophically loaded domain of mental representation. We argue that this is the case for recent claims about the convergence of representational properties across different AI models. In particular, we assess the arguments developed in The Platonic Representation Hypothesis, according to which this convergence is driven by a unified structure of reality. We examine this claim by introducing arguments and ideas from debates about mental representation in the philosophy of mind. We argue that these philosophical resources can clarify what is at stake in such claims, explain why alignment evidence alone is insufficient for strong metaphysical conclusions, and suggest directions for future research.
Reference graph
Works this paper leans on
-
[1]
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-V oss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford,...
work page 1901
-
[8]
URL https://distill.pub/2020/ circuits/zoom-in. Pitt, D. Mental representation. In Zalta, E. N. and Nodelman, U. (eds.),The Stanford Encyclopedia of Phi- losophy. Metaphysics Research Lab, Stanford Univer- sity,
work page 2020
-
[9]
Sucholutsky, I., Muttenthaler, L., Weller, A., Peng, A., Bobu, A., Kim, B., Love, B. C., Cueva, C. J., Grant, E., Groen, I., Achterberg, J., Tenenbaum, J. B., Collins, K. M., Hermann, K. L., Oktar, K., Greff, K., Hebart, M. N., Cloos, N., Kriegeskorte, N., Jacoby, N., Zhang, Q., Marjieh, R., Geirhos, R., Chen, S., Kornblith, S., Rane, S., Konkle, T., O’Co...
-
[10]
doi: 10.48550/arXiv.2310.13018. URL https://arxiv. org/abs/2310.13018. 6
-
[1981]
LeCun, Y ., Bengio, Y ., and Hinton, G
doi: 10.1086/288975. LeCun, Y ., Bengio, Y ., and Hinton, G. Deep learning.Na- ture, 521(7553):436–444,
-
[2018]
doi: 10.23915/distill.00010. 5 Representation in ML: Beyond Plato and Aristotle Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S. Zoom in: An introduction to circuits. Distill,
-
[2020]
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L. Sparse autoencoders find highly inter- pretable features in language models.arXiv preprint arXiv:2309.08600,
-
[2023]
doi: 10.48550/arXiv.2309. 08600. URL https://arxiv.org/abs/2309. 08600. Davidson, D. Knowing one’s own mind. InProceedings and Addresses of the American Philosophical Association, volume 60, pp. 441–458
Show all 10 references
-
[2024]
Laudan, L
URL https: //arxiv.org/abs/2405.07987. Laudan, L. A confutation of convergent realism.Philosophy of Science, 48(1):19–49,
-
[2026]
He, K., Chen, X., Xie, S., Li, Y ., Doll´ar, P., and Girshick, R
URLhttps://arxiv.org/abs/2602.14486. He, K., Chen, X., Xie, S., Li, Y ., Doll´ar, P., and Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16000–16009,
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.