REVIEW 3 major objections 4 minor 54 references
On the Semantics of Large Language Models
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read LLM representations can carry Frege's sense, not reference—so whether they have semantics depends on the theory of meaning.
desk verdict A thoughtful Fregean mapping of LLM semantics, let down by an asserted empirical premise about cross-model representational identity; still worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is Frege's distinction between sense and reference, together with the eight properties of sense the paper extracts from Frege's text, and the LLM's latent space: the distributed, high-dimensional vector representation of tokens and sentences produced by the model. The paper maps the eight properties onto properties of vector representations—objectivity and shareability onto the claim that similar data and training yield similar representations across models, context-dependence onto the transformer's context-sensitive embeddings, sense-without-reference onto representations of fictional or non-existent terms, and compositionality onto the way sentence representations build from word representations. The vector space thus functions as the technical analog of the intermediate level Frege places between signs and objects, and it is the carrier of the paper's argument that LLMs can have a form of semantics.
What would settle it
Compare the internal word vector of the same word in the same sentence across two LLMs trained on similar but not identical data; if, after aligning for trivial symmetries, the representations show no common structure, the claim that latent space carries an objective, shareable Fregean sense is falsified.
Extended reading notes
Core claim
The paper's central claim is that the internal representation of language in an LLM—a point or region in a high-dimensional latent space—shares the defining properties of Frege's sense, though not his reference. A word's vector is a mode of presentation rather than an object; it is the same across contexts only when the context is the same, it is the common property of many rather than a private mental image, and the sentence's representation is determined by the representations of its parts, matching Frege's compositionality of sense. The paper is careful about scope: it does not claim LLMs have reference, truth values, or knowledge by acquaintance, and it concedes that attributing linguistic competence to the model is the one contestable step. What follows is that if meaning is understood as sense in addition to reference, LLM representations can carry that kind of semantics; if meaning is reduced to reference, they cannot. The resolution of the debate is therefore shifted from a factual question about model internals to a philosophical question about which kind of meaning is constitutive of language.
Load-bearing premise
The whole fit depends on the empirical premise that LLM vector representations actually are objective and shared enough—similar across models and contexts and unchanged by alignment—to play the role of Frege's sense, rather than being idiosyncratic model artifacts.
Editorial extensions
If this is right
- If the analogy holds, an LLM can possess word- and sentence-level meaning even though its representations have no direct reference to the world, so the octopus-test conclusion loses force.
- The question 'do LLMs understand?' becomes theory-dependent: under a sense-based semantics the answer can be yes, under a reference-only semantics no.
- LLMs should be judged on their relation to truth and judgment rather than on whether individual word vectors refer; hallucination is then a failure of reference and truth, not of sense.
- Frege's list of sense properties gives a concrete checklist for evaluating any LLM's semantic capacity, moving the debate from intuition to specific criteria.
- The analysis also offers a technical handle on Frege's notoriously elusive notion of sense by giving it a mathematical home in latent space.
Reading between the lines
- A testable extension: if the analogy is right, similarly trained models should converge to similar sense geometries, so representational-alignment measurements across models and random seeds could confirm or refute the objectivity claim.
- Another consequence you could draw is that word-vector similarity or direction might serve as an empirical proxy for sense identity, letting researchers test Fregean claims about expressions like 'morning star' and 'evening star' in a distributional setting.
- The paper's holism—sense arising from the whole training corpus—suggests a continuum between Fregean sense and distributional semantics, possibly implying that sense is graded rather than discrete.
- If multimodal models ground representations in sensor data, the same argument might extend to a mediated form of reference, pushing LLM semantics beyond sense alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper examines whether large language models (LLMs) can be said to have semantics at the word and sentence level. It reviews core technical elements of LLMs (probabilistic language modeling, distributed representations, neural network architectures, scaling), then presents Russell's and Frege's classical theories of meaning. The central thesis is that LLM latent-space vector representations instantiate several properties of Frege's sense (Sinn): objectivity, shareability, context-dependence, compositionality, and independence from reference. Consequently, whether LLMs have semantics depends on the theory of meaning one adopts: under a purely referential theory they do not, but under a sense-inclusive theory they arguably do. The paper explicitly disclaims that LLMs possess reference or direct world-grounded semantics.
Significance. The paper offers a valuable conceptual bridge between classical philosophy of language and contemporary LLM research. It is clearly structured, engages primary Frege texts, and formulates a precise, conditional thesis that avoids both the naive 'stochastic parrot' dismissals and uncritical claims of human-like understanding. Its main strength is the careful enumeration of eight properties of Frege's sense and the initial mapping of these onto empirical observations about embeddings and contextual representations. However, the argument's force depends on empirical premises about the objectivity and cross-model stability of vector representations that are asserted rather than demonstrated. The paper is honest about its hedges, but these hedges pull the thesis back from 'instantiation' to 'similarity.' If the required empirical support were supplied, the contribution would be significant for the philosophy of AI; as it stands, it is a plausible but under-supported position paper.
major comments (3)
- [Section 2.5 and Section 4] The central Fregean property of being 'the common property of many' (Property 5, Section 3.1.2) requires that the same sense be shared across speakers. The paper tries to establish this for LLMs in Section 2.5 by asserting that 'when datasets and models are sufficiently large, models will likely generate similar representations' and 'the system will generate a representation similar to other models.' The supporting examples are behavioral: different LLMs give the same answer to 'What is the capital of France?' This conflates output-level functional similarity with representational identity. Two models can encode the same fact in non-isomorphic coordinate systems while producing identical text. The paper should either (a) cite representational-level evidence (e.g., linear-probe similarity, canonical correlation analysis, model stitching showing cross-model latent spaces are aligned under a natural correspondence), or (b) explicitly weaken the claim from 'instantiation' to 'analogy.' As written, the load-bearing premise that LLM vector spaces are objective and model-independent is unsupported.
- [Section 2.4] The claim that alignment (RLHF, etc.) 'does not alter significantly the internal representation but rather aims to control the output' is an empirical assertion that is central to the paper's thesis. If alignment changes the geometry of the latent space meaningfully, then the representations of commercially deployed LLMs are not simply 'somewhat independent of the specific model used' but are partly shaped by human preference data. The single footnote referring to Mollo & Millière (2023) for a contrary view is not sufficient engagement with the technical literature on the effects of RLHF on internal representations. The author should either supply supporting evidence or present this as a substantive assumption whose failure would weaken the argument. Without this, the 'objectivity' of the represented sense across actual LLMs remains unestablished.
- [Section 4] The paper asserts that the eight properties of Frege's sense enumerated in §3.1.2 are 'well reflected' in the vector space representation, but the mapping is not carried out point by point. In particular, compositionality (Property 8) is supported only indirectly by a reference to Lepori et al. (2023) on structural compositionality in neural networks; the paper does not show that the sentence-level representation is determined by the word-level representations in a way that matches Frege's compositionality of sense. Similarly, the claim that the representation is 'non-subjective' (Section 4) is asserted rather than argued. A systematic table or point-by-point discussion, distinguishing properties that are fully satisfied, partially satisfied, or merely analogous, would make the thesis easier to evaluate and would preempt the impression of cherry-picking properties to fit the empirical phenomena.
minor comments (4)
- [Section 3.1.2] The list of Fregean properties is drawn from a single primary text (Frege 1948) and the paper notes the debate about sense; however, the selection of these eight properties as definitive is motivated in part by the later comparison to LLMs. It would improve clarity to acknowledge that other readings of Frege would stress different aspects (e.g., sense as a mode of determining reference, as in Dummett), and to explain why the chosen list is the right one for the LLM comparison.
- [Section 2.5] The phrase 'it will always give the same answer' is too categorical; even a single LLM with stochastic sampling will not always give exactly the same answer. Rephrase as 'in typical greedy decoding, the answer is stable' or similar.
- [References] Correct the Turing entry ('Turing, M.' should be 'Turing, A.M.') and check the formatting of author initials throughout (e.g., 'V Y .' appears in several entries).
- [Section 4] The phrase 'the impact of the representation is revealed only through its use' seems to gesture toward a pragmatist or use-theoretic view, which is not developed. Either clarify the relation to Frege's sense (which for Frege is not defined by use) or remove the sentence.
Circularity Check
No significant circularity: the paper offers an interpretive comparison between Frege's sense and LLM latent-space representations, not a derivation that reduces to its own inputs.
full rationale
The paper's central claim is explicitly conditional and comparative rather than derivational: "If meaning however is associated with another kind of meaning such as Frege's sense in addition to reference, it can be argued that LLM representations can carry that kind of semantics" (Section 5). It lists eight properties of Frege's sense from primary sources (§3.1.2) and then argues that LLM vector-space representations "share certain properties with Frege's notion of sense" (§4). This is a philosophical analogy supported by qualitative evidence about embeddings, context-sensitivity, and compositionality, not a formal derivation in which a fitted parameter or defined quantity is later rediscovered as a prediction. The paper also explicitly hedges the identification: "This does not mean that we find exactly Frege's sense here" (§4). There are no fitted parameters, no load-bearing self-citations, no imported uniqueness theorems, and no renaming of a known result presented as a derivation. The most plausible circularity concern, that the conclusion is guaranteed by selecting the very properties used to define the analogy, does not rise to the level of a specific reduction because the paper does not claim to derive the properties; it claims only that the comparison is arguable and partial. The empirical premise that different LLMs share similar representations (§2.5) is asserted rather than demonstrated, but that is a correctness or evidential weakness, not circularity. Accordingly, no circular steps are identified.
Assumptions & free parameters
assumptions (4)
- domain assumption Frege's sense is accurately characterized by the eight properties listed in Section 3.1.2: mode of presentation, possible sense without reference, objectivity, common property of many, context-dependence, and compositionality of sentence sense from part senses.
- domain assumption LLM vector representations instantiate the Fregean sense properties such as objectivity, context-dependence, compositionality, and sense without reference.
- domain assumption Alignment techniques such as RLHF do not alter the core semantic capabilities or internal representations of LLMs.
- domain assumption Text-based LLMs lack direct reference to the external world.
Cite this review
Pith. "Pith review of On the Semantics of Large Language Models." pith.science (2026). https://pith.science/paper/KKVVONAX
@misc{pith2026250705448,
author = {Pith},
title = {Pith review of: On the Semantics of Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/KKVVONAX}},
note = {Machine review of arXiv:2507.05448}
}
read the original abstract
Large Language Models (LLMs) such as ChatGPT demonstrated the potential to replicate human language abilities through technology, ranging from text generation to engaging in conversations. However, it remains controversial to what extent these systems truly understand language. We examine this issue by narrowing the question down to the semantics of LLMs at the word and sentence level. By examining the inner workings of LLMs and their generated representation of language and by drawing on classical semantic theories by Frege and Russell, we get a more nuanced picture of the potential semantic capabilities of LLMs.
Figures
Reference graph
Works this paper leans on
-
[1]
Achiam, J. et al. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[2]
Anthropic. (2022). Claude: A Next-generation AI Assistant. https://claude.ai
work page 2022
-
[3]
Bedau, M. A. (1997). Weak emergence. Philosophical Perspectives , 11, pp. 375-399
work page 1997
-
[4]
Bender, E.M., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , (pp. 610-623)
work page 2021
-
[5]
Bender, E. M. & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 5185--5198
work page 2020
-
[6]
Bengio, Y., Ducharme, R. & Vincent, P. (2000). A neural probabilistic language model. Advances in Neural Information Processing Systems, 13
work page 2000
-
[7]
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., & Zhang, Y. (2023). Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712
arXiv 2023
-
[8]
Burge, T. (1990). Frege on sense and linguistic meaning. The Analytic Tradition: Meaning, Thought, and Knowledge, 242-269
work page 1990
Show all 54 references
-
[9]
Burns, C., Ye, H., Klein, D., & Steinhardt, J. (2022). Discovering latent knowledge in language models without supervision. arXiv preprint arXiv:2212.03827
2022 arXiv
-
[10]
Chang, Y., et al. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3), 1-45
2024
-
[11]
F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems , 30
2017
-
[12]
Conant, J. (1998). Wittgenstein on Meaning and Use. Philosophical Investigations , 21(3)
1998
-
[13]
W., Lee, K., & Toutanova, K
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805
2018 arXiv
-
[14]
Dummett, M. (1981). Frege: Philosophy of Language . Harvard University Press
1981
-
[15]
Eco, U. (1979). A Theory of Semiotics . Indiana University Press
1979
-
[16]
Evans, G. (1986). The Varieties of Reference . Philosophy, 61(238)
1986
-
[17]
Frege, G. (1879). Begriffsschrift, eine der arithmetischen nachgebildete Formelsprache des reinen Denkens. Louis Nebert, Halle a. S
-
[18]
Frege, G. (1884). Die Grundlagen der Arithmetik . Eine logisch mathematische Untersuchung über den Begriff der Zahl. Wilhelm Koebner, Breslau
-
[19]
Frege, G. (1891). Function und Begriff. Vortrag gehalten in der Sitzung vom 9. Januar 1891 der Jenaischen Gesellschaft für Medicin und Naturwissenschaft
-
[20]
Frege, G. (1892). Über Sinn und Bedeutung. Zeitschrift für Philosophie und philosophische Kritik . 100, 1892, S. 25–50
-
[21]
Frege, G. (1892). Über Begriff und Gegenstand. Vierteljahrsschrift für wissenschaftliche Philosophie . 16. Jahrgang, Nr. 2, 1892, S. 192–205
-
[22]
Frege, G. (1918). Der Gedanke: Eine logische Untersuchung. Beiträge zur Philosophie des Deutschen Idealismus . S. 58-77
1918
-
[23]
Frege, G. (1948). Sense and Reference. The Philosophical Review , 57(3), 209-230
1948
-
[24]
Gu, A., & Dao, T. (2023). Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752
2023 arXiv
-
[25]
Haack, S. (1978). Philosophy of Logics . Cambridge University Press
1978
-
[26]
Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1-3), 335-346
1990
-
[27]
Harris, Z. S. (1954). Distributional structure. Word, 10, 146-162
1954
-
[28]
Hassija, V., et al. (2024). Interpreting black-box models: a review on explainable artificial intelligence. Cognitive Computation, 16(1), 45-74
2024
-
[29]
Hinton, G. E. (1986). Learning distributed representations of concepts. Proceedings of the Eighth Annual Conference of the Cognitive Science Society . (Vol. 1, p. 12)
1986
-
[30]
Jelinek, F. (1990). Self-organized language modeling for speech recognition. Readings in Speech Recognition, 450-506
1990
-
[31]
& Martin, J.H
Jurafsky, D. & Martin, J.H. (2024). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. Online manuscript released August 20, 2024
2024
-
[32]
B., Chess, B., Child, R.,
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., ... & Amodei, D. (2020). Scaling laws for neural language models. arXiv preprint arXiv:2001.08361
2020 arXiv
-
[33]
LeCun, Y. (2023). Do large language models need sensory grounding for meaning and understanding? Philosophy of Deep Learning Talk, NYU, 3-24-2023
2023
-
[34]
Lepori, M., Serre, T., & Pavlick, E. (2023). Break it down: Evidence for structural compositionality in neural networks. Advances in Neural Information Processing Systems, 36, 42623-42660
2023
-
[35]
S., & Dean, J
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., & Dean, J. (2013). Distributed representations of words and phrases and their compositionality. Advances in Neural Information Processing Systems , 26
2013
-
[36]
Mistral AI. (2023). Le Chat Mistral: An Advanced Conversational AI. https://chat.mistral.ai/
2023
-
[37]
C., & Millière, R
Mollo, D. C., & Millière, R. (2023). The vector grounding problem. arXiv preprint arXiv:2304.01481
2023
-
[38]
OpenAI. (2022). ChatGPT: Optimizing Language Models for Dialogue. https://chatgpt.com/
2022
-
[39]
Peirce, C. S. (1867). Writings of Charles S. Peirce: Volume 2, 1867-1871: A Chronological Edition . Bloomington, IN, Indiana University Press
-
[40]
Pennington, J., Socher, R., & Manning, C. D. (2014). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 1532-1543)
2014
-
[41]
T., & Hill, F
Piantadosi, S. T., & Hill, F. (2022). Meaning without reference in large language models. arXiv preprint arXiv:2208.02957
2022 arXiv
-
[42]
D., Ermon, S., & Finn, C
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2024). Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems , 36
2024
-
[43]
M., & Laland, K
Reader, S. M., & Laland, K. N. (2002). Social intelligence, innovation, and enhanced brain size in primates. Proceedings of the National Academy of Sciences, 99(7), 4436-4441
2002
-
[44]
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI blog, 1(8), 9
2019
-
[45]
Rogers, A., Kovaleva, O., & Rumshisky, A. (2021). A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics , 8, 842-866
2021
-
[46]
Russell, B. (1903). The Principles of Mathematics , Cambridge: Cambridge University Press
1903
-
[47]
Russell, B. (1905). On denoting. Mind, 14(56) , pp. 479-493
1905
-
[48]
Russell, B. (1912). Truth and Falsehood. The Problems of Philosophy . London: Williams and Norgate
1912
-
[49]
Russell, B. (1918). The Philosophy of Logical Atomism: Lectures 1-2. The Monist , 28(4), 495-527
1918
-
[50]
Schütze, H. (1992). Word space. Advances in Neural Information Processing Systems, 5
1992
-
[51]
Shannon, C. E. (1951). Prediction and entropy of printed English. Bell System Technical Journal , 30(1), 50-64
1951
-
[52]
Turing, M. (1950). COMPUTING MACHINERY AND INTELLIGENCE . Mind, Vol. LIX, 236 , (pp. 433–460)
1950
-
[53]
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems , 30
2017
-
[54]
Wittgenstein, L. (1958). Philosophical Investigations. Trad. G.E.M. Anscombe, Oxford, Basil Blackwell
1958
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.