Pith. sign in

REVIEW 3 major objections 4 minor 54 references

On the Semantics of Large Language Models

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read LLM representations can carry Frege's sense, not reference—so whether they have semantics depends on the theory of meaning.

desk verdict A thoughtful Fregean mapping of LLM semantics, let down by an asserted empirical premise about cross-model representational identity; still worth refereeing. read the letter →

arxiv 2507.05448 v1 pith:KKVVONAX submitted 2025-07-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords semanticsFregesenseandreferencelargelanguagemodelsdistributedrepresentationslatentspacephilosophyofAItheorymeaning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether large language models can be said to mean anything, and answers that the question only makes sense relative to a theory of meaning. It argues that the distributed vector representations inside LLMs can plausibly instantiate Frege's notion of sense: a mode of presentation that is objective, shareable, context-dependent, and compositional, but independent of reference to the world. Because text-only LLMs have no direct reference, a theory that equates meaning with reference denies them semantics; the paper contends that this conclusion does not follow for a sense-based semantics. If the argument is right, LLMs can carry word- and sentence-level meaning even while lacking truth values and direct grounding, which reframes the 'octopus test' objection to LLM understanding. The stakes are practical as well as philosophical: how we judge LLM trustworthiness and hallucination depends on whether meaning is tied to reference or to sense.

What carries the argument

The load-bearing machinery is Frege's distinction between sense and reference, together with the eight properties of sense the paper extracts from Frege's text, and the LLM's latent space: the distributed, high-dimensional vector representation of tokens and sentences produced by the model. The paper maps the eight properties onto properties of vector representations—objectivity and shareability onto the claim that similar data and training yield similar representations across models, context-dependence onto the transformer's context-sensitive embeddings, sense-without-reference onto representations of fictional or non-existent terms, and compositionality onto the way sentence representations build from word representations. The vector space thus functions as the technical analog of the intermediate level Frege places between signs and objects, and it is the carrier of the paper's argument that LLMs can have a form of semantics.

What would settle it

Compare the internal word vector of the same word in the same sentence across two LLMs trained on similar but not identical data; if, after aligning for trivial symmetries, the representations show no common structure, the claim that latent space carries an objective, shareable Fregean sense is falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that the internal representation of language in an LLM—a point or region in a high-dimensional latent space—shares the defining properties of Frege's sense, though not his reference. A word's vector is a mode of presentation rather than an object; it is the same across contexts only when the context is the same, it is the common property of many rather than a private mental image, and the sentence's representation is determined by the representations of its parts, matching Frege's compositionality of sense. The paper is careful about scope: it does not claim LLMs have reference, truth values, or knowledge by acquaintance, and it concedes that attributing linguistic competence to the model is the one contestable step. What follows is that if meaning is understood as sense in addition to reference, LLM representations can carry that kind of semantics; if meaning is reduced to reference, they cannot. The resolution of the debate is therefore shifted from a factual question about model internals to a philosophical question about which kind of meaning is constitutive of language.

Load-bearing premise

The whole fit depends on the empirical premise that LLM vector representations actually are objective and shared enough—similar across models and contexts and unchanged by alignment—to play the role of Frege's sense, rather than being idiosyncratic model artifacts.

Editorial extensions

If this is right

  • If the analogy holds, an LLM can possess word- and sentence-level meaning even though its representations have no direct reference to the world, so the octopus-test conclusion loses force.
  • The question 'do LLMs understand?' becomes theory-dependent: under a sense-based semantics the answer can be yes, under a reference-only semantics no.
  • LLMs should be judged on their relation to truth and judgment rather than on whether individual word vectors refer; hallucination is then a failure of reference and truth, not of sense.
  • Frege's list of sense properties gives a concrete checklist for evaluating any LLM's semantic capacity, moving the debate from intuition to specific criteria.
  • The analysis also offers a technical handle on Frege's notoriously elusive notion of sense by giving it a mathematical home in latent space.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension: if the analogy is right, similarly trained models should converge to similar sense geometries, so representational-alignment measurements across models and random seeds could confirm or refute the objectivity claim.
  • Another consequence you could draw is that word-vector similarity or direction might serve as an empirical proxy for sense identity, letting researchers test Fregean claims about expressions like 'morning star' and 'evening star' in a distributional setting.
  • The paper's holism—sense arising from the whole training corpus—suggests a continuum between Fregean sense and distributional semantics, possibly implying that sense is graded rather than discrete.
  • If multimodal models ground representations in sensor data, the same argument might extend to a mediated form of reference, pushing LLM semantics beyond sense alone.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper examines whether large language models (LLMs) can be said to have semantics at the word and sentence level. It reviews core technical elements of LLMs (probabilistic language modeling, distributed representations, neural network architectures, scaling), then presents Russell's and Frege's classical theories of meaning. The central thesis is that LLM latent-space vector representations instantiate several properties of Frege's sense (Sinn): objectivity, shareability, context-dependence, compositionality, and independence from reference. Consequently, whether LLMs have semantics depends on the theory of meaning one adopts: under a purely referential theory they do not, but under a sense-inclusive theory they arguably do. The paper explicitly disclaims that LLMs possess reference or direct world-grounded semantics.

Significance. The paper offers a valuable conceptual bridge between classical philosophy of language and contemporary LLM research. It is clearly structured, engages primary Frege texts, and formulates a precise, conditional thesis that avoids both the naive 'stochastic parrot' dismissals and uncritical claims of human-like understanding. Its main strength is the careful enumeration of eight properties of Frege's sense and the initial mapping of these onto empirical observations about embeddings and contextual representations. However, the argument's force depends on empirical premises about the objectivity and cross-model stability of vector representations that are asserted rather than demonstrated. The paper is honest about its hedges, but these hedges pull the thesis back from 'instantiation' to 'similarity.' If the required empirical support were supplied, the contribution would be significant for the philosophy of AI; as it stands, it is a plausible but under-supported position paper.

major comments (3)
  1. [Section 2.5 and Section 4] The central Fregean property of being 'the common property of many' (Property 5, Section 3.1.2) requires that the same sense be shared across speakers. The paper tries to establish this for LLMs in Section 2.5 by asserting that 'when datasets and models are sufficiently large, models will likely generate similar representations' and 'the system will generate a representation similar to other models.' The supporting examples are behavioral: different LLMs give the same answer to 'What is the capital of France?' This conflates output-level functional similarity with representational identity. Two models can encode the same fact in non-isomorphic coordinate systems while producing identical text. The paper should either (a) cite representational-level evidence (e.g., linear-probe similarity, canonical correlation analysis, model stitching showing cross-model latent spaces are aligned under a natural correspondence), or (b) explicitly weaken the claim from 'instantiation' to 'analogy.' As written, the load-bearing premise that LLM vector spaces are objective and model-independent is unsupported.
  2. [Section 2.4] The claim that alignment (RLHF, etc.) 'does not alter significantly the internal representation but rather aims to control the output' is an empirical assertion that is central to the paper's thesis. If alignment changes the geometry of the latent space meaningfully, then the representations of commercially deployed LLMs are not simply 'somewhat independent of the specific model used' but are partly shaped by human preference data. The single footnote referring to Mollo & Millière (2023) for a contrary view is not sufficient engagement with the technical literature on the effects of RLHF on internal representations. The author should either supply supporting evidence or present this as a substantive assumption whose failure would weaken the argument. Without this, the 'objectivity' of the represented sense across actual LLMs remains unestablished.
  3. [Section 4] The paper asserts that the eight properties of Frege's sense enumerated in §3.1.2 are 'well reflected' in the vector space representation, but the mapping is not carried out point by point. In particular, compositionality (Property 8) is supported only indirectly by a reference to Lepori et al. (2023) on structural compositionality in neural networks; the paper does not show that the sentence-level representation is determined by the word-level representations in a way that matches Frege's compositionality of sense. Similarly, the claim that the representation is 'non-subjective' (Section 4) is asserted rather than argued. A systematic table or point-by-point discussion, distinguishing properties that are fully satisfied, partially satisfied, or merely analogous, would make the thesis easier to evaluate and would preempt the impression of cherry-picking properties to fit the empirical phenomena.
minor comments (4)
  1. [Section 3.1.2] The list of Fregean properties is drawn from a single primary text (Frege 1948) and the paper notes the debate about sense; however, the selection of these eight properties as definitive is motivated in part by the later comparison to LLMs. It would improve clarity to acknowledge that other readings of Frege would stress different aspects (e.g., sense as a mode of determining reference, as in Dummett), and to explain why the chosen list is the right one for the LLM comparison.
  2. [Section 2.5] The phrase 'it will always give the same answer' is too categorical; even a single LLM with stochastic sampling will not always give exactly the same answer. Rephrase as 'in typical greedy decoding, the answer is stable' or similar.
  3. [References] Correct the Turing entry ('Turing, M.' should be 'Turing, A.M.') and check the formatting of author initials throughout (e.g., 'V Y .' appears in several entries).
  4. [Section 4] The phrase 'the impact of the representation is revealed only through its use' seems to gesture toward a pragmatist or use-theoretic view, which is not developed. Either clarify the relation to Frege's sense (which for Frege is not defined by use) or remove the sentence.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper offers an interpretive comparison between Frege's sense and LLM latent-space representations, not a derivation that reduces to its own inputs.

full rationale

The paper's central claim is explicitly conditional and comparative rather than derivational: "If meaning however is associated with another kind of meaning such as Frege's sense in addition to reference, it can be argued that LLM representations can carry that kind of semantics" (Section 5). It lists eight properties of Frege's sense from primary sources (§3.1.2) and then argues that LLM vector-space representations "share certain properties with Frege's notion of sense" (§4). This is a philosophical analogy supported by qualitative evidence about embeddings, context-sensitivity, and compositionality, not a formal derivation in which a fitted parameter or defined quantity is later rediscovered as a prediction. The paper also explicitly hedges the identification: "This does not mean that we find exactly Frege's sense here" (§4). There are no fitted parameters, no load-bearing self-citations, no imported uniqueness theorems, and no renaming of a known result presented as a derivation. The most plausible circularity concern, that the conclusion is guaranteed by selecting the very properties used to define the analogy, does not rise to the level of a specific reduction because the paper does not claim to derive the properties; it claims only that the comparison is arguable and partial. The empirical premise that different LLMs share similar representations (§2.5) is asserted rather than demonstrated, but that is a correctness or evidential weakness, not circularity. Accordingly, no circular steps are identified.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim rests on interpretive assumptions about Frege's sense, empirical generalizations about LLM representations and training, and the premise that alignment does not change internal semantics. There are no fitted parameters or invented physical entities; the paper's burden is carried by unsupported empirical claims and a contested textual interpretation.

assumptions (4)
  • domain assumption Frege's sense is accurately characterized by the eight properties listed in Section 3.1.2: mode of presentation, possible sense without reference, objectivity, common property of many, context-dependence, and compositionality of sentence sense from part senses.
    The paper's comparison depends on this reading of Frege. The paper notes that 'there is no precise definition of sense, only descriptions' and that interpretations diverge (Dummett, Evans), but it says it 'adheres closely to Frege's own statements.' This is an interpretive choice, not a settled fact.
  • domain assumption LLM vector representations instantiate the Fregean sense properties such as objectivity, context-dependence, compositionality, and sense without reference.
    Section 4 asserts 'These points are well reflected in the vector space representation of language.' The paper provides illustrative arguments but no dedicated experiments; e.g., the objectivity claim rests on the assertion that similar datasets and training methods yield similar representations (Section 2.5).
  • domain assumption Alignment techniques such as RLHF do not alter the core semantic capabilities or internal representations of LLMs.
    Stated in Section 2.4: 'the core semantic capabilities of LLMs remain largely unaffected by current alignment techniques.' It is supported by a reference to earlier GPT models and a footnote citing a contrary view (Mollo and Millière 2023), but not by systematic evidence. This premise matters because the paper generalizes from base LLMs to assistants like ChatGPT.
  • domain assumption Text-based LLMs lack direct reference to the external world.
    This is the standard symbol-grounding premise taken from Bender and Koller and Harnad; the paper builds its sense-without-reference analysis on it without proving it for current models, noting multimodal models would still have mediated access.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the Semantics of Large Language Models." pith.science (2026). https://pith.science/paper/KKVVONAX

@misc{pith2026250705448,
  author       = {Pith},
  title        = {Pith review of: On the Semantics of Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KKVVONAX}},
  note         = {Machine review of arXiv:2507.05448}
}
read the original abstract

Large Language Models (LLMs) such as ChatGPT demonstrated the potential to replicate human language abilities through technology, ranging from text generation to engaging in conversations. However, it remains controversial to what extent these systems truly understand language. We examine this issue by narrowing the question down to the semantics of LLMs at the word and sentence level. By examining the inner workings of LLMs and their generated representation of language and by drawing on classical semantic theories by Frege and Russell, we get a more nuanced picture of the potential semantic capabilities of LLMs.

Figures

Figures reproduced from arXiv: 2507.05448 by the authors.

Figure 1
Figure 1. The original Transformer architecture with the encoder (left) - decoder (right) structure. Vaswani et al. [2017]. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

54 extracted references · 46 canonical work pages

  1. [1]

    Achiam, J. et al. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774

  2. [2]

    Anthropic. (2022). Claude: A Next-generation AI Assistant. https://claude.ai

  3. [3]

    Bedau, M. A. (1997). Weak emergence. Philosophical Perspectives , 11, pp. 375-399

  4. [4]

    & Shmitchell, S

    Bender, E.M., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , (pp. 610-623)

  5. [5]

    Bender, E. M. & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 5185--5198

  6. [6]

    & Vincent, P

    Bengio, Y., Ducharme, R. & Vincent, P. (2000). A neural probabilistic language model. Advances in Neural Information Processing Systems, 13

  7. [7]

    Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., & Zhang, Y. (2023). Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712

  8. [8]

    Burge, T. (1990). Frege on sense and linguistic meaning. The Analytic Tradition: Meaning, Thought, and Knowledge, 242-269

Show all 54 references
  1. [9]

    Burns, C., Ye, H., Klein, D., & Steinhardt, J. (2022). Discovering latent knowledge in language models without supervision. arXiv preprint arXiv:2212.03827

  2. [10]

    Chang, Y., et al. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3), 1-45

  3. [11]

    F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D

    Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems , 30

  4. [12]

    Conant, J. (1998). Wittgenstein on Meaning and Use. Philosophical Investigations , 21(3)

  5. [13]

    W., Lee, K., & Toutanova, K

    Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805

  6. [14]

    Dummett, M. (1981). Frege: Philosophy of Language . Harvard University Press

  7. [15]

    Eco, U. (1979). A Theory of Semiotics . Indiana University Press

  8. [16]

    Evans, G. (1986). The Varieties of Reference . Philosophy, 61(238)

  9. [17]

    Frege, G. (1879). Begriffsschrift, eine der arithmetischen nachgebildete Formelsprache des reinen Denkens. Louis Nebert, Halle a. S

  10. [18]

    Frege, G. (1884). Die Grundlagen der Arithmetik . Eine logisch mathematische Untersuchung über den Begriff der Zahl. Wilhelm Koebner, Breslau

  11. [19]

    Frege, G. (1891). Function und Begriff. Vortrag gehalten in der Sitzung vom 9. Januar 1891 der Jenaischen Gesellschaft für Medicin und Naturwissenschaft

  12. [20]

    Frege, G. (1892). Über Sinn und Bedeutung. Zeitschrift für Philosophie und philosophische Kritik . 100, 1892, S. 25–50

  13. [21]

    Frege, G. (1892). Über Begriff und Gegenstand. Vierteljahrsschrift für wissenschaftliche Philosophie . 16. Jahrgang, Nr. 2, 1892, S. 192–205

  14. [22]

    Frege, G. (1918). Der Gedanke: Eine logische Untersuchung. Beiträge zur Philosophie des Deutschen Idealismus . S. 58-77

  15. [23]

    Frege, G. (1948). Sense and Reference. The Philosophical Review , 57(3), 209-230

  16. [24]

    Gu, A., & Dao, T. (2023). Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752

  17. [25]

    Haack, S. (1978). Philosophy of Logics . Cambridge University Press

  18. [26]

    Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1-3), 335-346

  19. [27]

    Harris, Z. S. (1954). Distributional structure. Word, 10, 146-162

  20. [28]

    Hassija, V., et al. (2024). Interpreting black-box models: a review on explainable artificial intelligence. Cognitive Computation, 16(1), 45-74

  21. [29]

    Hinton, G. E. (1986). Learning distributed representations of concepts. Proceedings of the Eighth Annual Conference of the Cognitive Science Society . (Vol. 1, p. 12)

  22. [30]

    Jelinek, F. (1990). Self-organized language modeling for speech recognition. Readings in Speech Recognition, 450-506

  23. [31]

    & Martin, J.H

    Jurafsky, D. & Martin, J.H. (2024). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. Online manuscript released August 20, 2024

  24. [32]

    B., Chess, B., Child, R.,

    Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., ... & Amodei, D. (2020). Scaling laws for neural language models. arXiv preprint arXiv:2001.08361

  25. [33]

    LeCun, Y. (2023). Do large language models need sensory grounding for meaning and understanding? Philosophy of Deep Learning Talk, NYU, 3-24-2023

  26. [34]

    Lepori, M., Serre, T., & Pavlick, E. (2023). Break it down: Evidence for structural compositionality in neural networks. Advances in Neural Information Processing Systems, 36, 42623-42660

  27. [35]

    S., & Dean, J

    Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., & Dean, J. (2013). Distributed representations of words and phrases and their compositionality. Advances in Neural Information Processing Systems , 26

  28. [36]

    Mistral AI. (2023). Le Chat Mistral: An Advanced Conversational AI. https://chat.mistral.ai/

  29. [37]

    C., & Millière, R

    Mollo, D. C., & Millière, R. (2023). The vector grounding problem. arXiv preprint arXiv:2304.01481

  30. [38]

    OpenAI. (2022). ChatGPT: Optimizing Language Models for Dialogue. https://chatgpt.com/

  31. [39]

    Peirce, C. S. (1867). Writings of Charles S. Peirce: Volume 2, 1867-1871: A Chronological Edition . Bloomington, IN, Indiana University Press

  32. [40]

    Pennington, J., Socher, R., & Manning, C. D. (2014). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 1532-1543)

  33. [41]

    T., & Hill, F

    Piantadosi, S. T., & Hill, F. (2022). Meaning without reference in large language models. arXiv preprint arXiv:2208.02957

  34. [42]

    D., Ermon, S., & Finn, C

    Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2024). Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems , 36

  35. [43]

    M., & Laland, K

    Reader, S. M., & Laland, K. N. (2002). Social intelligence, innovation, and enhanced brain size in primates. Proceedings of the National Academy of Sciences, 99(7), 4436-4441

  36. [44]

    Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI blog, 1(8), 9

  37. [45]

    Rogers, A., Kovaleva, O., & Rumshisky, A. (2021). A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics , 8, 842-866

  38. [46]

    Russell, B. (1903). The Principles of Mathematics , Cambridge: Cambridge University Press

  39. [47]

    Russell, B. (1905). On denoting. Mind, 14(56) , pp. 479-493

  40. [48]

    Russell, B. (1912). Truth and Falsehood. The Problems of Philosophy . London: Williams and Norgate

  41. [49]

    Russell, B. (1918). The Philosophy of Logical Atomism: Lectures 1-2. The Monist , 28(4), 495-527

  42. [50]

    Schütze, H. (1992). Word space. Advances in Neural Information Processing Systems, 5

  43. [51]

    Shannon, C. E. (1951). Prediction and entropy of printed English. Bell System Technical Journal , 30(1), 50-64

  44. [52]

    Turing, M. (1950). COMPUTING MACHINERY AND INTELLIGENCE . Mind, Vol. LIX, 236 , (pp. 433–460)

  45. [53]

    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems , 30

  46. [54]

    Wittgenstein, L. (1958). Philosophical Investigations. Trad. G.E.M. Anscombe, Oxford, Basil Blackwell

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.