{"id":"91151914-12a8-4620-85c1-f315c18c579c","arxiv_id":"2412.14501","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Large language models process language in ways that align with inferentialist, anti-representationalist semantics, and they may be understood through a consensus theory of truth grounded in human feedback.","lead":"This paper argues that large language models are best understood through Robert Brandom's inferentialism, the view that meaning comes from roles in inference rather than from representing the world. It offers a philosophical frame for why chatbots like ChatGPT generate meaningful language without direct contact with reality.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The anti-representationalist conclusion depends on the unsupported equation of text-only training with linguistic idealism; text can transmit world-directed content, so the load-bearing premise fails.","rationale":"The paper is an honest, philosophically literate essay that clearly states its limitations; §9 in particular acknowledges unresolved tensions. However, the Abstract's strongest claim—that LLMs 'exhibit fundamentally anti-representationalist properties'—is not supported by the argument as written. The entire case for anti-representationalism reduces to the observation that text-only LLMs receive no direct sensory input (§7.1). That is the load-bearing premise because every other ISA-based alignment (inference, substitution, anaphora) is either explicitly qualified (§6.2 admits a mismatch on substitution) or merely analogical (attention is not anaphora). If the no-sensory-input premise does not entail the absence of world-directed content, the paper has not established anti-representationalism; it has only described a weak form of epistemic isolation. Empirical work showing world models in text-only networks would therefore refute the premise, while the lack of such work leaves the conclusion unsubstantiated. The reader's verdict already identified this exact weakness; my analysis agrees. The recommended test—linear decoding of a game state from a text-only transformer—is a direct, feasible check that would determine whether the premise holds. Pending that test, the appropriate verdict remains conditional: the paper should be reframed as proposing an inferentialist interpretation, not demonstrating that LLMs are fundamentally anti-representationalist.","tokens_in":21269,"tokens_out":5366,"duration_ms":45743,"concrete_test":"Train or use an existing text-only transformer on purely textual move sequences from the game of Othello (e.g., Othello-GPT; Li et al., 2023) and run a linear probe to decode the board state from hidden activations at each time step. If the world state (board configuration) is linearly decodable, then a language model with no sensory input nevertheless acquires internal world-directed representations, directly contradicting §7.1's claim that non-multimodal LLMs are 'devoid of any perceptual experience' and therefore linguistically idealist. A successful decoding would demonstrate that text alone can ground representational content, so the anti-representationalist conclusion would require a different argument. If decoding fails even in this controlled setting, the linguistic idealism premise gains support and the paper's conditional framing becomes more defensible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Abstract) that LLMs 'exhibit fundamentally anti-representationalist properties' rests on §7.1's identification of non-multimodal LLMs as 'linguistically idealist in the sense that they are devoid of any perceptual experience.' The load-bearing inference is from lack of sensory input to absence of world-directed representational content. This inference fails because training text is itself world-directed: it is produced by agents embedded in the world and is statistically correlated with states of affairs. Next-token prediction over such text can yield internal world models, as probing studies (e.g., board-game state decoding from purely textual move sequences) have shown. The authors' only argument against this is that 'there is no principled guarantee that the system maintains a representational relation with the world' (§7.1). But absence of a guarantee of reference is not absence of reference; it is at most epistemic fallibility, which also applies to humans. Moreover, the 'apples are rainbow-colored' thought experiment shows only that training data can be systematically intervened upon, not that ordinary text fails to carry worldly information. Because the paper offers no account of why text should cease to be about the world when consumed by a statistical learner rather than a human, the linguistic idealism premise is unsupported, and the anti-representationalist conclusion does not follow. The paper's own §9 tensions (propositionalism vs. sub-symbolic processing) further weaken the positive inferentialist reading, but the decisive gap is the premise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that Robert Brandom's inferential semantics, and its ISA (inference, substitution, anaphora) approach, provides a more appropriate foundational semantics for Transformer-based LLMs than distributional semantics or truth-conditional semantics. It claims that LLMs exhibit anti-representationalist properties, that non-multimodal LLMs are \"linguistically idealist\" because they lack perceptual input, and that a consensus theory of truth grounded in RLHF-style feedback is suitable for conversational LLMs. The paper acknowledges several tensions in Section 9, including the propositional commitments of inferentialism versus the sub-symbolic nature of LLMs.","tokens_in":21480,"tokens_out":5782,"duration_ms":46338,"significance":"If the central claim were established, the paper would make a meaningful contribution to the philosophy of LLMs by systematically connecting Brandom's frameworks to transformer architecture and by offering a philosophically grounded alternative to distributional semantics. The paper is commendable for stating its assumptions explicitly, including its reliance on a consensus theory of truth and its admission of unresolved tensions in Section 9. However, the significance is conditional: the main conclusion rests on an unsupported equation of text-only training with linguistic idealism, and the paper's own caveats undercut the strength of the word \"demonstrate\" in the abstract.","major_comments":[{"comment":"The paper's load-bearing inference is from the absence of perceptual sensors in non-multimodal LLMs to the claim that they are \"linguistically idealist\" and hence anti-representationalist. But the paper only argues that \"there is no principled guarantee\" that the system maintains a representational relation with the world (§7.1). Absence of a guarantee of reference is not absence of reference; it is at most epistemic fallibility, which applies to human cognition as well. Moreover, the authors do not address the well-known objection that text is produced by situated agents and is statistically correlated with states of affairs, so next-token prediction over text can yield internal world models. Because this premise is the decisive step for the abstract's central claim, the conclusion does not follow as stated.","section":"§7.1 and Abstract"},{"comment":"The paper cites \"factuality hallucination\" as \"consistent with the anti-representationalist nature of LLMs\" (§7.1). But a hallucination is defined precisely as a divergence from actual world facts; the notion of hallucination presupposes a standard of accuracy and hence world-directed content. Using hallucination as evidence of anti-representationalism is self-undermining, because it treats failures of representation as though they showed the absence of representation.","section":"§7.1"},{"comment":"The paper's own admissions undermine the \"demonstrate\" claim in the abstract. §6.2 concedes \"there is a mismatch between the inferentialist approach of extracting singular terms and predicates via substitutional inference and the way singular terms and predicates are handled within LLMs,\" and §9 concedes a \"fundamental mismatch\" between inferentialism's propositionalism and LLMs' sub-symbolic continuous processing. These concessions are not local; they concern the ISA mapping that is supposed to establish the anti-representationalist conclusion. The paper should be reframed as proposing a partial analogy or open hypothesis rather than demonstrating a result.","section":"§6.2 and §9"},{"comment":"The consensus theory of truth is asserted rather than derived. Equations (4) and (5) formalize reward-model training and KL-regularized policy optimization; nothing in this formalism shows that human preference feedback constitutes normative statuses, commitments, or scorekeeping in Brandom's sense. The conclusion that this theory is \"most suitable\" for conversational LLMs requires an explicit success criterion and an argument that RLHF instantiates normativity rather than mere optimization. As it stands, the analogy between RLHF and normative scorekeeping is asserted.","section":"§8"},{"comment":"The paper characterizes inferentialism as linguistically idealist and claims meaning is \"confined within language\" (§5.2), yet §9 notes that Brandom's inferentialism adopts conceptual realism, holding that the world is conceptually articulated. If Brandom's position includes world-responsiveness, then the anti-representationalist reading of LLMs cannot be directly exported from Brandom without addressing this tension. The paper should either argue that conceptual realism is not essential to the ISA account or explain how LLMs satisfy it.","section":"§5.2 and §9"}],"minor_comments":[{"comment":"\"as an suitable foundational semantics\" should be \"as a suitable foundational semantics.\"","section":"§1"},{"comment":"The concatenation/direct-sum symbol in Equation (1) is not rendered correctly in the text; the typesetting of the direct sum operator should be checked.","section":"§2, Eq. (1)"},{"comment":"\"The inner product between 'that' and 'pig' is large respectively\" is unclear; \"respectively\" is misplaced.","section":"§6.3"},{"comment":"\"In this section, I examine\" switches to first-person singular; use \"we\" for consistency with the rest of the paper.","section":"§8"},{"comment":"The reference \"V oita\" should be \"Voita,\" and \"thebert-base-uncased\" in §6.3 should be \"the bert-base-uncased.\"","section":"References"},{"comment":"The sentence about color and brightness is redundant and ambiguous; revise for clarity, for example by stating that the color indicates the sign and the brightness indicates the magnitude of the inner product.","section":"Figure 1 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the journal's readership in philosophy of AI, but the abstract overstates what the argument delivers. A major revision that weakens the claim to a conditional and engages with the world-model objection could make the paper publishable. The editors may also wish the authors to sharpen the distinction between their contribution and the earlier proposals by Havlík (2024) and Borg (2025)."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a genuine philosophical effort, not a dressed-up survey. It reads Brandom closely, distinguishes strong vs. hyper-inferentialism, deploys Quine against the Fodor-Lepore compositionality objection, and makes an honest attempt to connect inferentialist machinery to attention heads, induction heads, and RLHF. The novelty is real: Havlík and Borg suggested inferentialism as a candidate semantics for LLMs, but nobody has worked through the ISA triad (inference, substitution, anaphora) as systematically, nor proposed a consensus theory of truth grounded in RLHF. The paper also deserves credit for flagging its own obstacles in §9: propositionalism versus sub-symbolic processing, the privileging of assertion, and Brandom's conceptual realism. That is a fair-minded limitations section.\n\nThat said, the abstract's \"demonstrate\" is too strong. The load-bearing move in §7.1 is the identification of non-multimodal LLMs as \"linguistically idealist\" because they lack perceptual experience. The stress-test note is right: text is produced by agents embedded in the world and is statistically correlated with states of affairs, so next-token prediction can yield world-directed internal models—probing studies on board-game states from pure text are one example. The paper's \"no principled guarantee of representational relation\" is epistemic fallibility, not anti-representationalism. Humans also lack a guarantee that their representations track the world; that does not make humans anti-representationalists. The apples/rainbow thought experiment only shows training data can be manipulated, not that ordinary text ceases to be about the world.\n\nThe RLHF-to-truth analogy is also asserted rather than derived. Treating human preference scores as Brandomian \"sanctioning\" conflates reward maximization with normative commitment. A model optimizes a reward signal; it does not thereby endorse the correctness of its outputs in the space of reasons. This might be salvageable, but the paper would need to say much more about the distinction between improving a policy and undertaking a commitment.\n\nIs the paper worth a serious referee? Yes. It is a coherent, well-cited contribution to a live debate, and the ISA analysis is a useful scaffold even if the conclusion is not established. The conditional acceptance recommended by the reader is right: reframe the headline as a proposal, engage with the world-model literature, and either derive or visibly soften the RLHF-truth parallel. I would not cite it as established, but I would bring it to a reading group—it will generate an hour of good argument.","headline":"A serious, well-sourced inferentialist reading of LLMs whose central anti-representationalist claim overreaches: the step from text-only training to linguistic idealism does not hold, though the ISA analysis and RLHF-truth proposal are worth engaging.","tokens_in":22054,"tokens_out":2052,"would_cite":false,"duration_ms":20679,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Transformer-based LLMs should be understood through inferential semantics: they generate meaning through inferential roles and normative interaction rather than reference to external world representations.","keywords":["inferentialism","large language models","anti-representationalism","ISA approach","semantic externalism","compositionality","consensus theory of truth","linguistic idealism"],"falsifier":"A controlled experiment in which a text-only Transformer is trained on a corpus that contains no explicit statements about a hidden spatial layout, yet the model can nevertheless infer the layout from narrative structure and answer novel questions about it, would count against linguistic idealism. Similarly, ablating induction and suppression heads and observing whether substitution-based meaning survives would test the ISA mapping directly.","tokens_in":1709,"feed_emoji":"🤖","tokens_out":2539,"duration_ms":92586,"temperature":0.7,"pith_summary":"This paper argues that Transformer-based large language models are best understood not as systems that mirror the world but as systems whose linguistic meaning is generated by inferential roles and normative interaction among language users. It proposes inferential semantics, rather than distributional or truth-conditional semantics, as the foundational semantics for LLMs, and reads the models' architecture through the ISA approach (inference, substitution, anaphora). On this reading, LLMs are anti-representationalist and even linguistically idealist: because text-only models have no perceptual access to the world, their words relate only to other words. The paper also develops a consensus theory of truth for conversational LLMs, grounded in RLHF-style feedback as a normative practice. If the picture holds, standard assumptions such as strict compositionality and semantic externalism no longer apply to LLM meaning.","feed_headline":"LLMs make meaning by inference, not reference","feed_subtitle":"Text-only models can't point to the world; their words get meaning from inferential relations and human feedback.","key_machinery":"The ISA approach is the central lens: Inference, Substitution, and Anaphora are three ways inferentialism reinterprets representationalist concepts such as reference and truth. Inference gives words their content through material inference; substitution defines singular terms and predicates by symmetric and asymmetric replacement; anaphora handles demonstratives and coreference by pointing back into discourse. In the paper, the ISA mapping carries the argument: material inference maps to statistical pattern learning in Transformer weights, substitutional inference maps to induction and suppression heads that replicate or suppress earlier tokens, and anaphora maps to attention, as shown by the paper's attention-visualization example of 'it' attending to 'that pig.' The supporting machinery includes RLHF and DPO as scorekeeping mechanisms through which human preference feedback fixes what a model ought to say, and linguistic idealism, the thesis that a text-only model's world is confined to language.","core_discovery":"The central claim is that LLMs treat truth and reference as intra-linguistic devices rather than as relations to a mind-independent world. The paper demonstrates this through the three components of the ISA approach: inference in LLMs is material rather than formal, learned from statistical patterns in training data rather than encoded logical rules; substitution, which in inferentialism defines singular terms and predicates by symmetric and asymmetric inferential replacements, has no a priori counterpart in LLMs but may correspond to induction and suppression heads that replicate or suppress earlier tokens; anaphora is realized by attention and induction heads, so demonstratives like 'it' point back into the discourse rather than out to an object. Because training data and human preference feedback can be arbitrarily altered, the paper concludes that there is no principled relation connecting LLM outputs to worldly facts, making factuality hallucinations an expected consequence rather than a puzzle. It then argues that truth for conversational LLMs should be a consensus theory: what is correct is fixed by the normative feedback, including RLHF, shared between the model and its human or artificial interlocutors. The paper's own limitations section concedes an unresolved gap between inferentialism's commitment to discrete propositional content and LLMs' continuous sub-symbolic processing.","pith_inferences":["The paper leaves open that multimodal LLMs, which do receive perceptual input, would fall under strong inferentialism rather than full linguistic idealism; extending the ISA analysis to such models would sharpen where the anti-representationalist claim applies.","A testable extension is to probe for stable internal world models in text-only Transformers; robust evidence of such models would undercut pure linguistic idealism and push toward a hybrid semantics combining inferential roles with internal representations.","The consensus theory of truth implies that a model aligned with one community's preferences would inherit that community's norms; the paper does not address how conflicting human consensuses should be resolved when RLHF data are drawn from many populations."],"forward_implications":["If LLMs generate meaning through inferential roles, then claims like 'P is true' in an LLM output express commitment within a discourse rather than correspondence to a fact, so their truth is fixed by consensus in the interaction.","Strict compositionality, understood as the determination of complex meaning by constituent meanings, would not describe LLM semantics; meaning would instead be quasi-compositional, built from accumulated inferential patterns.","Semantic externalism, the view that meaning depends on facts outside the speaker, would fail for text-only models; their meanings would be determined by internal training distributions and interaction history.","Hallucination would not be a performance bug to be eliminated entirely but a structural consequence of a system whose content is not anchored to the world; mitigations would work by adjusting norms and training distributions.","Alignment techniques such as RLHF become constitutive of meaning rather than mere safety wrappers: they are the normative practice that fixes what the model ought to say."],"supporting_citations":[{"why":"Supplies the inferentialist framework, including the ISA approach, scorekeeping, and the normativity account the paper applies to LLMs.","marker":"Brandom (1994)"},{"why":"Provides the characterization of Transformers as distributional in method, the position the paper argues is insufficient and seeks to replace.","marker":"Grindrod (2024)"},{"why":"Defines the Transformer and self-attention architecture whose attention mechanism is mapped to anaphora and inferential relations.","marker":"Vaswani et al. (2017)"},{"why":"Supplies the InstructGPT/RLHF procedure that the paper interprets as normative scorekeeping and consensus-based truth.","marker":"Ouyang et al. (2022)"},{"why":"Advances the compositionality objection to inferential-role semantics that the paper re-frames using LLMs' inability to distinguish analytic from synthetic statements.","marker":"Fodor and Lepore (1991)"},{"why":"Provides the critique of the analytic/synthetic distinction that the paper uses to blunt the compositionality objection.","marker":"Quine (1951)"},{"why":"Gives the empirical basis for induction heads, which the paper links to substitution and anaphora in Transformers.","marker":"Olsson et al. (2022)"},{"why":"Provides the mechanistic framework for Transformer circuits used to describe heads as realizing inferential operations.","marker":"Elhage et al. (2021)"},{"why":"Presents the attention-visualization example that supplies the paper's concrete evidence of attention-based coreference resolution.","marker":"Vig (2019)"},{"why":"Supplies DPO, an alternative to RLHF, which the paper treats as another feedback-based normative mechanism.","marker":"Rafailov et al. (2023)"}],"fun_headline_variants":["LLMs get meaning from inference, not world","Hallucinations are normal: LLMs don't refer","Inferential semantics explains LLM meaning","LLM truth is consensus, not correspondence","LLMs infer meaning, don't refer to world"],"cache_read_input_tokens":24192,"weakest_assumption_plain":"The argument rests on the premise that a text-only model, having no perceptual contact with the world, can only relate words to other words; if text-only training can nonetheless build reliable internal models of the world, the anti-representationalist conclusion does not follow.","fun_headline_variants_meta":{"raw":{"variants":["LLMs get meaning from inference, not world","Hallucinations are normal: LLMs don't refer","Inferential semantics explains LLM meaning","LLM truth is consensus, not correspondence","LLMs infer meaning, don't refer to world"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000838,"raw_usage":{"total_tokens":3685,"prompt_tokens":1007,"completion_tokens":2678,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":623,"completion_tokens_details":{"reasoning_tokens":2621}},"tokens_in":623,"tokens_out":2678,"duration_ms":16533,"temperature":1.0,"reasoning_tokens":2621,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:10:24.536584+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled experiment in which a text-only Transformer is trained on a corpus that contains no explicit statements about a hidden spatial layout, yet the model can nevertheless infer the layout from narrative structure and answer novel questions about it, would count against linguistic idealism. Similarly, ablating induction and suppression heads and observing whether substitution-based meaning survives would test the ISA mapping directly.","supporting_citations":[],"review_version":1}