{"id":"0fe346af-ae9d-45f7-b4c7-bf8d35236f3a","arxiv_id":"2505.13561","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM success at general reasoning is best explained by the abstractness and data-efficiency of natural language, supporting the view that language itself enables domain-general inference.","lead":"This paper argues that the reasoning abilities of large language models support Daniel Dennett's claim that acquiring language fundamentally transforms a mind, because language compresses information so tightly that general inference becomes computationally tractable. It is a philosophical essay that uses AI results as a test case for a long-standing question about the role of language in human cognition.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Causal attribution to linguistic compression is confounded by scale, compute, and undisclosed architecture; the evidence does not isolate the representational medium.","rationale":"The reader's weakest_assumption already identifies the core confound: the observed difference between language-trained and non-language-trained AI systems may be due to scale, compute, and architecture rather than to language's representational properties. My stress-test agrees and sharpens the point by locating it in the central causal claim of §4. The paper explicitly acknowledges some of these limitations (footnote 13, §5.2), which is honest, but the central argument still depends on a causal attribution that the presented evidence cannot support. However, the paper is framed as a philosophical argument offering a possibility proof and falsifiable predictions, not as a controlled empirical study; its conclusion is appropriately hedged. Therefore the reader's CONDITIONAL verdict remains appropriate: the central claim is plausible and original, but its soundness is limited by the lack of controlled comparison and acknowledged confounds. I see no reason to move the verdict to REJECT or UNVERDICTED, because the argument is not internally inconsistent and the negative claim is explicitly falsifiable. The proposed concrete test would directly address the confound by manipulating the representational medium while holding scale and architecture fixed; until such evidence exists, the causal role of linguistic compression remains an open question rather than an established result. Credit where due: the paper's falsifiable predictions and candid acknowledgment of GPT-4's non-public components are genuine strengths, and they are consistent with a conditional rather than rejecting verdict.","tokens_in":18199,"tokens_out":5556,"duration_ms":67719,"concrete_test":"Train an open-weight transformer (Llama-3-8B class) on a non-linguistic but discrete, abstract token stream—for example, a large corpus of programmatically generated world-state descriptions or tokenized video frames reduced to abstract events—matched in token count and compute to a text-trained control, and evaluate both on the same battery of cross-domain reasoning tasks used to support §3's positive claim. If the non-linguistic variant reaches comparable reasoning, the claim that natural language's compression is uniquely enabling is falsified. If it fails despite matched scale, the language-specific explanation gains support; to separate compression from content, additionally vary text entropy (e.g., artificially compressing or permuting tokens) and measure reasoning performance as a function of compressibility.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central causal claim in §4 is that language's compression and abstraction make general inference computationally tractable, and this is inferred from the observation that only language-trained systems (LLMs) exhibit domain-general reasoning. The load-bearing problem is that the observation does not isolate the representational medium. LLMs differ from non-linguistic systems along several confounded dimensions: (i) training scale (trillions of tokens versus video or game corpora), (ii) compute, (iii) architecture (decoder-only transformers plus, per footnote 13, undisclosed non-public components in GPT-4), and (iv) training objective and post-training, such as RLHF and instruction tuning, which the paper does not discuss. The video-prediction comparison in §4 is a task/modality difference, not a controlled manipulation of compression. Thus the data are consistent with the alternative that LLM reasoning is driven by scale and architecture, with language serving merely as a discrete channel. The paper itself concedes in footnote 13 that GPT-4's inferential powers are 'likely due not solely to the next-token prediction using a transformer network but further, as yet non-public computational architecture,' and in §5.2 that for LLMs learning a language is not separable from learning large amounts of world information. These concessions directly undercut a clean causal attribution to language's representational properties. The falsifiable predictions in §3 target only the negative claim (no non-linguistic general reasoner) and would not disambiguate the mechanism: a non-linguistic system could fail while compression is still irrelevant, or succeed while language is still necessary. The central explanation is therefore underdetermined by the evidence presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that the success of large language models (LLMs) at domain-general reasoning, combined with the failure of non-linguistic AI systems to match it, provides empirical support for Daniel Dennett's thesis that adding language to a mind transforms it. The author defends a positive claim (LLMs exhibit powerful domain-general reasoning) and a negative claim (non-language-based AI systems do not), proposes an explanation in terms of the compression and abstraction of natural language making inference computationally tractable, and draws implications for cognitive science: language may be partly shaped for thought, may be a driver rather than a product of human cognitive powers, and LLMs serve as a possibility proof against an innate general language of thought.","tokens_in":18423,"tokens_out":5395,"duration_ms":51245,"significance":"If the central empirical claim were established, the paper would provide a novel and philosophically important argument connecting contemporary AI results to long-standing debates about language and thought. The paper is original in treating the 'Great AI Experiment' as evidence for the cognitive utility of linguistic compression, and it is commendably honest: it states a falsifiable prediction, acknowledges in footnote 13 that GPT-4's architecture is not fully public, and concedes in §5.2 that learning a language is not separable from learning world information. However, the argument rests on an empirical premise that is not adequately supported, and the proposed explanation is not independently operationalized. The paper is nevertheless a useful and provocative contribution for philosophers of language and cognitive science.","major_comments":[{"comment":"The negative claim that no non-language-based AI system exhibits domain-general reasoning is load-bearing, but it is supported only by absence of evidence and informal examples such as video prediction and game playing. Because the comparison classes differ from LLMs along multiple dimensions (training scale, compute, architecture, training objective, and post-training procedures such as instruction tuning), the observation does not isolate the representational medium. Footnote 13 concedes that GPT-4's inferential powers may rely on non-public computational architecture beyond next-token prediction with a transformer network, which further weakens the contrast. To make the premise credible, the paper should either control for these confounds or explicitly argue why scale and compute cannot explain the difference.","section":"§3 (The negative claim, pp. 14-16)"},{"comment":"The central explanation that linguistic compression makes inference computationally tractable is not operationalized. The War and Peace versus video comparison (p. 18) illustrates data rate for conveying a single proposition, but the tractability claim is about the difficulty of performing inference over possible continuations; the paper gives no measure of 'abstractness' or 'data-efficiency' and no independent account of why the number of possible words versus frames is the relevant variable. As stated, the explanation is not testable beyond the observation it is meant to explain; the paper should specify a formal or empirical criterion for compression and for inferential tractability.","section":"§4 (Natural language makes inference tractable, pp. 17-23)"},{"comment":"The concession that for LLMs 'learning a language is not separable from learning large amounts of information about the world encoded in language' undermines the clean causal attribution to the representational medium proposed in §4. If the relevant factor is the information content of linguistic data rather than its compressed format, then the LLM evidence supports a 'data channel' view, not the compression view. The paper should distinguish these alternatives and state what evidence would favor the compression explanation over the information-content explanation.","section":"§5.2 (Language as a driver rather than product, pp. 26-28)"},{"comment":"The paper claims to make a falsifiable empirical claim, but the stated criteria are too vague to test: a video prediction system 'systematically generate[s] the continuations of videos in a way that exhibits strong causal, agential and numerical reasoning' and game-playing systems with 'human-level capacities to get up to speed' are not defined with any benchmark or metric. The author should specify concrete evaluation tasks and thresholds so that the negative claim is actually falsifiable in practice.","section":"§3 (Falsifiable predictions, pp. 16-17)"}],"minor_comments":[{"comment":"BERT is described as an 'early LLM' and the same citation (Devlin et al. 2018) is given for both BERT and GPT-4; BERT is not a generative next-token-prediction model in the sense used in the paper, so the characterization and citation should be corrected.","section":"p. 3"},{"comment":"The citation to Mahowald et al. appears as a broken LaTeX artifact '(?, e.g)[MAHOWALD2024517' and should be fixed.","section":"p. 9"},{"comment":"'discreet' is used where 'discrete' is intended; this occurs, for example, in 'discreet representation' and 'discreet and efficient medium of representation.'","section":"pp. 17-18"},{"comment":"The bit-count comparison between the sentence and the video is illustrative but informally computed; the sentence is counted in raw ASCII while video size depends on compression, so the comparison should be flagged as order-of-magnitude rather than exact.","section":"p. 18"},{"comment":"'not a priori fact' should be 'not an a priori fact' (missing article).","section":"p. 20"},{"comment":"'langauge' is a typo for 'language.'","section":"p. 31"}],"recommendation":"major_revision","confidential_remarks":"This is a philosophy paper with a central empirical claim that is not yet established; the argument would be strengthened by engaging with scaling laws and the role of instruction tuning, and by including a response from an AI researcher. The paper is within the scope of the journal if the journal publishes work on philosophy of language and cognitive science. I recommend major revision: the philosophical argument is interesting, but the empirical premise and the explanatory mechanism need sharper formulation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a philosophical paper, not an empirical one, and it knows that. Its central thesis—that LLMs succeed at general reasoning because natural language's compression makes inference computationally tractable—is genuinely new and worth taking seriously. The argument is clear, the engagement with the literature is fair, and the author repeatedly flags where the evidence is weaker than the claim. That honesty is real and not just rhetorical.\n\nWhat works: the paper reframes the language-and-thought debate around the concept of computational tractability, which is a useful contribution. It gives a specific mechanism (data efficiency via abstraction) rather than a vague appeal to 'language transforms thought.' It also offers falsifiable predictions, even if they are weak. The discussion of Pinker's 'thick cables' point is sharp: compression might matter for inference even if communication bandwidth is not the bottleneck. And the paper is careful to distinguish LLM reasoning from human reasoning, acknowledging that LLMs lack innate structure and receive massive data.\n\nWhere it's soft: the load-bearing empirical premise—that only language-trained systems show domain-general reasoning—rests on anecdote and current AI system reports rather than a controlled comparison. The author concedes in footnote 13 that GPT-4's inferential powers may be due to non-public architecture, which undercuts the clean attribution to language. More fundamentally, LLMs differ from non-linguistic systems in scale, compute, training objective, and post-processing. The video-prediction comparison is illustrative but not a controlled manipulation of compression. So the observed success is consistent with the alternative that scale and architecture do the heavy lifting, with language just being the channel. The falsifiable predictions target only the negative claim (no non-linguistic general reasoner) and cannot disambiguate the mechanism. The paper acknowledges some of this in §5.2, but the acknowledgment shows the conclusion is more speculative than the earlier sections suggest.\n\nThe citation for Mahowald et al. is broken in the text, and there are typos, but those are minor.\n\nWho it's for: philosophers of cognitive science, and AI researchers interested in the philosophical implications of LLMs. It's a good discussion piece, not a definitive empirical study. It deserves serious peer review: a good reviewer will push for a more careful treatment of confounds and perhaps a more modest conclusion. I'd accept it with revisions, and I'd be happy to see it in print.","headline":"Original thesis on language as computational compression, but the causal evidence is underdetermined; worth publishing with revisions.","tokens_in":18986,"tokens_out":2521,"would_cite":true,"duration_ms":26684,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that the reasoning power of large language models is best explained by a property of natural language itself—compression and abstraction—and that this supports the view that adding language to a mind transforms it.","keywords":["large language models","language and thought","domain-general reasoning","linguistic compression","computational tractability","next-token prediction","cognitive science","language of thought hypothesis"],"falsifier":"A non-linguistic video-prediction system that, trained only on raw pixels, reliably generates continuations requiring causal, agential, and numerical reasoning, with no language-based component in its training or operation, would falsify the paper's central claim; so would a game-playing agent that reaches human-level competence on entirely new games after human-scale amounts of play.","tokens_in":17986,"feed_emoji":"💬","tokens_out":9098,"duration_ms":82602,"temperature":0.7,"pith_summary":"The paper reads the recent wave of AI systems as an experiment in cognitive science and draws a stark moral from it: among today's AI, only systems trained on natural language show powerful reasoning that spans many domains, while systems trained on video, games, or raw sensory data remain narrow specialists. The author's explanation is that language is an abstract, data-efficient encoding of the world, and that this compression is what makes general inference computationally tractable for a next-token-prediction network. If that is right, the success of large language models is empirical evidence for the radical thesis that adding language to a mind changes the kind of mind it is, not merely adds a communication channel. The paper also draws consequences for human cognition: language may be shaped partly for thought, and learning a language may itself unlock inferential abilities that a pre-linguistic mind does not have.","feed_headline":"Language itself is what unlocks AI's general reasoning","feed_subtitle":"If right, LLMs show that adding language changes the kind of mind a system has.","key_machinery":"The mechanism that carries the argument is the data-efficiency of linguistic encoding: natural language packs the information needed for inference into a compact, discrete, symbolic form, so a system that learns to predict text is solving a far smaller prediction problem than one that must predict raw video or act in a continuous sensorimotor world. The author compares a 30-second video clip to a text message or even War and Peace in bits, and argues that this compression, not expressive power alone, is what makes next-token prediction yield general reasoning. The central objects are the LLM itself—a transformer network trained by next-token prediction—and the abstract linguistic representations it consumes, which do the work of making inference computationally tractable.","core_discovery":"The paper's central claim is that the broad, domain-general reasoning shown by current large language models is made possible by the representational properties of natural language. Linguistic descriptions abstract away from irrelevant detail and encode the facts that matter for prediction in very few bits, so a network trained to predict text is effectively trained on a compressed, already-abstracted stream of information about the world. Non-linguistic AI systems, such as video-prediction or game-playing networks, must discover those abstractions themselves from high-dimensional raw data, and the paper contends this is why none of them has matched LLMs at general inference. The author takes this to support the thesis that adding language to a mind transforms that mind, and treats LLMs as a possibility proof that exposure to language alone can unlock substantial reasoning powers, while being careful that this does not by itself prove that human cognition works the same way.","pith_inferences":["If compression is the operative ingredient, then a controlled experiment training the same architecture on a non-linguistic but similarly abstract discrete encoding of events—for example, structured event logs—should produce more general reasoning than training on raw video, which would show that abstraction, not language per se, does the work.","The view predicts measurable human effects: adults trained on a formal symbolic system that compresses relational information should show improved cross-domain inference on problems expressible in that notation, a testable consequence the paper does not pursue.","Because the paper's negative claim about non-linguistic AI rests on absence of evidence, a natural extension is to build a benchmark battery of cross-domain non-verbal reasoning tasks and administer it to language-trained and non-language-trained models of matched scale, turning the philosophical argument into a quantitative comparison."],"forward_implications":["If the explanation is right, scaling up non-linguistic systems on raw video or sensor data will not by itself yield domain-general reasoning, because without an abstract encoding the prediction problem remains too large.","The success of large language models strengthens the case that a mental language is not merely a read-out of prior thought but can itself be a driver of new inferential capacities.","Human language may be shaped partly by its utility for thought, not only for communication, because compressed representational formats are valuable for inference inside a single head.","The argument implies that AI systems aiming at general reasoning should be built around abstract, symbolic representational media rather than raw high-dimensional input alone."],"supporting_citations":[{"why":"Supplies the thesis the paper tests: adding language to a mind changes the kind of mind it is.","marker":"Dennett 1996"},{"why":"Provides the transformer architecture whose next-token prediction is the training regime behind LLMs.","marker":"Vaswani et al. 2017"},{"why":"Documents the scale of LLM parameters and training tokens that underlies the empirical premise.","marker":"Touvron et al. 2023"},{"why":"Reports early evidence of GPT-4's broad reasoning capabilities across domains.","marker":"Bubeck et al. 2023"},{"why":"Distinguishes linguistic competence from reasoning and documents both LLM strengths and limitations, framing the positive claim.","marker":"Mahowald et al. 2024"},{"why":"States the opposing view that language is primarily a communication tool rather than a vehicle of thought.","marker":"Fedorenko et al. 2024"},{"why":"The language-of-thought hypothesis the author argues LLMs make less plausible.","marker":"Fodor 1975"},{"why":"Used to contrast human fast generalization with the slow, narrow learning of non-linguistic game-playing systems.","marker":"Lake et al. 2017"},{"why":"Provides the alternative view that language combines pre-existing domain-specific representations, which the paper partially adopts and modifies.","marker":"Spelke 2003"},{"why":"Documents the active research area of video prediction and its lack of general reasoning, supporting the negative claim.","marker":"Oprea et al. 2020"}],"fun_headline_variants":["Language unlocks LLM reasoning, paper claims","Why language makes AI inference tractable","LLMs as proof that language transforms thought","Paper: language gives AI its general reasoning","Language's efficiency drives LLM inference"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the gap in general reasoning between language-trained and non-language-trained AI systems is caused by language's compressed, abstract encoding rather than by the much larger scale of text data, greater compute, or non-public components in commercial models.","fun_headline_variants_meta":{"raw":{"variants":["Language unlocks LLM reasoning, paper claims","Why language makes AI inference tractable","LLMs as proof that language transforms thought","Paper: language gives AI its general reasoning","Language's efficiency drives LLM inference"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1206,"prompt_tokens":853,"completion_tokens":353,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":289}},"tokens_in":469,"tokens_out":353,"duration_ms":4117,"temperature":1.0,"reasoning_tokens":289,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:23:23.380053+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A non-linguistic video-prediction system that, trained only on raw pixels, reliably generates continuations requiring causal, agential, and numerical reasoning, with no language-based component in its training or operation, would falsify the paper's central claim; so would a game-playing agent that reaches human-level competence on entirely new games after human-scale amounts of play.","supporting_citations":[{"cited_title":"(2017) and Quilty-Dunn et al","cited_arxiv_id":null,"evidence_quote":"Used to contrast human fast generalization with the slow, narrow learning of non-linguistic game-playing systems."}],"review_version":1}