{"id":"492a29e8-5781-4404-b57a-a5508ec82103","arxiv_id":"2606.07722","paper_version":4,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A conceptual argument that LLM chatbots simulate only a reduced, 'onward' form of thinking and cannot be scaled into reliable analytical thinking partners.","lead":"This paper argues that chatbots are stuck at System 1-style thinking: they follow learned text patterns and cannot become reflective, analytical thinking partners, no matter how large the models get. It introduces a new conceptual frame, 'metaphorical problem propagation', to explain the limitation and warns against treating chatbots as serious problem-solving partners.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unmeasured 'onward text' premise and untested reconstruction hypothesis carry the impossibility claim; the paper itself disclaims ability to assess them (Sec. 10.5).","rationale":"Good-faith reading: the paper is explicitly a speculative essay introducing 'metaphorical problem propagation' as a bridge between Cognitive Linguistics, Aggregation Dynamics, Predictive Processing, and LLM behavior. Its conclusion is the strong negative claim that not only are current basic chatbots System-1-like, but no text-only training, even on analytical texts, can make them analytical thinking partners. For that conclusion to hold, it must be true that (i) the bulk of pretraining text is 'onward' in the paper's sense and (ii) the training dynamics necessarily embed an onward, System-1-only structure that cannot be overcome by other text. Both are asserted, not established. The reader's verdict of REJECT is appropriate. The present stress-test converges on the same weakest assumption: the dataset-characterization premise is a postulate (literally 'We now assume...' in Sec. 9) with no measurement. The paper's own admission in Sec. 10.5 — that the authors cannot assess the value of their model in the relevant technical literature — is an explicit missing-support passage that strengthens the reader's concern. To its credit, the paper does cite empirical studies (e.g., [17,48]) that are consistent with LLMs encoding discourse-level situation structure, and it clearly labels much of its own contribution as hypotheses. But consistency with two studies is not sufficient to establish the categorical impossibility claim. The cited studies do not measure 'onwardness' nor test whether alternative training distributions produce analytical behavior. No change of verdict is needed.","tokens_in":25676,"tokens_out":7868,"duration_ms":87162,"concrete_test":"Sample a stratified random set of documents from an actual pretraining corpus (e.g., C4 or The Pile) and have multiple annotators apply a pre-registered 'onwardness' rubric operationalized from the traits listed in §8.1 (goal-directed, solution-oriented, non-reflective, comfort-seeking). If the proportion of onward documents is not a clear majority (or inter-annotator agreement is low), the premise 'most texts in the dataset are of the onward type' (§9) fails, and the conclusion that chatbots are irredeemably System 1 loses its empirical foundation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion — no basic chatbot is or ever will be an analytical thinking partner — rests on two empirically unvalidated premises: (1) most LLM training text is 'onward text' (subjective traits listed in §8.1, asserted as 'we now assume that most of the texts ... are of the previously postulated onward type' in §9), and (2) LLM training reconstructs these as 'artificial metaphorical problem propagations' whose onward character confines the chatbot to System 1 (Sections 9–10). Premise (1) is never operationalized or measured; the cited survey [50] says narrative texts are abundant, but narrative ≠ onward. Premise (2) is asserted rather than derived; the cited evidence [17,48] shows LLM representations track discourse situations and brain responses, not that the encoded structures are irrevocably System-1-like. The paper's own §10.5 admits: 'we are too far removed from this highly technical field of research to assess the value of our model.' Moreover, §10.4's claim that even an analytical training set cannot help depends on an unargued necessity of embodied experience for analytical reasoning; no evidence rules out text-only acquisition of System-2-like competencies. The impossibility claim thus outruns the evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a speculative account of why basic LLM-based chatbots cannot be analytical thinking partners. It introduces 'metaphorical problem propagation' (MPP) as a synthesis of Aggregation Dynamics, Cognitive Linguistics, and Predictive Processing. It then argues that most LLM training text is 'onward text' (System 1-like, solution-oriented, non-reflective), that LLM training reconstructs 'artificial MPPs' from this text, and that consequently chatbots are confined to System-1-like thinking and cannot become analytical partners even with larger models or analytical training sets. The conclusion is an impossibility claim.","tokens_in":26001,"tokens_out":3701,"duration_ms":38001,"significance":"If the impossibility claim were established, it would be a significant contribution to the debate on LLM reasoning and human-AI interaction. The paper deserves credit for attempting to connect interpretability research, cognitive linguistics, and dual-process theory into a structured model, and for being explicit about its speculative character. However, the manuscript as it stands is a hypothesis sketch rather than a demonstration: the key premises are asserted, not measured, and the authors themselves disclaim the ability to assess their central model (§10.5). The paper's value lies in proposing questions and a vocabulary, not in providing a validated answer.","major_comments":[{"comment":"The premise that most training texts are 'onward text' is load-bearing but never operationalized. §8.1 lists subjective traits ('strategically oriented', 'tailored to an audience seeking comfort') with no metric; §9 simply says 'We now assume that most of the texts ... are of the previously postulated onward type.' The cited survey [50] states that narrative texts are abundant, but 'narrative' is not equivalent to 'onward' — indeed, narrative theory includes reflective, non-linear forms. Without empirical support for the dominance of onward text, the System-1 conclusion collapses.","section":"§8.1 and §9"},{"comment":"The inference from LLM training to 'artificial metaphorical problem propagation' is not derived. The evidence [17,48] shows that LLM activations map onto human brain responses and can represent evolving discourse situations; it does not show that these representations are 'onward' or lack reflective potential. The paper's own §10.5 admits 'we are too far removed from this highly technical field of research to assess the value of our model.' A central claim cannot rest on a model whose validity the authors disclaim.","section":"§9 and §10.5"},{"comment":"The claim that even an analytical training set cannot produce an analytical thinking partner depends on an unargued necessity of embodied experience for System-2 thinking. No evidence is provided that text-only training cannot acquire reflective competence; indeed, publications such as [68] indicate that prompting and training can elicit multi-step reasoning. The strong negative conclusion therefore outruns the support.","section":"§10.4"},{"comment":"Circularity: §10 opens 'Suppose an LLM encodes metaphorical problem propagation, as argued above,' and then derives the System-1 conclusion from that supposition. Since the 'onward' character was built into the MPP hypothesis in §8–9, the conclusion is a restatement of the hypothesis rather than an independent result. The conclusion should be framed as a conditional, not a categorical claim.","section":"§9–§10"}],"minor_comments":[{"comment":"'Bereska en Gavves' should be 'Bereska and Gavves' (Dutch 'en' in otherwise English text).","section":"§2.2"},{"comment":"Typo: 'Consulted om May 29, 2026' should be 'Consulted on May 29, 2026'.","section":"Reference [20]"},{"comment":"The term 'aggregation' is used extremely broadly ('jealousy of the gods' as a player); a formal definition or at least a clearer scope condition would help the reader follow the later argument.","section":"§3"},{"comment":"The last line of Figure 3, 'The', appears to be a fragment; either complete the sentence or remove the stray article.","section":"Figure 3"}],"recommendation":"reject","confidential_remarks":"The paper is more of a position piece than a research article. Its core claims are presented as hypotheses and the authors repeatedly hedge, but the abstract and conclusion state a categorical impossibility. If the journal is open to speculative contributions with clearly labeled conjectures, a major revision reframing the conclusion as a conditional hypothesis might be considered. However, given the unmeasured central premises and the authors' own admitted distance from the relevant technical research, I do not see a path to acceptance as a scientific paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a well-structured, honest speculative essay, not a research result. The authors repeatedly flag the load-bearing steps as hypotheses, and in 10.5 they admit they are too far removed from interpretability research to assess the value of their own model. That candor is refreshing, but it also means the central conclusion — no basic chatbot can be an analytical thinking partner, and no amount of text-only training changes that — is a thesis, not a derivation.\n\nWhat is actually new is the packaging: 'metaphorical problem propagation' weaves together the cognitive-linguistic conceptual system, Aggregation Dynamics' problem-solution networks, and Predictive Processing's prediction dynamics; 'onward text' names the largely goal-directed, non-reflective prose that plausibly dominates LLM training corpora. The paper makes that picture vivid, and the Los Angeles smog example is genuinely useful. It also engages with concept-based and mechanistic interpretability work, and it correctly acknowledges that the System 1/System 2 conclusion is already present in [56] and [80]. The contribution, then, is a structured restatement of a known position with a new vocabulary.\n\nThe soft spots sit at the load-bearing joints. The onward-text premise is asserted, not demonstrated: Section 8.1 lists subjective traits and then says 'We refer to this as onward text'; Section 9 simply opens 'We now assume that most of the texts ... are of the previously postulated onward type.' That assumption carries the whole argument, and the cited survey [50] shows only that narrative texts are abundant, not that they are irreflexive or System 1-like. Second, the reconstruction hypothesis — that LLM training encodes artificial metaphorical problem propagations — is supported only by loose analogies to [17] and [48], which show that LLM representations track discourse situations and brain responses. That evidence does not establish that the encoded structures are irredeemably System 1. Third, the reasoning is structurally circular: Section 10 opens 'Suppose an LLM encodes metaphorical problem propagation, as argued above,' but 'above' was an assumption. Finally, the claim that an analytical training set could not help depends on an unargued necessity of embodied experience for analytical thought; the authors do not rule out text-only acquisition of System 2-like competencies, they assert it cannot work.\n\nWho should read this? People thinking about whether LLM-based chatbots can ever become genuine reasoning partners. The 'onward text' hypothesis is testable in principle, and that is a real virtue. I would not cite it as evidence, but I would bring it to a reading group and I would send it to a venue that welcomes speculative cognitive claims, with reviewers instructed to push for operationalization. As a referee I would recommend major revision: sample or measure the text distribution, state the reconstruction hypothesis in falsifiable form, and retreat from a categorical impossibility claim to a conditional one.","headline":"A clearly written speculative essay that is honest about its own status, but the central impossibility claim rests on an unmeasured 'onward text' premise and an untested reconstruction hypothesis; worth a serious discussion, not a conclusive result.","tokens_in":26437,"tokens_out":2793,"would_cite":false,"duration_ms":32661,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A chatbot, as a conversation partner, is not and cannot become an analytical thinking partner through text-only training.","keywords":["chatbots","large language models","metaphorical problem propagation","onward text","System 1 and System 2 thinking","analogical reasoning","predictive processing","problem-solving conversations"],"falsifier":"Compile a training corpus of deliberately analytical, reflective texts — for example, philosophical debates, contradictory position papers, and open-ended problem analyses that withhold conclusions — train an LLM on it, and test whether it reliably catches its own inconsistencies, revises assumptions on feedback, and withholds answers under uncertainty. If it does so robustly, the paper's claim that text-only training cannot produce analytical thinking would be falsified.","tokens_in":25556,"feed_emoji":"🤖","tokens_out":6768,"duration_ms":63546,"temperature":0.7,"pith_summary":"This paper argues that chatbots, built on large language models, are fundamentally System 1-style conversation partners: fluent, fast, and problem-to-solution oriented, but incapable of the reflective, analytical thinking that defines System 2 reasoning. It introduces a model of human thought called 'metaphorical problem propagation' — networks of problem positions and solution steps shaped by metaphor and prediction — and hypothesizes that LLM training reconstructs reduced, 'onward' forms of these from text. Because most training texts are solution-oriented, non-reflective prose, and because the resulting model lacks an embodied, experience-based conceptual system, the chatbot cannot assess or correct its own outputs the way a human thinking partner can. The conclusion is that further development of LLMs alone will not turn chatbots into analytical thinking partners; their role must be understood as guided cognitive assistance, not independent analysis.","feed_headline":"Even bigger LLMs won't make chatbots analytical thinkers","feed_subtitle":"A new analysis traces chatbot limits to the 'onward text' they train on, and argues for vigilant human guidance.","key_machinery":"The central object is the 'metaphorical problem propagation' — a model of the human thought space as a network of problem positions and solution steps, structured by metaphors and powered by predictive processing. It carries the argument by linking human cognition, human text, and LLM behavior: thought is modeled as propagation through this space; text is a reduced verbalization of it; LLM training is said to reconstruct artificial versions of it from text. The 'onward text' hypothesis is the load-bearing premise: most training text is assumed to be solution-oriented, coherent, non-experimental prose that moves from problem to conclusion without reflection, and this determines the System 1-l","core_discovery":"The central claim, stated in Section 10.4, is that 'a chatbot, as a conversational partner, is not an analytical thinking partner, nor can it become one with its current architecture and through text-only training.' The paper reaches this by combining four perspectives: a systems theory that views life as problem-solving, a metaphor-based account of concepts, a predictive view of perception and action, and a dual-process psychology distinguishing fast, automatic thinking from slow, analytical thinking. It introduces 'metaphorical problem propagation' as the mental space in which humans frame problems, propose solutions, and generate predictions. The paper then hypothesizes that most text use","pith_inferences":["A testable extension is to measure the 'onward-ness' of a training corpus and correlate it with an LLM's performance on analytical reasoning tasks; the paper's model predicts a strong negative correlation.","If the 'onward text' hypothesis is correct, injecting deliberately reflective, contradictory, and open-ended texts into pretraining should measurably shift chatbot behavior, even if — as the paper argues — it cannot fully overcome the structural limitation.","The paper's view implies that the gap between chatbot and human cognition is narrower in routine problem-solving domains (where humans also rely on System 1) and wider in open-ended analytical tasks, a differentiation the authors leave implicit.","The 'artificial metaphorical problem propagation' account suggests a possible evaluation metric: quantify the extent to which a chatbot's responses stay 'at the problem front' versus exploring alternative branches or revisiting assumptions."],"forward_implications":["Chatbots are best understood as System 1 conversation partners: they provide plausible, fast answers that confirm a user's worldview rather than critiquing it.","They can still serve as creative aids — for example, finding analogies or 'dark knowledge' — but only when guided by an alert prompt writer.","Scaling up models or adding more text of the same kind will not produce analytical thinking; the bottleneck is the lack of experiential grounding and the reduced nature of text.","Safety and reliability efforts should focus on barriers, monitoring, and user vigilance, not on the expectation that chatbots will become trustworthy reasoning partners.","The model gives interpretability researchers a target: look for concept clusters and activation paths that correspond to problem-solution 'riverbeds.'"],"fun_headline_variants":["LLMs can't think analytically—even with more data","Chatbots: conversational mimics, not analytical thinkers","Why bigger language models won't create thinking partners","The innovation illusion: chatbots lack human cognitive flexibility","Metaphor gap: why LLMs can't match human problem-solving"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that most text in LLM training datasets is 'onward text' — solution-oriented, non-reflective System 1-style prose — and that the training process encodes these as artificial metaphorical problem propagations; if training data were not dominated by such text, or if the training did not encode it in this way, the conclusion that chatbots are irredeemably System 1 would collapse.","fun_headline_variants_meta":{"raw":{"variants":["LLMs can't think analytically—even with more data","Chatbots: conversational mimics, not analytical thinkers","Why bigger language models won't create thinking partners","The innovation illusion: chatbots lack human cognitive flexibility","Metaphor gap: why LLMs can't match human problem-solving"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1238,"prompt_tokens":824,"completion_tokens":414,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":336}},"tokens_in":568,"tokens_out":414,"duration_ms":4441,"temperature":1.0,"reasoning_tokens":336,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T04:45:35.570604+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compile a training corpus of deliberately analytical, reflective texts — for example, philosophical debates, contradictory position papers, and open-ended problem analyses that withhold conclusions — train an LLM on it, and test whether it reliably catches its own inconsistencies, revises assumptions on feedback, and withholds answers under uncertainty. If it does so robustly, the paper's claim that text-only training cannot produce analytical thinking would be falsified.","supporting_citations":[],"review_version":4}