{"id":"1caa09e9-2c43-40c8-9dfd-f017dcd0ddd2","arxiv_id":"2501.10361","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM hallucination is reframed as excess extrapolation, with a historical genealogy from Wiener and Kolmogorov's time-series prediction to modern chatbots.","lead":"This paper argues that large language models should be seen as machines of extrapolation, predicting the next most likely token, and that hallucinations are not failures but signs the model is extrapolating too well. It also traces the concept of extrapolation back to Norbert Wiener's missile guidance work and Andrey Kolmogorov's parallel 1939 studies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim requires that fabricated citations lie outside the training set's convex hull; the paper never shows this, so 'hallucination = extrapolation' is unsupported.","rationale":"The reader's weakest assumption is precisely the hinge I identify: the conflation of next-token generation with extrapolation beyond the convex hull. The paper is a cultural-studies essay, not an empirical study, so the historical sections can stand on their own; but the central claim about LLM hallucination is an empirical claim about model behavior. §3 defines extrapolation via the convex hull, and §4 asserts that made-up references are extrapolations, but no membership test is supplied. The proposed adversarial test would settle whether typical hallucinations are outside the hull. Unless such a test is run (or the claim is softened to 'some hallucinations'), the paper's headline claim is conditional at best. I therefore agree with the reader's identification and move the verdict to CONDITIONAL rather than REJECT, because the historical and conceptual framing remains valuable and the empirical assumption is testable.","tokens_in":12917,"tokens_out":3797,"duration_ms":38727,"concrete_test":"Apply Yousefzadeh's convex-hull membership test (Yousefzadeh 2021; cf. Cao & Yousefzadeh 2023) to a sample of documented ChatGPT hallucinations, e.g., the fabricated references in Alkaissi & McFarlane 2023. For a target model, compute the convex hull of the training corpus in the model's embedding/feature space and test whether each hallucinated citation falls inside or outside it. If a majority fall inside the hull, or can be reconstructed as contiguous n-grams from training data, the paper's extrapolation claim is falsified for those cases.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim—'artificial hallucination has another name: extrapolation' (§1, p. 3)—requires that a hallucinated reference be a prediction beyond the training set's convex hull. In §3 the author defines extrapolation as output outside 'the convex hull' of the training set and interpolation as inside it. Yet no evidence is given that fabricated citations actually lie outside that hull. A made-up reference is typically a recombination of real author names, journal titles, and DOIs that appear in the training data; at the token and embedding level such a string may lie inside the convex hull, i.e., it may be an interpolation between memorized exemplars. Hallucination can also arise from decoding errors, sampling randomness, or attention failures, none of which is extrapolation. The appeal to Wiener's stationary-time-series theory (§1) is metaphorical: LLM token distributions are not stationary and the transformer is not a linear predictor. Thus the dichotomy on which the reframing rests—hallucination = extrapolation vs. malfunction—is unsupported. If hallucinations are often memorized or interpolated, calling them extrapolation is a category error.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that large language models are best understood as machines of extrapolation, and that the phenomenon usually called \"hallucination\" is in fact an instance of the model extrapolating beyond its training data. The argument proceeds in two strands. The first is historical: the paper traces the concept of extrapolation through Norbert Wiener's wartime work on time series and missile prediction, connects it to Andrey Kolmogorov's parallel work, and reads the subsequent reprinting of Wiener's monograph as a Cold War public-relations maneuver. The second strand is conceptual: using the convex hull of a training set as a geometric criterion, the paper claims that outputs outside that hull are extrapolations, and that fabricated references produced by LLMs are such outputs. The paper concludes that hallucination is not a malfunction but evidence that the model is working as designed, potentially too well.","tokens_in":13139,"tokens_out":2642,"duration_ms":89855,"significance":"If the central claim were established, the paper would provide a provocative and genuinely useful reframing of AI hallucination, with consequences for how researchers, users, and regulators describe and evaluate LLM failures. The historical research on Wiener and Kolmogorov is a genuine contribution: it draws attention to underappreciated connections between wartime prediction theory and modern machine learning, and it does so with careful attention to archival detail and to the sociopolitical contexts of both scientists. The paper also productively connects contemporary LLM discussions to a longer intellectual history, which is valuable for interdisciplinary AI scholarship. However, the load-bearing conceptual claim—that hallucinated outputs are extrapolations—is not supported by any empirical or formal analysis. The definition of extrapolation via convex hulls is introduced but never operationalized on LLM outputs, and the paper does not engage with alternative mechanisms for hallucination such as memorization, interpolation, decoding errors, or sampling randomness.","major_comments":[{"comment":"The paper's central identification of hallucination with extrapolation is unsupported. In §3, extrapolation is defined as an output lying outside the convex hull of the training set, and interpolation as lying inside it. In §4, fabricated references are offered as examples of extrapolation. Yet the paper provides no evidence that a hallucinated citation actually lies outside the convex hull of the training data. At the token, embedding, or semantic level, a fabricated reference is typically a recombination of real author names, journal titles, and DOIs, which may well lie inside the convex hull of memorized exemplars. The definition in §3 is therefore never connected to the phenomenon in §4; the claim \"artificial hallucination has another name: extrapolation\" (p. 3) remains an assertion rather than a demonstrated equivalence.","section":"§3 and §4"},{"comment":"The analogy to Wiener's stationary time series extrapolation is metaphorical rather than substantive. Section 1 describes token generation as a time series and cites Wiener's theory, while §3 switches to a convex-hull definition that has no direct relationship to Wiener's linear stationary prediction theory. LLM token distributions are not stationary, transformers are not linear predictors, and the mean-square error framework of Wiener's work does not apply to next-token sampling in modern LLMs. The paper does not bridge this gap, so the historical appeal cannot bear the weight of the conceptual conclusion.","section":"§1 and §3"},{"comment":"The claim that a hallucinated output is \"nothing more unreal than what was statistically probable in a series\" (p. 2) dismisses without argument a substantial body of evidence that hallucination can arise from mechanisms other than extrapolation, such as memorization of training data, sampling randomness, decoding errors, or attention failures. Even if some hallucinations are extrapolative, the paper needs to argue that these are the typical or defining cases; otherwise, the reframing is a category error. The paper offers no such argument, and its dichotomy between \"malfunction\" and \"extrapolation\" is too coarse to capture the multi-causal nature of hallucination.","section":"§1 and §4"}],"minor_comments":[{"comment":"There are numerous typographical and grammatical errors, e.g., \"There could had been many paths\" (p. 2), \"recreats\" (p. 8), \"Guatarri\" (p. 10), \"Norvi\" (works cited entry for Russell and Norvig), and \"Marconian\" (p. 23). These should be corrected in a revision.","section":"Throughout"},{"comment":"The footnote states that \"interpolation and extrapolation are the same statistical method,\" but the main text in §3 clearly distinguishes them. This apparent contradiction should be resolved.","section":"Footnote 2"},{"comment":"The paper cites Wiener's \"A Scientist Rebels\" as published in 1947 in the text and as 1946 in the works cited list; the access date and issue details should be reconciled.","section":"§2, historical citations"},{"comment":"The claim that the 1949 reprinting of Wiener's monograph was a deliberate public-relations pivot away from the atomic bomb is presented as interpretation, but the evidence offered (the reprint itself, the preface) does not clearly support this causal claim; the claim should be labeled as speculative or supported with additional archival evidence.","section":"§2"},{"comment":"The discussion of Balestriero et al. (2021) and Cao and Yousefzadeh (2023) would benefit from a clearer explanation of the difference between \"always amounts to extrapolation\" and the author's more restricted claim, especially since the latter is the author's own prior work and is not summarized in enough detail for a reader to evaluate the distinction.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"The paper is clearly positioned in cultural studies / critical AI rather than in a formal ML or statistics venue. The historical contribution is real, but the central conceptual claim needs substantially more support. I would suggest that the editor consider whether the paper might be reframed as a historical and conceptual proposal—\"hallucinations as usefully seen through the lens of extrapolation\"—rather than as an established equivalence. The heavy reliance on the author's own prior work (Cao and Yousefzadeh 2023) for the key interpolation/extrapolation distinction is also worth flagging; the argument would be strengthened by summarizing the method of that work or by citing independent sources. As it stands, the paper overreaches in its central claim while being underdeveloped in the technical bridge between the convex-hull definition and actual LLM outputs."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a cultural-studies essay, not a technical paper. The historical reconstruction is genuinely new and worth reading; the central claim—'artificial hallucination has another name: extrapolation'—is a suggestive metaphor that the paper never actually supports.\n\nThe paper's real contribution is the genealogy. It traces extrapolation from Wiener's 1941 missile-prediction work to his communication engineering, reconstructs the parallel Kolmogorov 1939 paper, and shows how the concept became a bridge between wartime science and cybernetics. That synthesis is not present in the literature I know, and the author handles the Wiener/Shannon rivalry and Wiener's post-Hiroshima moral turn with nuance. For a humanities audience, this is solid, careful historical writing.\n\nWhere it gets soft: the conceptual argument. In §3 the author defines extrapolation via the convex hull of a training set, and asserts that a hallucinated citation is extrapolation because it leaps 'beyond the boundaries of its training data.' But the paper offers no evidence that fabricated references actually lie outside the convex hull. In fact, a made-up citation is typically a recombination of real author names, journal titles, and DOIs that appear in the training data; at the embedding level, such a string might well be inside the hull. The author never grapples with that possibility. Neither does the paper engage with the stationarity assumptions in Wiener's theory; LLM token distributions are not stationary time series, so the 'extrapolation' is at best an analogy. The author also leans on her own prior work with Yousefzadeh for the key distinction, which is fine but not a substitute for support here.\n\nNone of this sinks the historical half. The essay clearly knows its genre: the claim is an interpretive reframing, not a theorem. But the paper presents the convex hull definition as if it does the empirical work, and it doesn't. That gap is the main thing a referee would want fixed.\n\nWho this is for: people interested in the cultural history of AI and the rhetorical framing of hallucination. It would be a reasonable contribution to Big Data & Society or an STS venue. I'd send it to peer review, not desk reject it—but I'd expect the reviewers to push hard on the convex hull point.\n\nRecommendation: worth engaging; the history is valuable, and the conceptual claim is worth arguing with.","headline":"A genuine historical genealogy of extrapolation, with a central conceptual claim that is more slogan than demonstrated.","tokens_in":13609,"tokens_out":2633,"would_cite":true,"duration_ms":44307,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models are machines of extrapolation; hallucinated references are statistically probable continuations beyond the training data.","keywords":["large language models","extrapolation","hallucination","cybernetics","Norbert Wiener","Andrey Kolmogorov","time series","convex hull"],"falsifier":"Take a known fabricated citation and locate its representation relative to the model's training set: if the invented reference falls inside the envelope of training examples, or is closer to memorized verbatim passages than to the boundary, then hallucination is interpolation rather than extrapolation and the paper's central reframing fails.","tokens_in":12723,"feed_emoji":"🤖","tokens_out":9362,"duration_ms":91398,"temperature":0.7,"pith_summary":"This paper argues that large language models are best understood as machines of extrapolation: they predict the next value in a series by leaping beyond the data they were trained on. On this view, the fabricated references that get called 'hallucinations' are not malfunctions but evidence that the model is working according to its design, producing statistically probable continuations that may not correspond to anything in the real world. The paper supports this conceptual claim with a historical one, tracing extrapolation back to Norbert Wiener's wartime work on time-series prediction and Andrey Kolmogorov's nearly simultaneous 1939 paper, and connecting Kolmogorov's later compression ideas to the deep-learning lineage behind modern chatbots. If the argument is right, the urgent question is not how to stop models from hallucinating but how to calibrate and disclose the extrapolation they are built to perform.","feed_headline":"LLM hallucinations are extrapolation, not failure","feed_subtitle":"A cultural-history reading says made-up citations are the model working as designed, leaping beyond its training data.","key_machinery":"The object that carries the argument is extrapolation, defined as the statistical operation of predicting the next value in a series, joined with the convex-hull picture of machine learning: a training set forms a minimal enclosure, and any output outside that enclosure is extrapolation while output inside it is interpolation. Wiener's extrapolation of stationary time series supplies the mathematical template—a message, like a missile trajectory, is a sequence of values whose future can be predicted within error bounds—and the paper transplants that template onto transformer-based next-token generation. The convex hull gives the argument its geometry: GPT's fake references are outputs that reach beyond the hull of factual training data while remaining statistically likely continuations of the prompt.","core_discovery":"The central discovery the paper is trying to establish is that artificial hallucination has another name: extrapolation. Each token an LLM emits is, in the paper's account, the next value in a time series; the model has learned statistical correlations between elements and projects a continuation beyond its training set. Under this description, a made-up citation is a perfectly sensible sequence—a plausible title, a likely author, a journal that would fit, a pseudo-DOI—rather than an error in the machine. The paper therefore claims that 'hallucination' pathologizes a functional behavior: the chatbot is not broken but extrapolating, and often extrapolating too well for the factual task it was asked to do.","pith_inferences":["If the reframing holds, the practical safety question becomes 'how far beyond its training hull is a given output, and does the model signal that it has left the hull?' rather than 'how do we stop it from lying?'","The convex-hull criterion suggests a quantitative test of the paper's claim: hallucinated references should sit near the boundary of the training-data distribution, whereas memorized or interpolated outputs should sit deep inside it.","The copyright-evasion point implies extrapolation is also an economic design feature, not just an epistemic one, so transparency proposals will have to contend with the incentive to keep provenance deniable.","One extension of the historical story would be to compare hallucination rates of interpolation-heavy systems, such as retrieval-augmented models, against pure next-token generators; a lower rate would support the claim that hallucination tracks extrapolation capacity."],"forward_implications":["Hallucinated references would be reclassified as design behavior rather than malfunction, so suppressing them means constraining a core capability, not repairing a defect.","Extrapolation accounts for the same capacity that produces poems and novel prose and that lets chatbots avoid reproducing copyrighted text, giving the model a kind of intentional opacity.","The interpolation/extrapolation distinction offers a geometric language for transparency: models could in principle report when an output lies outside the training hull.","Evaluation would shift from asking whether every output matches a fact in the training set to asking how well the model extrapolates within acceptable error bounds.","The historical lineage places LLM-era debates about trust in machines on a continuum with mid-century cybernetics, where predicting missiles and predicting messages were already the same mathematical problem."],"supporting_citations":[{"why":"Introduces the term 'artificial hallucination' that the paper reinterprets as extrapolation.","marker":"Alkaissi and McFarlane 2023"},{"why":"Provides the extrapolation concept and the equation of messages with time series that anchors the argument.","marker":"Wiener 1942/1964"},{"why":"Supplies the parallel extrapolation result that supports the historical claim of independent simultaneous discovery.","marker":"Kolmogorov 1939"},{"why":"Recent machine-learning claim that high-dimensional learning always amounts to extrapolation, which the paper qualifies.","marker":"Balestriero et al. 2021"},{"why":"Earlier work connecting extrapolation to convex hulls and AI transparency, used to define extrapolation geometrically.","marker":"Cao and Yousefzadeh 2023"},{"why":"Classic textbook definition of machine learning as detecting and extrapolating patterns, invoked to show extrapolation is constitutive.","marker":"Russell and Norvig 2010"},{"why":"Supplies the 'iconic interpretation' concept used to diagnose anthropocentric expectations behind the hallucination label.","marker":"Weatherby and Justie 2022"}],"fun_headline_variants":["LLM hallucinations are extrapolation, not bugs","Extrapolation explains LLM hallucinations","From guided missiles to guided prompts: LLM extrapolation","Why LLMs make up facts? It's extrapolation, not error","Hallucinations are why LLMs work: extrapolation explained"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a chatbot's next-token generation really is a leap beyond its training data in the same sense that Wiener extrapolated a stationary time series, rather than a recombination of memorized or interpolated fragments.","fun_headline_variants_meta":{"raw":{"variants":["LLM hallucinations are extrapolation, not bugs","Extrapolation explains LLM hallucinations","From guided missiles to guided prompts: LLM extrapolation","Why LLMs make up facts? It's extrapolation, not error","Hallucinations are why LLMs work: extrapolation explained"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000771,"raw_usage":{"total_tokens":3364,"prompt_tokens":848,"completion_tokens":2516,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":2437}},"tokens_in":464,"tokens_out":2516,"duration_ms":17122,"temperature":1.0,"reasoning_tokens":2437,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:23:25.224757+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a known fabricated citation and locate its representation relative to the model's training set: if the invented reference falls inside the envelope of training examples, or is closer to memorized verbatim passages than to the boundary, then hallucination is interpolation rather than extrapolation and the paper's central reframing fails.","supporting_citations":[],"review_version":1}