{"id":"444946f4-6f2b-4ff5-8422-d2f3df22404f","arxiv_id":"2502.08663","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"The authors report that Minkowski distances between keyword embeddings differ significantly between Llama2 and Llama3 responses, and use this to classify hallucinations with up to 66% accuracy, though the setup conflates model identity with hallucination.","lead":"This paper claims that hallucinated LLM answers have measurably different distances in the embedding space than correct answers, and uses that difference to build a detector. The proposed detector reaches 66% accuracy, but only on a synthetic dataset where one model always answers wrongly and another always answers correctly.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim is unsupported because hallucination status is perfectly confounded with model identity: §3.2 labels Llama2 responses \"hallucinated\" and Llama3 responses \"non-hallucinated\" purely by training cutoff, so observed embedding-distance differences could be model style, not…","rationale":"The reader's weakest-assumption analysis identifies precisely the load-bearing flaw: the experimental design in Section 3.2 conflates model identity with hallucination status. I agree that this is fatal to the central claim as stated. The paper's own limitations section concedes that some Llama3 responses may be mislabeled, but the deeper issue is not a small label-noise fraction—it is that the labels are definitionally tied to the generator. Llama2 and Llama3 differ in training data, architecture, alignment, and decoding behavior; even when both outputs are mapped through the same BERT embedder, the lexical and stylistic properties of the generated text can differ systematically. The claimed \"scale-free\" robustness across norms, keyword counts, and response counts does not help, because all those variations are measured under the same confounded labeling. The test phase compounds the problem: training and test sets are drawn from the same 64 questions and the same two models, so the detector can succeed by recognizing model provenance rather than hallucination. A credible test of the structural-difference hypothesis requires a same-model control, such as eliciting both correct and hallucinated responses from one model and checking whether the distance distributions still separate. The appendix examples also support the concern: the response labeled non-hallucinated contains questionable factual material, and the responses labeled hallucinated look qualitatively different from typical fluent hallucinations, suggesting the comparison may be capturing coherence or prompt-adherence rather than factuality. I see no machine-checked proofs, released code, or external benchmark evaluations that would independently support the claim. Therefore the reader's REJECT verdict should stand; no adjustment is needed.","tokens_in":21409,"tokens_out":3612,"duration_ms":34265,"concrete_test":"Run the exact same pipeline (KeyBERT keyword extraction, BERT embeddings, Minkowski distances, KDE classification) with hallucination labels generated by a single model: use Llama3 with fixed sampling parameters to answer 64 questions, half with correct factual premises and half with fabricated premises designed to elicit hallucinated answers, and compare within-class distance distributions with the identical Wilcoxon, KL, and accuracy analyses. If separation persists when the generator is held fixed, the structural claim survives; if it collapses, the reported differences are attributable to model identity rather than hallucination.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that hallucinated content has statistically significant, scale-free structural differences in embedding space—rests entirely on dataset labels assigned in Section 3.2. There, responses are labeled solely by generator: Llama2 (training cutoff September 2022) is asserted to hallucinate answers about events between September 2022 and September 2023, while Llama3 is asserted to answer correctly. The authors write: \"This choice ensures that the answers generated by Llama2 hallucinated, while those generated by Llama3 did not.\" Because model identity and hallucination status are perfectly confounded, every reported difference—Wilcoxon p<0.01, KL divergence, median distance gaps, and the 66% test accuracy—could instead reflect stylistic, lexical, or decoding differences between Llama2 and Llama3 outputs after BERT embedding. The detection experiment does not break the confound: train and test responses come from the same 64 questions and the same two models, so the KDE-based classifier can exploit model identity rather than hallucination. The appendix further undermines the labels: the \"non-hallucinated\" response NH-1 contains dubious or false claims, while the \"hallucinated\" responses H-1 and H-2 are an unrelated assignment-style prompt and a quote collage, not fluent factually wrong answers. The limitations section acknowledges the label risk, but a \"little subset\" caveat does not remove the confound. A same-model control is required before the structural-difference claim can be accepted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a probabilistic framework for detecting LLM hallucinations by analyzing Minkowski distances between BERT embeddings of keywords extracted from model responses. The authors construct a synthetic dataset using Llama2 and Llama3, labeling Llama2 responses to questions about the September 2022–September 2023 period as hallucinated and Llama3 responses as non-hallucinated, purely based on training cutoff dates. They then compute intra-class distance distributions, report statistically significant differences via Wilcoxon tests and KL divergence, claim these differences are scale-free across distance norms, keyword counts, and response counts, and build a KDE-based classifier that achieves 66% accuracy for one configuration. The paper's central claim is that this work is the first to show structural differences between hallucinated and correct LLM responses.","tokens_in":21667,"tokens_out":5118,"duration_ms":43601,"significance":"If the central claim were valid, a purely geometric, external-knowledge-free signal for hallucination would be practically valuable and could support a simple and cheap detection method. The paper also provides a concrete, reproducible pipeline (questions, models, embeddings, and statistical tests) that could serve as a baseline for embedding-distance-based hallucination detection. However, the significance of the claimed results is entirely undermined by the experimental design: the ground-truth labels are perfectly confounded with model identity, so every reported distributional difference and the classifier's accuracy may reflect model-specific style, vocabulary, or decoding behavior rather than hallucination itself. Because the appendix examples further contradict the label definitions, the paper does not establish the existence of a structural signal specific to hallucinated content.","major_comments":[{"comment":"The dataset labels are assigned solely by generator identity. The authors write, \"This choice ensures that the answers generated by Llama2 hallucinated, while those generated by Llama3 did not,\" but this equates model identity with hallucination status: Llama2 responses are the 'hallucinated' class and Llama3 responses are the 'non-hallucinated' class for every question. Consequently, every distance-distribution difference reported in §4.1 and used in §3.4 is confounded with model-specific properties (e.g., token distributions, response length, decoding style) after BERT embedding. The caveat in §5.2 that \"a little subset of Llama3 responses might be hallucinated as well\" weakens the labels further but does nothing to remove the perfect correlation between class and model; no same-model comparison is offered.","section":"§3.2, Dataset"},{"comment":"The detection experiment shares both the same 64 questions and the same two generating models between training and test. The KDE likelihood scoring therefore has access to a training set in which the two classes are exactly the two models, so a classifier can separate the test responses by reproducing the Llama2/Llama3 embedding geometry rather than by detecting hallucination. The 66% accuracy in Table 3 is thus not evidence for a structural signal specific to hallucinated content. A control experiment using a single model with verified factual errors versus correct responses from that same model would be required to support the paper's central claim.","section":"§3.4 and §4.2, Test"},{"comment":"The qualitative examples contradict the label definitions. NH-1 asserts, among other things, that COVID-19 caused economic decline and that a phone call between President Trump and Chinese Vice Premier Liu He eased tariffs, claims that are false for the September 2022–September 2023 period; NH-2 is a nonsensical enumeration of countries. Conversely, H-1 is an unrelated assignment-style prompt and H-2 is a collage of motivational quotations, neither of which is a fluent, factually wrong response to the question about the global economy. These examples show that the automatic labels do not track the intended construct of hallucination versus correct content.","section":"Appendix A.2, Question vs. Responses"},{"comment":"The baseline accuracies are copied from Du et al. (2024), where they were reported on TruthfulQA, while the proposed method is evaluated on the authors' synthetic Llama2/Llama3 dataset. The two evaluation settings differ in topic distribution, label source, and response length, so the table does not support the statements that the method is \"second\" or \"comparable with the best results in the field.\" The baselines need to be run on the same evaluation protocol and the same test split before any comparison is meaningful.","section":"§4.2, Table 4, Comparison with state of the art"}],"minor_comments":[{"comment":"The Wilcoxon tests are performed on hundreds of thousands of paired distances per configuration, so p<0.01 carries little information; report an effect size (e.g., rank-biserial correlation) and a confidence interval.","section":"§4.1"},{"comment":"The \"scale-free\" property is inferred from visual similarity across boxplots; state a quantitative invariance criterion (e.g., consistent sign and approximate proportionality of median differences across conditions) and test it explicitly.","section":"§4.1 and Appendix A.4"},{"comment":"Clarify whether the 4-bit quantization applies to model weights or to the generation process, and report the generation temperature, top-p, top-k, and maximum token settings; reproducibility requires these details.","section":"§3.2"},{"comment":"The score comparison \"Shall > Snohall\" is equivalent to comparing average log-likelihoods only because the two sums have equal length qr; state this equivalence explicitly to avoid confusion.","section":"§3.4"},{"comment":"Typos and wording issues include: Table 1 caption \"Minkoswki\" should be \"Minkowski\"; §6 uses \"hallucinationed\" twice; reference Skala (2013) title contains \"Oexpected\"; and §4.1's sentence about the 279% KL increase lacks the baseline it increases with respect to.","section":"Typos and wording"}],"recommendation":"reject","confidential_remarks":"The paper's central claim is entirely undermined by the model-confounded labeling, and the qualitative examples in the appendix actively contradict the labels. A same-model control with verified factual errors versus correct responses, plus a fair re-run of baselines on the same data, would require a substantial rework rather than a local revision. I also note that the comparison in Table 4 appears to be reproduced from the HaloScope paper without running those baselines on the authors' dataset; the fit to this venue would be improved by addressing these issues in a future submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper is mechanically competent: the Minkowski distance computations, KDE fitting, Wilcoxon tests, and KL divergence reporting all look correct, and the writing is clear. Second, the central claim—that hallucinated responses have statistically significant, scale-free structural differences in embedding space—is not supported by the experimental design. Section 3.2 labels Llama2 responses as hallucinated and Llama3 responses as non-hallucinated purely by training cutoff date, so model identity and hallucination status are perfectly confounded. Every reported difference could just be Llama2's style, vocabulary, or decoding after BERT embedding. This is not a minor caveat; it is the load-bearing premise.\n\nWhat is genuinely new: the specific combination of Minkowski distances between BERT keyword embeddings, KDE likelihood scoring, and the 'structural differences' framing has not appeared before in hallucination detection. The authors also deserve credit for acknowledging the label risk in Section 5.2, though they don't resolve it. And the 'scale-free' checks across r, n, and p are a reasonable robustness exercise.\n\nThe soft spots are serious. The confound alone would sink the claim, but there is more. The appendix examples undermine the labels: the 'non-hallucinated' response NH-1 contains dubious or false claims, while the 'hallucinated' responses H-1 and H-2 are an unrelated assignment-style prompt and a quote collage, not fluent factually wrong answers. The test set comes from the same 64 questions and same two models used to build the KDE densities, so the classifier can exploit model identity. The best (r=8, n=1, p=0.5) configuration is selected post hoc from 210 configurations, and the baseline comparison in Table 4 uses mismatched datasets, cherry-picking numbers from HaloScope's evaluation. No code or data are released.\n\nThe 66% accuracy is modest and, given the confound, not evidence for the structural claim. A proper test would need the same model generating both correct and hallucinated responses—for instance, questions before and after the training cutoff, or a labeled factual/incorrect dataset like TruthfulQA.\n\nWho should read this: anyone designing hallucination detection experiments, as a cautionary example of how easy it is to confuse model identity with hallucination status. It does not deserve to be cited as evidence for structural differences. It does deserve a serious referee, because the idea is worth testing properly and a reviewer can force the required control. My recommendation: engage with it, but only with the expectation of major revision.","headline":"Mechanically careful but confounded: the central structural-difference claim reduces to 'Llama2 and Llama3 differ in embedding distances,' so the paper needs a same-model control before it can be taken seriously.","tokens_in":22253,"tokens_out":2477,"would_cite":false,"duration_ms":21194,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T50","62H30","62G07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that hallucinated and genuine LLM responses are structurally distinct in embedding space, and that a probabilistic distance-based classifier can flag hallucinations at up to 66% accuracy without external knowledge.","keywords":["LLM hallucination","embedding distance","Minkowski distance","hallucination detection","kernel density estimation","BERT embeddings","structural difference"],"falsifier":"Run the identical distance analysis within a single model on verifiable correct and incorrect answers from the same generator; if the two distance distributions do not separate under the same Wilcoxon comparison, the reported structural difference is a property of the two generators, not of hallucination itself.","tokens_in":21151,"feed_emoji":"📐","tokens_out":12198,"duration_ms":88585,"temperature":0.7,"pith_summary":"This paper claims that hallucinated and genuine LLM responses are structurally distinct in embedding space, and that the distinction can power a detector with no external knowledge or fact-checking. The authors generate 64 questions about events between September 2022 and September 2023, collect multiple answers from Llama2 (training cutoff September 2022) and Llama3 (cutoff December 2023), and compare the distributions of pairwise Minkowski distances among the BERT embeddings of each response's extracted keywords. They report a statistically significant difference between hallucinated and non-hallucinated distance distributions (Wilcoxon test, $p < 0.01$) that persists across every tested norm ($p \\in \\{0.5, 1, 2\\}$), keyword count ($n \\in \\{1, \\ldots, 10\\}$), and response count ($r \\in \\{4, \\ldots, 16\\}$). A probabilistic classifier built on these distributions reaches 66% accuracy in its best configuration. If the structural claim holds, hallucination detection becomes a geometry problem rather than a fact-checking problem.","feed_headline":"Embedding distances separate hallucinated from correct LLM answers","feed_subtitle":"Pairwise distance statistics flag hallucinated answers at 66% accuracy, no fact-checking needed.","key_machinery":"The load-bearing object is the Minkowski distance between BERT embeddings of response keywords. Each response is reduced by KeyBERT to its $n$ most important keywords ($n = 1, \\ldots, 10$), those keywords are embedded into 768-dimensional BERT vectors, and all pairwise distances within each class are computed under three Minkowski norms ($p \\in \\{0.5, 1, 2\\}$). Those pairwise distances become the unit of analysis: KL divergence and median difference quantify the gap between hallucinated and non-hallucinated distance distributions, the Wilcoxon test establishes that the gap is statistically significant, and Gaussian kernel density estimation converts the training distributions into likelihood models. At test time, each new response's distances to all training responses of each class are scored under the two KDE log-likelihoods, and the class with the higher summed log-likelihood wins. The same machinery yields both the claimed scale-free structural result and the detector.","core_discovery":"On the paper's own terms, the central discovery is that hallucinated content carries a measurable structural signature: the pairwise embedding distances of hallucinated responses and those of correct responses come from different distributions, and the difference is large enough to be statistically detectable and stable across parameters. The claim is grounded in an artificial dataset in which Llama2 necessarily hallucinates answers about a year (September 2022 to September 2023) beyond its training cutoff, while Llama3, whose training covers that year, produces correct answers. Distances are computed between BERT embeddings of the $n$ most important keywords extracted from each response, using Minkowski norms $p = 0.5$, $p = 1$, and $p = 2$. The authors report that the distributional gap, measured by KL divergence and median difference, widens as the number of keywords and responses grows, remains significant under a Wilcoxon test for every configuration, and is scale-free in the sense that the qualitative separation does not depend on the norm, the keyword count, or the number of responses. The paper presents this as the first demonstration that hallucination has a geometric structure in embedding space.","pith_inferences":["If the structural separation is genuine and not just a between-model artifact, it suggests hallucinated text is produced by a different statistical regime within a model, not merely text that happens to be false; that would make embedding geometry a useful probe for interpretability work.","A decisive extension is to repeat the exact distance protocol within a single model, comparing factually correct answers with factually wrong answers from the same generator; this would separate the hallucination signal from model style.","The 66% ceiling with a simple KDE rule suggests headroom: combining the distance score with semantic entropy or self-evaluation could yield a stronger joint detector than any single signal.","Testing the same claims with other embedding models (not only BERT) and other model families would show whether the scale-free structural difference transfers or is tied to these specific encoders."],"forward_implications":["A hallucination detector can operate on response geometry alone, needing no external knowledge base, retriever, or fact-checker.","Because the separation is claimed for every tested norm, keyword count, and response count, a deployment can trade cost for accuracy without expecting the signal to disappear.","Fractional distances ($p = 0.5$) amplify the distribution gap and produced the best accuracy (0.66 at $r=8$, $n=1$), pointing to high-dimensional distance behavior as a tuning lever.","The accuracy comparison against published detectors places the geometry-only approach above token-probability and entropy baselines but below HaloScope on the reported TruthfulQA numbers.","The cutoff-gap data-generation recipe, asking about the interval covered by one model's training but not another's, scales to any pair of models and can produce labeled hallucination data without manual annotation."],"supporting_citations":[{"why":"Supplies the BERT embedding that maps each response's keywords to the 768-dimensional vectors between which distances are computed.","marker":"(Kenton & Toutanova, 2019)"},{"why":"Supplies KeyBERT, the keyword extractor that converts each response into the 1-10 keywords whose embeddings are compared.","marker":"(Grootendorst, 2020)"},{"why":"Motivates the use of fractional Minkowski norms (p=0.5) to handle sparsity in high-dimensional spaces.","marker":"(Aggarwal et al., 2001)"},{"why":"Provides the robust kernel density estimation used to turn training distance distributions into likelihood models for scoring test responses.","marker":"(Kim & Scott, 2012)"},{"why":"The comparison source for published detector accuracies on TruthfulQA and the strongest baseline (HaloScope) against which the 66% result is positioned.","marker":"(Du et al., 2024)"},{"why":"The closest prior work also operating on embeddings, from which the paper distinguishes its claimed first structural-difference analysis.","marker":"(Chen et al., 2024)"},{"why":"Supplies TruthfulQA, the closest public dataset used to contextualize the accuracy comparison.","marker":"(Lin et al., 2022b)"}],"fun_headline_variants":["Embedding distances expose hallucination structure in LLMs","Hallucinations show distinct embedding distance patterns","Geometric signature flags LLM hallucinations at 66% accuracy","First proof: hallucinated text has a geometric embedding signature","Embedding distance stats separate true vs fabricated LLM answers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The labels that split responses into hallucinated and genuine come entirely from which model generated them: Llama2 is assumed to hallucinate every answer about the post-cutoff year, and Llama3 is assumed to answer every such question correctly.","fun_headline_variants_meta":{"raw":{"variants":["Embedding distances expose hallucination structure in LLMs","Hallucinations show distinct embedding distance patterns","Geometric signature flags LLM hallucinations at 66% accuracy","First proof: hallucinated text has a geometric embedding signature","Embedding distance stats separate true vs fabricated LLM answers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000494,"raw_usage":{"total_tokens":2433,"prompt_tokens":963,"completion_tokens":1470,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":1392}},"tokens_in":579,"tokens_out":1470,"duration_ms":10791,"temperature":1.0,"reasoning_tokens":1392,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T15:51:41.493557+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical distance analysis within a single model on verifiable correct and incorrect answers from the same generator; if the two distance distributions do not separate under the same Wilcoxon comparison, the reported structural difference is a property of the two generators, not of hallucination itself.","supporting_citations":[],"review_version":1}