Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper claims that hallucinated LLM outputs can be caught in a single pass by measuring when a model's internal representation of its input and of its generated answer drift apart, and introduces HIDE, a training-free score that does…

desk verdict Solid empirical method for single-pass hallucination detection, but the HSIC decoupling story rests on a pairing assumption that doesn't hold; the AUC results stand anyway. read the letter →

arxiv 2506.17748 v1 pith:HD2JQ2UV submitted 2025-06-21 cs.CL cs.AI

classification cs.CLcs.AI
keywords hallucinationdetectionlanguagemodelshiddenstatesHSICstatisticaldependencesingle-passfaithfulnessfactuality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to show that a language model's hallucinated answers can be identified in a single generation by measuring how much its internal representation of the output has come loose from its representation of the input. It introduces HIDE, which computes a kernel-based statistical-dependence score, HSIC, between hidden-state vectors of salient prompt tokens and salient generated tokens; a low score signals that the output is decoupled from the context and therefore likely hallucinated. Across four question-answering benchmarks and six open-source language models, HIDE beats other single-pass detectors by roughly 29 percent average relative AUC-ROC and matches or slightly exceeds multi-pass methods while using about 51 percent less computation time. The practical appeal is that hallucination detection becomes cheap enough for real-time use without extra training or repeated sampling.

What carries the argument

The load-bearing object is HSIC, the Hilbert-Schmidt Independence Criterion: a kernel-based measure that is zero exactly when two random variables are independent, estimated here from the hidden states of about twenty semantically salient input tokens and twenty salient output tokens taken from a middle decoder layer. The paper rescales the standard unbiased estimator's denominators to produce a bounded, numerically stable score $\widehat{\text{HSIC}}_{\text{HIDE}}$ that remains well-defined even for one or two tokens, is asymptotically unbiased, and separates faithful from hallucinated outputs near a threshold of about 0.12. The RBF kernel supplies the characteristic-kernel guarantee that lets a near-zero score be interpreted as genuine independence.

What would settle it

Shuffle the order of the selected output tokens relative to the selected input tokens and recompute the HIDE score on the same examples; if the score still separates hallucinated from faithful outputs, then it is not measuring input-output dependence between paired tokens, and the decoupling mechanism is not what carries the result.

Watch

Extended reading notes

Core claim

The central claim is that hallucination is visible as a measurable statistical decoupling between the model's internal representations of its input and the output it generates. Faithful answers keep the hidden states of selected input tokens and selected output tokens statistically dependent, while hallucinated answers let that dependence drop toward zero. HIDE quantifies this with an adapted HSIC estimator built from RBF-kernel Gram matrices of the selected hidden states, and flags an output as hallucinated when the resulting score falls below a threshold. The evidence offered is that HIDE outperforms the best single-pass baseline by an average relative improvement of about 29 percent in AUC-ROC across four datasets and six models, is competitive with multi-pass methods at about 3 percent relative improvement, and costs roughly 51 percent less computation time.

Load-bearing premise

The method stands on the assumption that the tokens picked from the prompt and the tokens picked from the answer can be treated as paired samples of the same underlying relationship, but the selection procedure chooses the two sets independently and pairs them only by rank order.

Editorial extensions

If this is right

  • If HIDE works as claimed, hallucination detection no longer requires multiple generations per input, removing the main latency barrier to real-time detection.
  • With roughly 29 percent average relative AUC-ROC gain over the best single-pass baseline, a HIDE-style score could replace or augment perplexity- and energy-based checks in production question-answering systems.
  • Because performance stays nearly flat across decoder layers and across most kernel choices, deployment needs little per-model tuning; a token budget near 15 to 20 and a threshold near 0.12 act as near-optimal defaults.
  • The method's known failure modes are short single-token answers, which produce an uninformative zero score, and outputs that copy the prompt verbatim, which inflate the score and can mask a hallucination.
  • The method is white-box and training-free, so any model with accessible hidden states can use it without additional data or fine-tuning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the paper never tests the pairing assumption directly; the selected input and output tokens are chosen independently and matched only by rank order, so the score's empirical power may partly come from confounds such as lexical overlap or output length rather than from statistical dependence between paired representations.
  • A testable next step is to shuffle the order of the selected output tokens before computing HSIC; if the score still separates hallucinated from faithful outputs, then the decoupling story is not what carries the result.
  • The stability of the optimal threshold around 0.12 across datasets suggests a calibration-free deployment, but those thresholds were fit using ground-truth labels, so a fair test is to fix the threshold on one benchmark and measure transfer to unseen domains.
  • The decoupling hypothesis predicts a causal signature: interventions that re-ground generation, such as context-aware decoding, should raise the HIDE score on the same prompts; measuring that would connect the detector to the mechanism it claims to exploit.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. Chatterjee, Goel, and Chakraborty propose HIDE, a single-pass and training-free hallucination detector. For an input prompt and one LM-generated output, HIDE selects a budget of salient tokens from the input and output via KeyBERT, extracts their hidden states from a chosen decoder layer, and computes an adapted HSIC-like score between the two representation sets. The score is thresholded to flag hallucinations. The paper's stated mechanism is that hallucinations correspond to a 'statistical decoupling' between the LM's internal representations of input context and generated output. The evaluation covers SQuAD, RACE, NQ, and TriviaQA with six Llama and Gemma models and compares HIDE against perplexity, energy, LN-entropy, lexical similarity, and Eigenscore under sentence-similarity, ROUGE-L, and exact-match correctness measures. The main reported results are an average ~29% relative AUC-ROC improvement over the best single-pass baseline, an average ~3% relative improvement over multi-pass methods, and ~51% lower computation time than multi-pass methods. Ablations examine token budget, layer choice, kernel choice, HSIC estimator, and token-selection strategy.

Significance. The empirical evaluation is broad and the AUC improvements over single-pass baselines are consistent, which makes HIDE a potentially useful practical contribution if the findings hold. The paper also ships open-source code and reports threshold-independent ranking metrics (AUC-ROC, PCC), which are not affected by threshold tuning. The main weakness is that the theoretical foundation does not support the stated mechanism: the HIDE score is computed from independently selected input/output tokens paired by rank, without a joint distribution, so Eq. (4) is not an HSIC estimate of input-output dependence in the sense of Lemma 1. The empirical score may still discriminate hallucinations as a heuristic, but the 'statistical decoupling' claim and the interpretability story need to be either made rigorous or explicitly downgraded. The threshold selection in Section 6.4 also uses evaluation labels, and some per-setting comparisons with multi-pass methods are overstated.

major comments (3)
  1. [Section 4.1, Section 4.4.1, Algorithm 1] Definition 2 and Section 4.4.1 construct the samples X and Y by independently selecting the top-neff KeyBERT tokens from the input and output and pairing them by rank order. Lemma 1 guarantees HSIC=0 iff independence only when (x_i,y_i) are drawn from a joint distribution P_XY; for a single generation no such joint distribution over input token types and output token types is defined, and the rank pairing is an arbitrary coupling. Definition 2 is also not well-defined for repeated token types, since the same token type can occur with different hidden states (Appendix C shows duplicated selected tokens). Consequently the score in Eq. (4) is not an estimator of HSIC between the LM's input and output representations, and the central claim that low scores indicate 'statistical decoupling' does not follow from the theory presented. The paper itself calls the score a 'heuristic' in Section 4.2, yet the abstract and Section 6 present the decoupling hypothesis as the mechanism. The same objection applies to the SVD alignment strategy in Section 4.4.1, where rows of independently projected matrices are paired by singular-direction order. Please either define a genuine input-output coupling (for example, attention-based token pairing) with a corresponding joint distribution, or explicitly reframe HIDE as an empirical heuristic and remove the HSIC/independence interpretation from the abstract and conclusions.
  2. [Section 6.4, Table 3] The binary decision threshold is tuned on the evaluation labels. In Section 6.4, for each (model, dataset) pair the threshold tau is chosen by maximizing G-Mean using exact-match ground-truth labels, and the average value tau=0.12 is then recommended as an operating point. This is test-label leakage for the binary decision procedure: any accuracy, G-Mean, or deployment-oriented statement built on this threshold is not an unbiased estimate of performance. The AUC-ROC and PCC results are threshold-independent and therefore retain their validity, but the threshold analysis should be performed on a held-out split (or the threshold should be derived from a separate calibration set), and the practical binary-decision claims should be limited accordingly.
  3. [Section 6.2, Table 2] Section 6.2 says that on TriviaQA Eigenscore achieves 'marginally higher AUC-ROC and PCC values,' but Table 2 shows much larger gaps in four of the six models: for Llama-3.2-3B the AUCs values are 76.13 (Eigenscore) versus 58.36 (HIDE); for Llama-3.2-3B-Instruct, 78.35 versus 58.43; for Llama-3-8B, 81.06 versus 65.65; and for Llama-3-8B-Instruct, 82.83 versus 61.70. The aggregate ~3% relative improvement over multi-pass methods is therefore driven by gains on other datasets, and the abstract's 'competitive with multi-pass methods' claim needs qualification by dataset and by output length, especially for short-answer factuality tasks.
minor comments (4)
  1. [Appendix A, Lemma 4] The proof relies on the unquantified approximation HSICHIDE roughly equal to (1 - 3/n) times HSICu and does not bound the difference; the asymptotic bias and consistency statements may be true, but the presented argument is not yet a proof. Please replace it with a precise rate-of-convergence argument.
  2. [Throughout] There are several typos and formatting issues: 'the the' in Section 1, 'actuality hallucinations' in Section 6.2, 'repeatation' in Section 8.2, and 'LLama' in the Figure 6 caption.
  3. [Table B.1] Under exact-match labels, Energy outperforms HIDE in some settings (for example, Gemma-2-9B-Instruct on SQuAD: 81.68 versus 72.62); the main text's 'outperforms other single-pass methods in almost all settings' claim should state which correctness measure it refers to.
  4. [Figure 1, Section 4.2] The score is described as an 'adapted variant of unbiased HSIC,' but Definition 3 is explicitly biased; consider calling it an HSIC-based score throughout to avoid confusing readers.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the HIDE score is independently computed and the headline AUC-ROC claims are threshold-independent.

full rationale

The derivation chain is self-contained. The HIDE score is computed from hidden states via an adapted HSIC estimator (Eq. 4), and the reported AUC-ROC and PCC results (Tables 2 and B.1) rank this continuous score against ground-truth labels without any fitted parameter entering the score itself. The threshold tau_avg = 0.12 is obtained by maximizing G-Mean on the evaluation labels (Section 6.4, Table 3), but it is used only for the binary decision rule and illustrative examples; it does not enter the AUC claims, so no fitted input is renamed as a prediction. Hyperparameters such as neff = 20, the middle-layer choice, and the RBF kernel are motivated by ablations on the same datasets; this is ordinary hyperparameter selection and does not reduce the central claim by construction. The only self-citation (Chakraborty and Masud 2024) is a background remark about creativity and hallucination, and it is not load-bearing for the detection method. The possible theoretical gap noted by a reader, namely that input and output tokens are selected independently by KeyBERT and paired by rank rather than drawn from a defined joint distribution, is a validity concern about the mechanism rather than a circularity, and therefore does not affect this verdict.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The core measurement depends on four hand-set or tuned parameters (token budget, kernel bandwidth, layer, threshold), and on the unexamined assumption that keyword-ranked token sets form a valid paired sample. No new physical or architectural entities are introduced.

free parameters (4)
  • token budget neff = 20
    Set to 20 based on ablation on SQuAD and NQ (Figure 3), showing diminishing returns beyond 15-20 tokens.
  • RBF kernel bandwidth gamma = unspecified (moderate value; degrades for gamma >= 0.1)
    Default bandwidth not given numerically; ablation (Figure 5b) shows stability for gamma in 1e-9 to 1e-3.
  • decoder layer l = middle layer (l_mid)
    Fixed to middle layer based on prior work (Azaria and Mitchell 2023; Skean et al. 2025); ablation (Figure 4) shows layer-agnostic behavior.
  • decision threshold tau = 0.12 (average across settings)
    Chosen per (dataset, model) by maximizing G-Mean on exact-match test labels, then averaged (Table 3, Section 6.4). This is a post-hoc fit to the evaluation data.
assumptions (4)
  • standard math RBF kernel is characteristic on R^d, so HSIC equals zero iff independence (Lemmas 1 and 2)
    Invoked in Section 3.1 to justify HSIC as an independence measure.
  • domain assumption Hidden states of LMs contain signals predictive of hallucination
    Based on cited prior work (Azaria and Mitchell 2023; Ji et al. 2024; Chen et al. 2024a) and on the paper's own experiments.
  • domain assumption KeyBERT keyword extraction selects tokens representative of meaning
    Used in Algorithm 1 without validation; only compared against SVD alignment in Section 7.5, not against other extractors.
  • ad hoc to paper The adapted HSIC estimator approximates the unbiased V-statistic as (1 - 3/n) times HSIC_u for n >= 4
    Stated in the Lemma 4 proof (Appendix A); used to claim asymptotic unbiasedness and consistency, but the approximation is loose for small n and is not an exact identity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations." pith.science (2026). https://pith.science/paper/HD2JQ2UV

@misc{pith2026250617748,
  author       = {Pith},
  title        = {Pith review of: HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HD2JQ2UV}},
  note         = {Machine review of arXiv:2506.17748}
}
read the original abstract

Contemporary Language Models (LMs), while impressively fluent, often generate content that is factually incorrect or unfaithful to the input context - a critical issue commonly referred to as 'hallucination'. This tendency of LMs to generate hallucinated content undermines their reliability, especially because these fabrications are often highly convincing and therefore difficult to detect. While several existing methods attempt to detect hallucinations, most rely on analyzing multiple generations per input, leading to increased computational cost and latency. To address this, we propose a single-pass, training-free approach for effective Hallucination detectIon via Decoupled rEpresentations (HIDE). Our approach leverages the hypothesis that hallucinations result from a statistical decoupling between an LM's internal representations of input context and its generated output. We quantify this decoupling using the Hilbert-Schmidt Independence Criterion (HSIC) applied to hidden-state representations extracted while generating the output sequence. We conduct extensive experiments on four diverse question answering datasets, evaluating both faithfulness and factuality hallucinations across six open-source LMs of varying scales and properties. Our results demonstrate that HIDE outperforms other single-pass methods in almost all settings, achieving an average relative improvement of ~29% in AUC-ROC over the best-performing single-pass strategy across various models and datasets. Additionally, HIDE shows competitive and often superior performance with multi-pass state-of-the-art methods, obtaining an average relative improvement of ~3% in AUC-ROC while consuming ~51% less computation time. Our findings highlight the effectiveness of exploiting internal representation decoupling in LMs for efficient and practical hallucination detection.

Figures

Figures reproduced from arXiv: 2506.17748 by the authors.

Figure 1
Figure 1. Our proposed method, HIDE is a single-pass, training-free approach for hallucination detection. For a given input prompt, HIDE extracts representative tokens from both the input and the LM-generated output. It then computes a HIDE score using an adapted variant of unbiased HSIC on the hidden state embeddings of these token sets from a specific LM layer. This score, reflecting input-output representational dependence… view at source ↗
Figure 2
Figure 2. Comparison of the average computation time required by different hallucination [PITH_FULL_IMAGE:figures/full_fig_p016_2.png] view at source ↗
Figure 4
Figure 4. Comparison of AUCs scores when using representation from differ￾ent layers of Llama-3-8B and Gemma-2-9B for HIDE￾score calculation. We ob￾serve that the performance of HIDE is almost agnostic to the chosen layer in case of both models, for both SQuAD and NQ datasets. SQuAD NQ Dataset 0 20 40 60 80 A U C s RBF Linear Polynomial Cosine Sigmoid Laplacian Exponential Periodic Matern (a) 10 1 10 3 10 5 10 7 10 9 for RBF … view at source ↗
Figures from the paper (4 more)
Figure 5
Figure 5. Figure 5: Comparison of AUCs scores across (a) different kernel functions, and (b) differ￾ent values of the hyperparameter γ for the RBF kernel used for HIDE-score calculation on the SQuAD and NQ datasets using Llama-3-8B. low-level or high-level semantic features exclusively. T…
Figure 6
Figure 6. Figure 6: Comparison of AUCs scores using the biased estimator HSIC \b and our adapted estima￾tor HSIC \HIDE for HIDE score computation, for SQuAD and NQ datasets using LLama-3-8B and Gemma￾2-9B. 0.0×10 11 0.8×10 11 1.6×10 11 2.4×10 11 HSICb Score 0 250 500 750 1000 1250 1500 17…
Figure 7
Figure 7. Figure 7: Distribution of [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 8
Figure 8. Figure 8: Comparison of AUCs scores for keyword￾based and SVD-based token selection strategies, for SQuAD and NQ datasets using LLama-3- 8B and Gemma-2-9B. both approaches perform nearly identically across datasets and architectures. The SVD method guarantees mathematical orthog…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CrossHallu: Do Hallucination Signals Generalize Across Languages and Domains in Large Language Model's Internals?

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Hallucination signals from LLM internals transfer across English–Arabic and Arabic domains for most models, depending on class separability and feature-space language alignment.

Reference graph

Works this paper leans on

75 extracted references · 58 canonical work pages · cited by 1 Pith paper

  1. [1]

    Azaria, Amos and Tom Mitchell. 2023. The internal state of an llm knows when it’s lying. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 967--976

  2. [2]

    Bender, Emily M., Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT '21, page 610–623, Association for Computing Machinery, New York, NY, USA

  3. [3]

    Brown, Tom, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott G...

  4. [4]

    Chakraborty, Tanmoy and Sarah Masud. 2024. The promethean dilemma of ai at the intersection of hallucination and creativity. Commun. ACM, 67(10):26–28

  5. [5]

    Chang, Haw-Shiuan and Andrew McCallum. 2022. Softmax bottleneck makes language models unable to represent multi-mode word distributions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics ( Volume 1: Long Papers ) , pages 8048--8073, Association for Computational Linguistics, Dublin, Ireland

  6. [6]

    Chen, Chao, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye. 2024 a . Inside: Llms' internal states retain the power of hallucination detection. In The Twelfth International Conference on Learning Representations

  7. [7]

    Chen, Jifan, Grace Kim, Aniruddh Sriram, Greg Durrett, and Eunsol Choi. 2024 b . Complex claim verification with evidence retrieved in the wild. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 3569--3587

  8. [8]

    Chern, I-Chun, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, and Pengfei Liu. 2023. Factool: Factuality detection in generative ai -- a tool augmented framework for multi-task and multi-domain scenarios. arXiv preprint arXiv:2307.13528

Show all 75 references
  1. [9]

    Chiang, Cheng-Han and Hung-Yi Lee. 2023. Can large language models be an alternative to human evaluations? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15607--15631

  2. [10]

    Chowdhery, Aakanksha, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vin...

  3. [11]

    Chuang, Yung-Sung, Yujia Xie, Hongyin Luo, Yoon Kim, James R Glass, and Pengcheng He. 2024. Dola: Decoding by contrasting layers improves factuality in large language models. In The Twelfth International Conference on Learning Representations

  4. [12]

    Duan, Hanyu, Yi Yang, and Kar Yan Tam. 2024. Do llms know about hallucination? an empirical investigation of llm's hidden states. arXiv preprint arXiv:2402.09733

  5. [13]

    Durmus, Esin, He He, and Mona Diab. 2020. FEQA : A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5055--5070, Association for Co...

  6. [14]

    Dziri, Nouha, Andrea Madotto, Osmar Za \"i ane, and Avishek Joey Bose. 2021. Neural path hunter: Reducing hallucination in dialogue systems via path grounding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2197--2214, Associat...

  7. [15]

    Falke, Tobias, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019. Ranking generated summaries by correctness: An interesting but challenging application for natural language inference. In Proceedings of the 57th Annual Meeting of the Association ...

  8. [16]

    Fan, Angela, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889--898, Association for Computational Linguistics, Melbourne, Australia

  9. [17]

    Grattafiori, Aaron, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Ar...

  10. [18]

    Gretton, Arthur, Ralf Herbrich, Alexander Smola, Olivier Bousquet, and Bernhard Sch \"o lkopf. 2005. Kernel methods for measuring independence. Journal of Machine Learning Research, 6(70):2075--2129

  11. [19]

    Hernandez, Danny, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, et al. 2022. Scaling laws and interpretability of learning from repeated data. arXiv preprint arXiv:2205.10487

  12. [20]

    Holtzman, Ari, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In International Conference on Learning Representations

  13. [21]

    Hong, Giwon, Aryo Pradipta Gema, Rohit Saxena, Xiaotang Du, Ping Nie, Yu Zhao, Laura Perez - Beltrachini, Max Ryabinin, Xuanli He, Cl \' e mentine Fourrier, and Pasquale Minervini. 2024. The hallucinations leaderboard - an open effort to measure hallucinations in large languag...

  14. [22]

    Huang, Lei, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Trans. Inf. ...

  15. [23]

    Huo, Siqing, Negar Arabzadeh, and Charles LA Clarke. 2023. Retrieving supporting evidence for llms generated answers. arXiv preprint arXiv:2306.13781

  16. [24]

    Ji, Ziwei, Delong Chen, Etsuko Ishii, Samuel Cahyawijaya, Yejin Bang, Bryan Wilie, and Pascale Fung. 2024. LLM internal states reveal hallucination risk faced with a query. In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, page...

  17. [25]

    Ji, Ziwei, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Comput. Surv., 55(12)

  18. [26]

    Joshi, Mandar, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ...

  19. [27]

    Kadavath, Saurav, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fo...

  20. [28]

    Kasai, Jungo, Keisuke Sakaguchi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir Radev, Noah A Smith, Yejin Choi, Kentaro Inui, et al. 2023. Realtime qa: What's the answer right now? Advances in neural information processing systems, 36:49025--49043

  21. [29]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Kwiatkowski, Tom, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Sla...

  22. [30]

    Laban, Philippe, Wojciech Kry \'s ci \'n ski, Divyansh Agarwal, Alexander R Fabbri, Caiming Xiong, Shafiq Joty, and Chien-Sheng Wu. 2023. Llms as factual reasoners: Insights from existing benchmarks and beyond. arXiv preprint arXiv:2305.14540

  23. [31]

    Ladhak, Faisal, Esin Durmus, Mirac Suzgun, Tianyi Zhang, Dan Jurafsky, Kathleen McKeown, and Tatsunori Hashimoto. 2023. When do pre-training biases propagate to downstream tasks? a case study in text summarization. In Proceedings of the 17th Conference of the European Chapter ...

  24. [32]

    Lai, Guokun, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. RACE : Large-scale R e A ding comprehension dataset from examinations. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 785--794, Association for Computatio...

  25. [33]

    Lee, Katherine, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022. Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (...

  26. [34]

    Lee, Minhyeok. 2023. A mathematical investigation of hallucination and creativity in gpt models. Mathematics, 11(10)

  27. [35]

    Li, Zuchao, Shitou Zhang, Hai Zhao, Yifei Yang, and Dongjie Yang. 2023. Batgpt: A bidirectional autoregessive talker from generative pre-trained transformer. arXiv preprint arXiv:2307.00360

  28. [36]

    Manning, Christopher Ré, Diana Acosta-Navas, Drew A

    Liang, Percy, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Nava...

  29. [37]

    Lin, Chin-Yew. 2004. ROUGE : A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74--81, Association for Computational Linguistics, Barcelona, Spain

  30. [38]

    Lin, Stephanie, Jacob Hilton, and Owain Evans. 2022. T ruthful QA : Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214--3252, Association for Computational ...

  31. [39]

    Lin, Zhen, Shubhendu Trivedi, and Jimeng Sun. 2024. Generating with confidence: Uncertainty quantification for black-box large language models. Transactions on Machine Learning Research

  32. [40]

    Lin, Zi, Jeremiah Zhe Liu, and Jingbo Shang. 2022. Towards collaborative neural-symbolic graph semantic parsing via uncertainty. In Findings of the Association for Computational Linguistics : ACL 2022 , pages 4160--4173, Association for Computational Linguistics, Dublin, Ireland

  33. [41]

    Liu, Bingbin, Jordan Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2023. Exposing attention glitches with flip-flop language modeling. Advances in Neural Information Processing Systems, 36:25549--25583

  34. [42]

    Liu, Weitang, Xiaoyun Wang, John Owens, and Yixuan Li. 2020. Energy-based Out -of-distribution Detection . In Advances in Neural Information Processing Systems , volume 33, pages 21464--21475, Curran Associates, Inc

  35. [43]

    Liu, Yijin, Xianfeng Zeng, Chenze Shao, Fandong Meng, and Jie Zhou. 2024. Instruction position matters in sequence generation with large language models. In Findings of the Association for Computational Linguistics ACL 2024, pages 11652--11663

  36. [44]

    Luo, Zheheng, Qianqian Xie, and Sophia Ananiadou. 2023. Chatgpt as a factual inconsistency evaluator for text summarization. arXiv preprint arXiv:2303.15621

  37. [45]

    Maleki, Negar, Balaji Padmanabhan, and Kaushik Dutta. 2024. AI Hallucinations: A Misnomer Worth Clarifying . In 2024 IEEE Conference on Artificial Intelligence (CAI), pages 133--138, IEEE Computer Society, Los Alamitos, CA, USA

  38. [46]

    Malinin, Andrey and Mark Gales. 2021. Uncertainty estimation in autoregressive structured prediction. In International Conference on Learning Representations

  39. [47]

    Manakul, Potsawee, Adian Liusie, and Mark Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9004--9017

  40. [48]

    Maynez, Joshua, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906--1919, Association for Computational Lingu...

  41. [49]

    Miao, Ning, Yee Whye Teh, and Tom Rainforth. 2023. Selfcheck: Using llms to zero-shot check their own step-by-step reasoning

  42. [50]

    Min, Sewon, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. Factscore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empiri...

  43. [51]

    Nan, Feng, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, and Bing Xiang. 2021. Entity-level factual consistency of abstractive text summarization. In Proceedings of the 16th Conference of the European Chapter of the Ass...

  44. [52]

    Narayanan Venkit, Pranav, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao Huang, and Shomir Wilson. 2023. Nationality bias in text generation. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 116--122, Associat...

  45. [53]

    Onoe, Yasumasa, Michael Zhang, Eunsol Choi, and Greg Durrett. 2022. Entity cloze by date: what lms know about unseen entities. In Findings of the Association for Computational Linguistics : NAACL 2022 , pages 693--702, Association for Computational Linguistics, Seattle, United States

  46. [54]

    Paullada, Amandalynne, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna. 2021. Data and its (dis) contents: A survey of dataset development and use in machine learning research. Patterns, 2(11)

  47. [55]

    Rajpurkar, Pranav, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383--2392

  48. [56]

    Reimers, Nils and Iryna Gurevych. 2019. Sentence- BERT : Sentence embeddings using S iamese BERT -networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNL...

  49. [57]

    Ren, Jie, Jiaming Luo, Yao Zhao, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, and Peter J Liu. 2023. Out-of-distribution detection and selective generation for conditional language models. In The Eleventh International Conference on Learning Representations

  50. [58]

    Scialom, Thomas, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. Q uest E val: Summarization asks for fact-based evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processi...

  51. [59]

    Sharma, Mrinank, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al. 2024. Towards understanding sycophancy in language models. In 12th International Conference on Learning Representat...

  52. [60]

    Shi, Weijia, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Wen-tau Yih. 2024. Trusting your evidence: Hallucinate less with context-aware decoding. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Ling...

  53. [61]

    Simhi, Adi, Itay Itzhak, Fazl Barez, Gabriel Stanovsky, and Yonatan Belinkov. 2025. Trust me, i'm wrong: High-certainty hallucinations in llms. arXiv preprint arXiv:2502.12964

  54. [62]

    Skean, Oscar, Md Rifat Arefin, Dan Zhao, Niket Nikul Patel, Jalal Naghiyev, Yann LeCun, and Ravid Shwartz-Ziv. 2025. Layer by layer: Uncovering hidden representations in language models. In Forty-second International Conference on Machine Learning

  55. [63]

    Song, Le, Alex Smola, Arthur Gretton, Justin Bedo, and Karsten Borgwardt. 2012. Feature selection via dependence maximization. Journal of Machine Learning Research, 13(47):1393--1434

  56. [64]

    Lanckriet

    Sriperumbudur, Bharath K., Arthur Gretton, Kenji Fukumizu, Bernhard Sch \"o lkopf, and Gert R.G. Lanckriet. 2010. Hilbert space embeddings and metrics on probability measures. Journal of Machine Learning Research, 11(50):1517--1561

  57. [65]

    Gretton, K

    Sriperumbudur, BK., A. Gretton, K. Fukumizu, G. Lanckriet, and B. Sch \"o lkopf. 2008. Injective hilbert space embeddings of probability measures. In Proceedings of the 21st Annual Conference on Learning Theory, pages 111--122, Max-Planck-Gesellschaft, Omnipress, Madison, WI, USA

  58. [66]

    Team, Gemma, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, L \'e onard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ram \'e , et al. 2024. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118

  59. [67]

    Touvron, Hugo, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, ...

  60. [68]

    Wang, Chaojun and Rico Sennrich. 2020. On exposure bias, hallucination and domain shift in neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 3544--3552, Association for Computational Linguistics, Online

  61. [69]

    Weidinger, Laura, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, Willia...

  62. [70]

    Xiong, Miao, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. 2024. Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. In The Twelfth International Conference on Learning Representations

  63. [71]

    Yang, Zhilin, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen. 2018. Breaking the softmax bottleneck: A high-rank rnn language model. In International Conference on Learning Representations

  64. [72]

    Zhang, Yue, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. 2023. Siren's song in the ai ocean: A survey on hallucination in large language models. arXiv pre...

  65. [73]

    Zheng, Shen, Jie Huang, and Kevin Chen-Chuan Chang. 2023. Why does chatgpt fall short in providing truthful answers? arXiv preprint arXiv:2304.10513

  66. [74]

    , " * write output.state after.block = add.period write newline

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence a...

  67. [75]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.