REVIEW 58 references
CHARM trains graph neural networks on token-attention graphs built from LLM computational traces and outperforms prior hallucination detectors on five benchmarks at token and response level.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 13:52 UTC pith:UJSBRHU7
Neural Message-Passing on Attention Graphs for Hallucination Detection
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
CHARM feeds these graphs into a graph neural network. The network updates token representations by passing messages along the attention edges, then reads out either a token-level hallucination score or a whole-response score. The authors show that two published attention heuristics, Lookback Lens and LLM-Check, are special cases of what CHARM can compute, under assumptions such as no attention thresholding and bounded text lengths.
On token-level benchmarks (NQ, CNN) CHARM with attention only improves AUROC by about 1.5 to 4 points over the strongest baselines. On response-level benchmarks (Movies, WinoBias, Math) it is best on Movies and Math, and adding activations helps on WinoBias and Math. Removing the graph structure lowers performance, and the method is robust to dropping most edges. Code is not yet released.
Core claim
The abstract states: 'We show that CHARM provably subsumes prior attention-based heuristics and, experimentally, it consistently outperforms other leading approaches across diverse benchmarks.' Concretely, Tables 1 and 2 report CHARM(att) or CHARM(att+act) reaching the best AUROC and AUPR among compared methods on NQ, CNN, Movies, WinoBias, and Math, with the graph ablation in Table 3 attributing part of the gain to message passing on the attention-induced topology.
Load-bearing premise
The graph topology built in Section 3 is assumed to carry the predictive signal: edges with attention below tau=0.05 are dropped, prompt-to-prompt edges are removed (Section D.1.2), and only the remaining directed attention connections are used for message passing. If hallucination cues live primarily in low-attention edges or in prompt-internal attention, the representation discards them. The paper ablates the overall graph structure (Table 3) and the threshold tau (Table 4), but never tests whether removing prompt-to-prompt edges is safe, so this modeling choice is the most fragile load-bearing premise.
Editorial analysis
A structured set of objections, weighed in public.
Axiom & Free-Parameter Ledger
free parameters (3)
- Attention threshold tau =
0.05
- Activation layer for CHARM(att+act-24) =
24
- GNN hyperparameters (learning rate, hidden dimension, number of layers, dropout, weight decay, schedulers, batch norm, r =
selected on validation AUPR per dataset
axioms (6)
- domain assumption Decoder-only transformer attention matrices are lower-triangular with positive entries after softmax normalization.
- domain assumption Teacher-forcing generation reproduces the computational traces that would occur during actual decoding.
- standard math MLP Universal Approximation Theorem holds for the required continuous functions.
- ad hoc to paper Attention scores are clipped away from zero for the LLM-Check expressiveness proof.
- ad hoc to paper Bounded prompt and response lengths in Proposition 1.
- domain assumption Deleting prompt-to-prompt edges preserves hallucination-relevant structure.
read the original abstract
Large Language Models (LLMs) often generate incorrect or unsupported content, known as hallucinations. Existing detection methods rely on heuristics or simple models over isolated computational traces such as activations, or attention maps. We unify these signals by representing them as attributed graphs, where tokens are nodes, edges follow attentional flows, and both carry features from attention scores and activations. Our approach, CHARM, casts hallucination detection as a graph learning task and tackles it by applying GNNs over the above attributed graphs. We show that CHARM provably subsumes prior attention-based heuristics and, experimentally, it consistently outperforms other leading approaches across diverse benchmarks. Our results shed light on the relevant role played by the graph structure and on the benefits of combining computational traces, whilst showing CHARM exhibits promising zero-shot performance on cross-dataset transfer.
Figures
Reference graph
Works this paper leans on
-
[1]
On the bottleneck of graph neural networks and its practical implications.International Conference on Learning Representations, 2021
Uri Alon and Eran Yahav. On the bottleneck of graph neural networks and its practical implications.International Conference on Learning Representations, 2021
2021
-
[2]
The internal state of an llm knows when it’s lying
Amos Azaria and Tom Mitchell. The internal state of an llm knows when it’s lying. InFindings of the Association for Computational Linguistics: EMNLP 2023, pages 967–976, 2023
2023
-
[3]
Learning on llm output signatures for gray-box behavior analysis
Guy Bar-Shalom, Fabrizio Frasca, Derek Lim, Yoav Gelberg, Yftah Ziser, Ran El-Yaniv, Gal Chechik, and Haggai Maron. Learning on llm output signatures for gray-box behavior analysis. arXiv:2503.14043, 2025
arXiv 2025
-
[4]
Araújo, Alex Vitvitskyi, Razvan Pascanu, and Petar Veličković
Federico Barbero, Andrea Banino, Steven Kapturowski, Dharshan Kumaran, João G.M. Araújo, Alex Vitvitskyi, Razvan Pascanu, and Petar Veličković. Transformers need glasses! information over-squashing in language tasks. InAdvances in Neural Information Processing Systems, volume 37, pages 98111–98142, 2024
2024
-
[5]
Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al. Relational inductive biases, deep learning, and graph networks.arXiv:1806.01261, 2018
Pith/arXiv arXiv 2018
-
[6]
Alexandra Bazarova, Aleksandr Yugay, Andrey Shulga, Alina Ermilova, Andrei Volodichev, Konstantin Polev, Julia Belikova, Rauf Parchiev, Dmitry Simakov, Maxim Savchenko, Andrey Savchenko, Serguei Barannikov, and Alexey Zaytsev. Hallucination detection in llms via topological divergence on attention graphs.arXiv:2504.10063, 2025
Pith/arXiv arXiv 2025
-
[7]
Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022
Yonatan Belinkov. Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022
2022
-
[8]
Experiment tracking with weights and biases, 2020
Lukas Biewald. Experiment tracking with weights and biases, 2020. Software available from wandb.com
2020
-
[9]
Hallucination detection in llms using spectral features of attention maps.arXiv:2502.17598, 2025
Jakub Binkowski, Denis Janiak, Albert Sawczyn, Bogdan Gabrys, and Tomasz Kajdanowicz. Hallucination detection in llms using spectral features of attention maps.arXiv:2502.17598, 2025
arXiv 2025
-
[10]
Discovering latent knowledge in language models without supervision.arXiv:2212.03827, 2022
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. Discovering latent knowledge in language models without supervision.arXiv:2212.03827, 2022
Pith/arXiv arXiv 2022
-
[11]
Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization
Meng Cao, Yue Dong, and Jackie Cheung. Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors,Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3340–3354, Dublin, Ireland, May
-
[12]
Inside: Llms’ internal states retain the power of hallucination detection.arXiv:2402.03744, 2024
Chao Chen, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye. Inside: Llms’ internal states retain the power of hallucination detection.arXiv:2402.03744, 2024. 12
Pith/arXiv arXiv 2024
-
[13]
Yung-Sung Chuang, Linlu Qiu, Cheng-Yu Hsieh, Ranjay Krishna, Yoon Kim, and James R. Glass. Lookback lens: Detecting and mitigating contextual hallucinations in large language models using only attention maps. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1419–1436. Association for Computational Linguistics, 2024
2024
-
[14]
Batu El, Deepro Choudhury, Pietro Liò, and Chaitanya K. Joshi. Towards mechanistic inter- pretability of graph transformers via attention graphs. InICLR 2025 Workshop on Explainable AI for Science (XAI4Science), 2025
2025
-
[15]
Fast graph representation learning with PyTorch Geometric
Matthias Fey and Jan Eric Lenssen. Fast graph representation learning with PyTorch Geometric. InICLR Workshop on Representation Learning on Graphs and Manifolds, 2019
2019
-
[16]
Schoenholz, Patrick F
Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. Neural message passing for quantum chemistry. InInternational Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Research, pages 1263–1272. PMLR, 2017
2017
-
[17]
Bronstein, and Kirill Veselkov
Guadalupe Gonzalez, Shunwang Gong, Ivan Laponogov, Michael M. Bronstein, and Kirill Veselkov. Predictinganticancerhyperfoodswithgraphconvolutionalnetworks.Human Genomics, 15(33), 2021
2021
-
[18]
Nuno M Guerreiro, Elena Voita, and André FT Martins. Looking for a needle in a haystack: A comprehensive study of hallucinations in neural machine translation.arXiv:2208.05309, 2022
Pith/arXiv arXiv 2022
-
[19]
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.ACM Transactions on Information Systems, 2023
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.ACM Transactions on Information Systems, 2023
2023
-
[20]
Yuheng Huang, Jiayang Song, Zhijie Wang, Shengming Zhao, Huaming Chen, Felix Juefei-Xu, and Lei Ma. Look before you leap: An exploratory study of uncertainty measurement for large language models.arXiv:2307.10236, 2023
Pith/arXiv arXiv 2023
-
[21]
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, 2023
2023
-
[22]
Mistral 7b.arXiv:2310.06825, 2023
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b.arXiv:2310.06825, 2023
Pith/arXiv arXiv 2023
-
[23]
Language models (mostly) know what they know.arXiv:2207.05221, 2022
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. Language models (mostly) know what they know.arXiv:2207.05221, 2022
Pith/arXiv arXiv 2022
-
[24]
Kipf and Max Welling
Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. InInternational Conference on Learning Representations (ICLR), 2017
2017
-
[25]
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation.arXiv:2302.09664, 2023. 13
Pith/arXiv arXiv 2023
-
[26]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. Natural questions: A benchmark for question answering research.Transact...
2019
-
[27]
Inference- time intervention: Eliciting truthful answers from a language model.Advances in Neural Information Processing Systems, 36, 2024
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Inference- time intervention: Eliciting truthful answers from a language model.Advances in Neural Information Processing Systems, 36, 2024
2024
-
[28]
Deep learning-guided discovery of an antibiotic targeting acinetobacter baumannii
Gary Liu, Denise Catacutan, Khushi Rathod, Kyle Swanson, Wengong Jin, Jody Mohammed, Anush Chiappino-Pepe, Saad Syed, Meghan Fragis, Kenneth Rachwalski, Jakob Magolan, Michael Surette, Brian Coombes, Tommi Jaakkola, Regina Barzilay, James Collins, and Jonathan Stokes. Deep learning-guided discovery of an antibiotic targeting acinetobacter baumannii. Natur...
2023
-
[29]
A token-level reference-free hallucination detection benchmark for free-form text generation
Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan. A token-level reference-free hallucination detection benchmark for free-form text generation. arXiv:2104.08704, 2021
Pith/arXiv arXiv 2021
-
[30]
Decoupled weight decay regularization.arXiv:1711.05101, 2017
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.arXiv:1711.05101, 2017
Pith/arXiv arXiv 2017
-
[31]
Potsawee Manakul, Adian Liusie, and Mark JF Gales. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models.arXiv:2303.08896, 2023
Pith/arXiv arXiv 2023
-
[32]
Samuel Marks and Max Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets.arXiv:2310.06824, 2023
Pith/arXiv arXiv 2023
-
[33]
Bronstein
Federico Monti, Fabrizio Frasca, Davide Eynard, Damon Mannion, and Michael M. Bronstein. Fake news detection on social media using geometric deep learning. InICLR 2019 Workshop on Representation Learning on Graphs and Manifolds, 2019
2019
-
[34]
Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, and Yonatan Belinkov. Llms know more than they show: On the intrinsic representation of llm hallucinations.arXiv:2410.02707, 2024
Pith/arXiv arXiv 2024
-
[35]
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou, editors,Proceedings of the 2021 Confe...
2021
-
[36]
Pytorch: An imperative style, high- performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high- perf...
2019
-
[37]
Approximation theory of the mlp model in neural networks.Acta Numerica, 8: 143–195, 1999
Allan Pinkus. Approximation theory of the mlp model in neural networks.Acta Numerica, 8: 143–195, 1999
1999
-
[38]
Learning representations of irregular particle-detector geometry with distance-weighted graph networks.The European Physical Journal C, 79(7), 2019
Shah Rukh Qasim, Jan Kieseler, Yutaro Iiyama, and Maurizio Pierini. Learning representations of irregular particle-detector geometry with distance-weighted graph networks.The European Physical Journal C, 79(7), 2019
2019
-
[39]
Detecting and mitigating hallucinations in multilingual summarisation
Yifu Qiu, Yftah Ziser, Anna Korhonen, Edoardo Ponti, and Shay Cohen. Detecting and mitigating hallucinations in multilingual summarisation. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empirical Methods in Natural Lan- guage Processing, pages 8914–8932, Singapore, December 2023. Association for Computational ...
2023
-
[40]
Weakly supervised detection of hallucinations in llm activations.arXiv:2312.02798, 2023
Miriam Rateike, Celia Cintas, John Wamburu, Tanya Akumu, and Skyler Speakman. Weakly supervised detection of hallucinations in llm activations.arXiv:2312.02798, 2023
Pith/arXiv arXiv 2023
-
[41]
Vipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, SM Tonmoy, Aman Chadha, Amit P Sheth, and Amitava Das. The troubling emergence of hallucination in large language models–an extensive definition, quantification, and prescriptive remediations. arXiv:2310.04988, 2023
Pith/arXiv arXiv 2023
-
[42]
Liu, and Christopher D
Abigail See, Peter J. Liu, and Christopher D. Manning. Get to the point: Summarization with pointer-generator networks. InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1073–1083, Vancouver, Canada, 2017
2017
-
[43]
Adi Simhi, Jonathan Herzig, Idan Szpektor, and Yonatan Belinkov. Constructing benchmarks and interventions for combating hallucinations in llms.arXiv:2404.09971, 2024
Pith/arXiv arXiv 2024
-
[44]
On early detection of hallucinations in factual question answering
Ben Snyder, Marius Moisescu, and Muhammad Bilal Zafar. On early detection of hallucinations in factual question answering. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 2721–2732, 2024
2024
-
[45]
Llm-check: Investigating detection of hallucinations in large language models
Gaurang Sriramanan, Siddhant Bharti, Vinu Sankar Sadasivan, Shoumik Saha, Priyatham Kattakinda, and Soheil Feizi. Llm-check: Investigating detection of hallucinations in large language models. InAdvances in Neural Information Processing Systems, volume 37, pages 34188–34216, 2024
2024
-
[46]
Stokes, Kevin Yang, Kyle Swanson, Wengong Jin, Andres Cubillos-Ruiz, Nina M
Jonathan M. Stokes, Kevin Yang, Kyle Swanson, Wengong Jin, Andres Cubillos-Ruiz, Nina M. Donghia, Craig R. MacNair, Shawn French, Lindsey A. Carfrae, Zohar Bloom-Ackermann, Victoria M. Tran, Anush Chiappino-Pepe, Ahmed H. Badran, Ian W. Andrews, Emma J. Chory, George M. Church, Eric D. Brown, Tommi S. Jaakkola, Regina Barzilay, and James J. Collins. A dee...
2020
-
[47]
Bench- marking hallucination in large language models based on unanswerable math word problem
Yuhong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Hui Zhao. Bench- marking hallucination in large language models based on unanswerable math word problem. arXiv:2403.03558, 2024
Pith/arXiv arXiv 2024
-
[48]
Chamberlain, Xiaowen Dong, and Michael M
Jake Topping, Francesco Di Giovanni, Benjamin P. Chamberlain, Xiaowen Dong, and Michael M. Bronstein. Understanding over-squashing and bottlenecks on graphs via curvature. InInterna- tional Conference on Learning Representations, 2022. 15
2022
-
[49]
Llama 2: Open foundation and fine-tuned chat models.arXiv:2307.09288, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Harts...
Pith/arXiv arXiv 2023
-
[50]
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation.arXiv:2307.03987, 2023
Pith/arXiv arXiv 2023
-
[51]
Alex Vitvitskyi, João G. M. Araújo, Marc Lackenby, and Petar Veličković. What makes a good feedforward computational graph? InProceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research. PMLR, 2025
2025
-
[52]
Martindale, and Marine Carpuat
Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, and Marine Carpuat. Un- derstanding and detecting hallucinations in neural machine translation via model introspection. Transactions of the Association for Computational Linguistics, 11:546–564, 2023
2023
-
[53]
Fan Yin, Jayanth Srinivasa, and Kai-Wei Chang. Characterizing truthfulness in large language model generations with local intrinsic dimension.arXiv:2402.18048, 2024
Pith/arXiv arXiv 2024
-
[54]
Mert Yuksekgonul, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar, Ranjita Naik, Hamid Palangi, Ece Kamar, and Besmira Nushi. Attention satisfies: A constraint-satisfaction lens on factual errors of language models.arXiv:2309.15098, 2023
Pith/arXiv arXiv 2023
-
[55]
Enhancing uncertainty-based hallucination detection with stronger focus
Tianhang Zhang, Lin Qiu, Qipeng Guo, Cheng Deng, Yue Zhang, Zheng Zhang, Chenghu Zhou, Xinbing Wang, and Luoyi Fu. Enhancing uncertainty-based hallucination detection with stronger focus. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 915–932. Association for Computational Linguistics, 2023
2023
-
[56]
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. Gender bias in coreference resolution: Evaluation and debiasing methods. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15–20, New Orleans, Louisiana,
2018
-
[57]
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. Representation engineering: A top-down approach to ai transparency.arXiv:2310.01405, 2023. 16 A Expressiveness: Claims and Proofs Proposition (informal) 1.Equipped with a single-layer message-passing stackfmp...
Pith/arXiv arXiv 2023
-
[2018]
Association for Computational Linguistics
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.