Pith. sign in

REVIEW 58 references

CHARM trains graph neural networks on token-attention graphs built from LLM computational traces and outperforms prior hallucination detectors on five benchmarks at token and response level.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

arxiv 2509.24770 v2 pith:UJSBRHU7 submitted 2025-09-29 cs.LG

Neural Message-Passing on Attention Graphs for Hallucination Detection

classification cs.LG
keywords attentioncharmdetectiongraphsactivationsattributedcomputationalgraph
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Large language models sometimes produce fluent but wrong text. Existing detectors either ask the model again, which is slow, or look at a single internal signal such as an activation vector or an attention matrix. This paper argues that those signals should be read together, as a graph. Each token becomes a node. An edge from token i to token j exists when token i attends to token j, and the edge carries the attention weights. Each node carries the token's self-attention and optionally an activation vector.

CHARM feeds these graphs into a graph neural network. The network updates token representations by passing messages along the attention edges, then reads out either a token-level hallucination score or a whole-response score. The authors show that two published attention heuristics, Lookback Lens and LLM-Check, are special cases of what CHARM can compute, under assumptions such as no attention thresholding and bounded text lengths.

On token-level benchmarks (NQ, CNN) CHARM with attention only improves AUROC by about 1.5 to 4 points over the strongest baselines. On response-level benchmarks (Movies, WinoBias, Math) it is best on Movies and Math, and adding activations helps on WinoBias and Math. Removing the graph structure lowers performance, and the method is robust to dropping most edges. Code is not yet released.

Core claim

The abstract states: 'We show that CHARM provably subsumes prior attention-based heuristics and, experimentally, it consistently outperforms other leading approaches across diverse benchmarks.' Concretely, Tables 1 and 2 report CHARM(att) or CHARM(att+act) reaching the best AUROC and AUPR among compared methods on NQ, CNN, Movies, WinoBias, and Math, with the graph ablation in Table 3 attributing part of the gain to message passing on the attention-induced topology.

Load-bearing premise

The graph topology built in Section 3 is assumed to carry the predictive signal: edges with attention below tau=0.05 are dropped, prompt-to-prompt edges are removed (Section D.1.2), and only the remaining directed attention connections are used for message passing. If hallucination cues live primarily in low-attention edges or in prompt-internal attention, the representation discards them. The paper ablates the overall graph structure (Table 3) and the threshold tau (Table 4), but never tests whether removing prompt-to-prompt edges is safe, so this modeling choice is the most fragile load-bearing premise.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Axiom & Free-Parameter Ledger

3 free parameters · 6 axioms · 0 invented entities

The listed axioms are the unproved background facts and modeling choices on which the central claim rests. The expressiveness theorems add bounded-length, zero-threshold, and clipping assumptions that the deployed system does not exactly satisfy. No invented physical entities are introduced; the attributed graph and CHARM are analytical constructs, not entities with independent falsifiable handles.

free parameters (3)
  • Attention threshold tau = 0.05
    Equation (1); chosen for default runs. Table 4 shows AUPR is stable from tau=0.001 to 0.1 while tau=0.5 degrades, so the value is a modeling choice rather than a fitted constant.
  • Activation layer for CHARM(att+act-24) = 24
    Section 5.1: layer 24 is used because Act-24 performed best among activation probes; on Math the authors also try layer 32 and obtain a small gain, so the layer choice is data-dependent.
  • GNN hyperparameters (learning rate, hidden dimension, number of layers, dropout, weight decay, schedulers, batch norm, r = selected on validation AUPR per dataset
    Appendix C.2, Table 7; these are standard validation-tuned hyperparameters and are not part of the theoretical claims.
axioms (6)
  • domain assumption Decoder-only transformer attention matrices are lower-triangular with positive entries after softmax normalization.
    Section 3 Preliminaries; the graph edge definition and the no-zero-attention claim in Proposition 1's proof rely on this structure.
  • domain assumption Teacher-forcing generation reproduces the computational traces that would occur during actual decoding.
    Section B.1: traces are extracted by teacher-forcing precomputed generations; if decoding-time distributions differ, the detector may not transfer to free generation.
  • standard math MLP Universal Approximation Theorem holds for the required continuous functions.
    Section A invokes Pinkus [37] to approximate the Lookback Lens ratio and the log function.
  • ad hoc to paper Attention scores are clipped away from zero for the LLM-Check expressiveness proof.
    Proposition 2 assumes a lower bound alpha_min > 0; the paper also clamps with epsilon=10^-6 for baselines, but this property is not guaranteed by the LLM itself.
  • ad hoc to paper Bounded prompt and response lengths in Proposition 1.
    The proof needs a compact domain for the ratio function; deployed applications can exceed the chosen bounds.
  • domain assumption Deleting prompt-to-prompt edges preserves hallucination-relevant structure.
    Section D.1.2: all prompt-to-prompt connections are removed for computational reasons and in alignment with Lookback Lens; no ablation tests this choice.

pith-pipeline@v1.3.0-alltime-deepseek · 21400 in / 14774 out tokens · 117589 ms · 2026-08-04T13:52:00.456154+00:00 · methodology

0 comments
read the original abstract

Large Language Models (LLMs) often generate incorrect or unsupported content, known as hallucinations. Existing detection methods rely on heuristics or simple models over isolated computational traces such as activations, or attention maps. We unify these signals by representing them as attributed graphs, where tokens are nodes, edges follow attentional flows, and both carry features from attention scores and activations. Our approach, CHARM, casts hallucination detection as a graph learning task and tackles it by applying GNNs over the above attributed graphs. We show that CHARM provably subsumes prior attention-based heuristics and, experimentally, it consistently outperforms other leading approaches across diverse benchmarks. Our results shed light on the relevant role played by the graph structure and on the benefits of combining computational traces, whilst showing CHARM exhibits promising zero-shot performance on cross-dataset transfer.

Figures

Figures reproduced from arXiv: 2509.24770 by Fabrizio Frasca, Guy Bar-Shalom, Haggai Maron, Yftah Ziser.

Figure 1
Figure 1. Figure 1: Overview of CHARM. We extract attention and activation matrices from LLM computations and build an attributed graph from them: edges and their features are derived from off-diagonal attention scores; node features are based on activations, and diagonal attention values. The resulting graph is processed by a GNN-based architecture, which outputs either token-level hallucination scores (as illustrated) or a … view at source ↗
Figure 2
Figure 2. Figure 2: HD with CHARM. The input is an attributed graph (shown in the bottom left). First fmp obtains refined node / token representations via msg-passing. Next, fpool aggregates these if response level predictions are required. Finally, a projection head, fpred, outputs the detection score. Starting from the original node (token) features XV , each layer in fmp calculates and updates hidden node representations b… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

58 extracted references · 23 linked inside Pith

  1. [1]

    On the bottleneck of graph neural networks and its practical implications.International Conference on Learning Representations, 2021

    Uri Alon and Eran Yahav. On the bottleneck of graph neural networks and its practical implications.International Conference on Learning Representations, 2021

  2. [2]

    The internal state of an llm knows when it’s lying

    Amos Azaria and Tom Mitchell. The internal state of an llm knows when it’s lying. InFindings of the Association for Computational Linguistics: EMNLP 2023, pages 967–976, 2023

  3. [3]

    Learning on llm output signatures for gray-box behavior analysis

    Guy Bar-Shalom, Fabrizio Frasca, Derek Lim, Yoav Gelberg, Yftah Ziser, Ran El-Yaniv, Gal Chechik, and Haggai Maron. Learning on llm output signatures for gray-box behavior analysis. arXiv:2503.14043, 2025

  4. [4]

    Araújo, Alex Vitvitskyi, Razvan Pascanu, and Petar Veličković

    Federico Barbero, Andrea Banino, Steven Kapturowski, Dharshan Kumaran, João G.M. Araújo, Alex Vitvitskyi, Razvan Pascanu, and Petar Veličković. Transformers need glasses! information over-squashing in language tasks. InAdvances in Neural Information Processing Systems, volume 37, pages 98111–98142, 2024

  5. [5]

    Battaglia, Jessica B

    Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al. Relational inductive biases, deep learning, and graph networks.arXiv:1806.01261, 2018

  6. [6]

    Hallucination detection in llms via topological divergence on attention graphs.arXiv:2504.10063, 2025

    Alexandra Bazarova, Aleksandr Yugay, Andrey Shulga, Alina Ermilova, Andrei Volodichev, Konstantin Polev, Julia Belikova, Rauf Parchiev, Dmitry Simakov, Maxim Savchenko, Andrey Savchenko, Serguei Barannikov, and Alexey Zaytsev. Hallucination detection in llms via topological divergence on attention graphs.arXiv:2504.10063, 2025

  7. [7]

    Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022

    Yonatan Belinkov. Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022

  8. [8]

    Experiment tracking with weights and biases, 2020

    Lukas Biewald. Experiment tracking with weights and biases, 2020. Software available from wandb.com

  9. [9]

    Hallucination detection in llms using spectral features of attention maps.arXiv:2502.17598, 2025

    Jakub Binkowski, Denis Janiak, Albert Sawczyn, Bogdan Gabrys, and Tomasz Kajdanowicz. Hallucination detection in llms using spectral features of attention maps.arXiv:2502.17598, 2025

  10. [10]

    Discovering latent knowledge in language models without supervision.arXiv:2212.03827, 2022

    Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. Discovering latent knowledge in language models without supervision.arXiv:2212.03827, 2022

  11. [11]

    Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization

    Meng Cao, Yue Dong, and Jackie Cheung. Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors,Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3340–3354, Dublin, Ireland, May

  12. [12]

    Inside: Llms’ internal states retain the power of hallucination detection.arXiv:2402.03744, 2024

    Chao Chen, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye. Inside: Llms’ internal states retain the power of hallucination detection.arXiv:2402.03744, 2024. 12

  13. [13]

    Yung-Sung Chuang, Linlu Qiu, Cheng-Yu Hsieh, Ranjay Krishna, Yoon Kim, and James R. Glass. Lookback lens: Detecting and mitigating contextual hallucinations in large language models using only attention maps. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1419–1436. Association for Computational Linguistics, 2024

  14. [14]

    Batu El, Deepro Choudhury, Pietro Liò, and Chaitanya K. Joshi. Towards mechanistic inter- pretability of graph transformers via attention graphs. InICLR 2025 Workshop on Explainable AI for Science (XAI4Science), 2025

  15. [15]

    Fast graph representation learning with PyTorch Geometric

    Matthias Fey and Jan Eric Lenssen. Fast graph representation learning with PyTorch Geometric. InICLR Workshop on Representation Learning on Graphs and Manifolds, 2019

  16. [16]

    Schoenholz, Patrick F

    Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. Neural message passing for quantum chemistry. InInternational Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Research, pages 1263–1272. PMLR, 2017

  17. [17]

    Bronstein, and Kirill Veselkov

    Guadalupe Gonzalez, Shunwang Gong, Ivan Laponogov, Michael M. Bronstein, and Kirill Veselkov. Predictinganticancerhyperfoodswithgraphconvolutionalnetworks.Human Genomics, 15(33), 2021

  18. [18]

    Looking for a needle in a haystack: A comprehensive study of hallucinations in neural machine translation.arXiv:2208.05309, 2022

    Nuno M Guerreiro, Elena Voita, and André FT Martins. Looking for a needle in a haystack: A comprehensive study of hallucinations in neural machine translation.arXiv:2208.05309, 2022

  19. [19]

    A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.ACM Transactions on Information Systems, 2023

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.ACM Transactions on Information Systems, 2023

  20. [20]

    Look before you leap: An exploratory study of uncertainty measurement for large language models.arXiv:2307.10236, 2023

    Yuheng Huang, Jiayang Song, Zhijie Wang, Shengming Zhao, Huaming Chen, Felix Juefei-Xu, and Lei Ma. Look before you leap: An exploratory study of uncertainty measurement for large language models.arXiv:2307.10236, 2023

  21. [21]

    Survey of hallucination in natural language generation

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, 2023

  22. [22]

    Mistral 7b.arXiv:2310.06825, 2023

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b.arXiv:2310.06825, 2023

  23. [23]

    Language models (mostly) know what they know.arXiv:2207.05221, 2022

    Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. Language models (mostly) know what they know.arXiv:2207.05221, 2022

  24. [24]

    Kipf and Max Welling

    Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. InInternational Conference on Learning Representations (ICLR), 2017

  25. [25]

    Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation.arXiv:2302.09664, 2023

    Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation.arXiv:2302.09664, 2023. 13

  26. [26]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. Natural questions: A benchmark for question answering research.Transact...

  27. [27]

    Inference- time intervention: Eliciting truthful answers from a language model.Advances in Neural Information Processing Systems, 36, 2024

    Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Inference- time intervention: Eliciting truthful answers from a language model.Advances in Neural Information Processing Systems, 36, 2024

  28. [28]

    Deep learning-guided discovery of an antibiotic targeting acinetobacter baumannii

    Gary Liu, Denise Catacutan, Khushi Rathod, Kyle Swanson, Wengong Jin, Jody Mohammed, Anush Chiappino-Pepe, Saad Syed, Meghan Fragis, Kenneth Rachwalski, Jakob Magolan, Michael Surette, Brian Coombes, Tommi Jaakkola, Regina Barzilay, James Collins, and Jonathan Stokes. Deep learning-guided discovery of an antibiotic targeting acinetobacter baumannii. Natur...

  29. [29]

    A token-level reference-free hallucination detection benchmark for free-form text generation

    Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan. A token-level reference-free hallucination detection benchmark for free-form text generation. arXiv:2104.08704, 2021

  30. [30]

    Decoupled weight decay regularization.arXiv:1711.05101, 2017

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.arXiv:1711.05101, 2017

  31. [31]

    Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models.arXiv:2303.08896, 2023

    Potsawee Manakul, Adian Liusie, and Mark JF Gales. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models.arXiv:2303.08896, 2023

  32. [32]

    The geometry of truth: Emergent linear structure in large language model representations of true/false datasets.arXiv:2310.06824, 2023

    Samuel Marks and Max Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets.arXiv:2310.06824, 2023

  33. [33]

    Bronstein

    Federico Monti, Fabrizio Frasca, Davide Eynard, Damon Mannion, and Michael M. Bronstein. Fake news detection on social media using geometric deep learning. InICLR 2019 Workshop on Representation Learning on Graphs and Manifolds, 2019

  34. [34]

    Llms know more than they show: On the intrinsic representation of llm hallucinations.arXiv:2410.02707, 2024

    Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, and Yonatan Belinkov. Llms know more than they show: On the intrinsic representation of llm hallucinations.arXiv:2410.02707, 2024

  35. [35]

    Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics

    Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou, editors,Proceedings of the 2021 Confe...

  36. [36]

    Pytorch: An imperative style, high- performance deep learning library

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high- perf...

  37. [37]

    Approximation theory of the mlp model in neural networks.Acta Numerica, 8: 143–195, 1999

    Allan Pinkus. Approximation theory of the mlp model in neural networks.Acta Numerica, 8: 143–195, 1999

  38. [38]

    Learning representations of irregular particle-detector geometry with distance-weighted graph networks.The European Physical Journal C, 79(7), 2019

    Shah Rukh Qasim, Jan Kieseler, Yutaro Iiyama, and Maurizio Pierini. Learning representations of irregular particle-detector geometry with distance-weighted graph networks.The European Physical Journal C, 79(7), 2019

  39. [39]

    Detecting and mitigating hallucinations in multilingual summarisation

    Yifu Qiu, Yftah Ziser, Anna Korhonen, Edoardo Ponti, and Shay Cohen. Detecting and mitigating hallucinations in multilingual summarisation. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empirical Methods in Natural Lan- guage Processing, pages 8914–8932, Singapore, December 2023. Association for Computational ...

  40. [40]

    Weakly supervised detection of hallucinations in llm activations.arXiv:2312.02798, 2023

    Miriam Rateike, Celia Cintas, John Wamburu, Tanya Akumu, and Skyler Speakman. Weakly supervised detection of hallucinations in llm activations.arXiv:2312.02798, 2023

  41. [41]

    The troubling emergence of hallucination in large language models–an extensive definition, quantification, and prescriptive remediations

    Vipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, SM Tonmoy, Aman Chadha, Amit P Sheth, and Amitava Das. The troubling emergence of hallucination in large language models–an extensive definition, quantification, and prescriptive remediations. arXiv:2310.04988, 2023

  42. [42]

    Liu, and Christopher D

    Abigail See, Peter J. Liu, and Christopher D. Manning. Get to the point: Summarization with pointer-generator networks. InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1073–1083, Vancouver, Canada, 2017

  43. [43]

    Constructing benchmarks and interventions for combating hallucinations in llms.arXiv:2404.09971, 2024

    Adi Simhi, Jonathan Herzig, Idan Szpektor, and Yonatan Belinkov. Constructing benchmarks and interventions for combating hallucinations in llms.arXiv:2404.09971, 2024

  44. [44]

    On early detection of hallucinations in factual question answering

    Ben Snyder, Marius Moisescu, and Muhammad Bilal Zafar. On early detection of hallucinations in factual question answering. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 2721–2732, 2024

  45. [45]

    Llm-check: Investigating detection of hallucinations in large language models

    Gaurang Sriramanan, Siddhant Bharti, Vinu Sankar Sadasivan, Shoumik Saha, Priyatham Kattakinda, and Soheil Feizi. Llm-check: Investigating detection of hallucinations in large language models. InAdvances in Neural Information Processing Systems, volume 37, pages 34188–34216, 2024

  46. [46]

    Stokes, Kevin Yang, Kyle Swanson, Wengong Jin, Andres Cubillos-Ruiz, Nina M

    Jonathan M. Stokes, Kevin Yang, Kyle Swanson, Wengong Jin, Andres Cubillos-Ruiz, Nina M. Donghia, Craig R. MacNair, Shawn French, Lindsey A. Carfrae, Zohar Bloom-Ackermann, Victoria M. Tran, Anush Chiappino-Pepe, Ahmed H. Badran, Ian W. Andrews, Emma J. Chory, George M. Church, Eric D. Brown, Tommi S. Jaakkola, Regina Barzilay, and James J. Collins. A dee...

  47. [47]

    Bench- marking hallucination in large language models based on unanswerable math word problem

    Yuhong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Hui Zhao. Bench- marking hallucination in large language models based on unanswerable math word problem. arXiv:2403.03558, 2024

  48. [48]

    Chamberlain, Xiaowen Dong, and Michael M

    Jake Topping, Francesco Di Giovanni, Benjamin P. Chamberlain, Xiaowen Dong, and Michael M. Bronstein. Understanding over-squashing and bottlenecks on graphs via curvature. InInterna- tional Conference on Learning Representations, 2022. 15

  49. [49]

    Llama 2: Open foundation and fine-tuned chat models.arXiv:2307.09288, 2023

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Harts...

  50. [50]

    A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation.arXiv:2307.03987, 2023

    Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation.arXiv:2307.03987, 2023

  51. [51]

    Alex Vitvitskyi, João G. M. Araújo, Marc Lackenby, and Petar Veličković. What makes a good feedforward computational graph? InProceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research. PMLR, 2025

  52. [52]

    Martindale, and Marine Carpuat

    Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, and Marine Carpuat. Un- derstanding and detecting hallucinations in neural machine translation via model introspection. Transactions of the Association for Computational Linguistics, 11:546–564, 2023

  53. [53]

    Characterizing truthfulness in large language model generations with local intrinsic dimension.arXiv:2402.18048, 2024

    Fan Yin, Jayanth Srinivasa, and Kai-Wei Chang. Characterizing truthfulness in large language model generations with local intrinsic dimension.arXiv:2402.18048, 2024

  54. [54]

    Attention satisfies: A constraint-satisfaction lens on factual errors of language models.arXiv:2309.15098, 2023

    Mert Yuksekgonul, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar, Ranjita Naik, Hamid Palangi, Ece Kamar, and Besmira Nushi. Attention satisfies: A constraint-satisfaction lens on factual errors of language models.arXiv:2309.15098, 2023

  55. [55]

    Enhancing uncertainty-based hallucination detection with stronger focus

    Tianhang Zhang, Lin Qiu, Qipeng Guo, Cheng Deng, Yue Zhang, Zheng Zhang, Chenghu Zhou, Xinbing Wang, and Luoyi Fu. Enhancing uncertainty-based hallucination detection with stronger focus. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 915–932. Association for Computational Linguistics, 2023

  56. [56]

    Gender bias in coreference resolution: Evaluation and debiasing methods

    Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. Gender bias in coreference resolution: Evaluation and debiasing methods. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15–20, New Orleans, Louisiana,

  57. [57]

    updating

    Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. Representation engineering: A top-down approach to ai transparency.arXiv:2310.01405, 2023. 16 A Expressiveness: Claims and Proofs Proposition (informal) 1.Equipped with a single-layer message-passing stackfmp...

  58. [2018]

    Association for Computational Linguistics