REVIEW 5 major objections 6 minor 48 references
Addressing Hallucinations in Language Models with Knowledge Graph Embeddings as an Additional Modality
T0 review · 5 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Injecting knowledge-graph embeddings as a second modality cuts LLM hallucinations.
desk verdict A plausible KG-modality adapter for hallucination reduction, with a genuinely useful dataset, but the evaluation undercuts the central claim by not validating the oracle-to-predicted embedding transfer. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the two-stage injection of KG embeddings into the LLM's input sequence. First, a Text2Graph mapper—RoBERTa-large with an unfrozen encoder and a linear head, trained with MSE against PyTorch-BigGraph TransE embeddings of Wikidata entities—converts a text span into a predicted KG embedding. Second, a single linear adapter, trained with cross-entropy language modeling on WikiEntities texts, projects that KG embedding into the LLM embedding space, and the result is concatenated to the token embeddings between special tokens `<GRAPH_START>` and `<GRAPH_END>`. The mapper is independent of the LLM; the adapter is retrained per model. WikiEntities supplies the paired data: 3.2M Wikipedia texts with Wikidata entity spans, IDs, and corresponding embedding lookups.
What would settle it
Measure the entity-linking precision and recall of the Text2Graph mapper on the evaluation inputs (HaluEval questions, True-False statements, FEVER claims) and correlate those errors with downstream accuracy; also compare the full pipeline against an oracle version that uses gold KG embeddings on every task. If downstream gains largely vanish when gold embeddings are used on tasks currently scored with predicted embeddings, the claim that predicted KGs drive the improvement is undercut.
Extended reading notes
Core claim
The central discovery is that a frozen LLM can consume knowledge-graph embeddings as a discrete additional modality—wrapped in special tokens `<GRAPH_START>` and `<GRAPH_END>`—and use them to judge and generate more factually. The authors show that the Text2Graph mapper, a RoBERTa-large encoder with a linear head trained under MSE loss on WikiEntities, yields KG embeddings that, when projected by a trained linear adapter, improve hallucination detection accuracy on HaluEval (e.g., LLaMA 2-7B from 0.468 to 0.546 average) and True-False (Mistral 7B from 0.81 to 0.97 average), and FEVER verification (Mistral from 0.665 to 0.743), while MMLU, GSM8k, and other general benchmarks stay approximately flat. The method explicitly avoids external retrieval and keeps the LLM frozen; only the adapter and the mapper are trained. For the HaluEval QA task, the authors use ground-truth entity embeddings from the dataset rather than Text2Graph predictions, while other tasks rely on predicted KG embeddings.
Load-bearing premise
The training pipeline feeds the adapter ground-truth Wikidata entity embeddings, but at test time the adapter receives predicted KG embeddings from Text2Graph; the paper reports no entity-linking accuracy or embedding prediction error on the evaluation sets, and for the HaluEval QA task it falls back to provided embeddings, so the transfer from oracle to noisy inputs is an unmeasured load-bearing premise.
Editorial extensions
If this is right
- Adding KG embeddings as a modality improves hallucination detection and fact verification by about 2 to 10 percentage points absolute on HaluEval, True-False, and FEVER across the three tested LLMs, with the largest gains on LLaMA 2-7B and Mistral 7B.
- The approach can be adapted to any LLM by training only a linear adapter, since the Text2Graph mapper is model-agnostic and the underlying KG embedding space stays fixed.
- Because the LLM is frozen and no retrieval is used, the method adds factual grounding without changing general-task behavior; the reported MMLU, GSM8k, Winogrande, HellaSwag, and ARC scores remain roughly flat.
- The WikiEntities dataset of over 3 million Wikipedia texts annotated with Wikidata entities and spans can itself serve as training or evaluation data for entity linking models.
- In the HaluEval QA task, using predicted KG embeddings slightly outperforms using the provided entity embeddings for Mistral (0.521 vs 0.516), suggesting the mapper adds signal beyond exact entity lookup.
Reading between the lines
- If the transfer from oracle to predicted embeddings holds up, the same adapter-modality trick could be applied to other structured knowledge sources (e.g., relation triples or numeric facts) and other base models, since the mapper and adapter are lightweight and model-agnostic.
- A natural, testable extension is to report entity-linking accuracy and KG embedding prediction error on the evaluation sets; without those numbers, the observed gains could partly be driven by the dataset's entity distribution rather than by the KG semantics.
- The method could be combined with retrieval-augmented generation: retrieved entities could be fed as KG embeddings rather than text, potentially reducing the token cost of RAG while keeping factual grounding.
- Because the adapter is trained with ground-truth embeddings on Wikipedia text, its robustness on out-of-domain or time-sensitive facts—where Wikidata embeddings may be stale—is an open question the paper does not address.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a method to reduce hallucinations in large language models by injecting knowledge graph (KG) embeddings as an additional modality. The pipeline consists of a Text2Graph mapper (RoBERTa-large plus a linear layer) that predicts PyTorch-BigGraph TransE embeddings of Wikidata entities mentioned in the input text, and a linear adapter that projects these KG embeddings into the LLM's embedding space, inserted between trainable <GRAPH_START> and <GRAPH_END> tokens. The LLM itself remains frozen; only the mapper and adapter are trained. The authors introduce WikiEntities, a dataset of over 3.2 million Wikipedia texts annotated with Wikidata entities and their embeddings, and use it to train the adapter on ground-truth entity embeddings. They evaluate the method on HaluEval, True-False, and FEVER for hallucination detection, reporting improvements for Mistral 7B, LLaMA 2-7B, and LLaMA 3-8B, while also reporting results on MMLU, GSM8k, TruthfulQA, Winogrande, HellaSwag, and ARC to show no large degradation.
Significance. If the central claim holds, the paper offers a novel direction for incorporating factual knowledge into LLMs without external retrieval, and the WikiEntities dataset could be a valuable resource for entity linking and multimodal adaptation. The method is lightweight in that it only trains a mapper and a linear adapter. However, the current evidence is undermined by the unmeasured gap between the oracle embeddings used to train the adapter and the noisy predictions used at inference, an ambiguous QA setup, a lack of statistical significance testing, and an evaluation that is entirely within the Wikipedia/Wikidata distribution. These issues must be addressed before the claimed improvements can be considered reliable.
major comments (5)
- [§3.3 and §4.1] The adapter is trained on ground-truth Wikidata entity embeddings from WikiEntities (oracle embeddings), while at inference on True-False and FEVER it is fed embeddings predicted by Text2Graph. The paper never reports entity linking accuracy, embedding prediction error, or a direct comparison of adapter behavior under predicted versus oracle inputs. This transfer is load-bearing: without evidence that predicted embeddings are close enough to the oracle, the reported gains could be due to the special tokens, the linear projection, or Text2Graph only linking easy, unambiguous entities. Please report these metrics and include a control with random KG embeddings to confirm that the KG content, not the injection format, drives the improvement.
- [§4.1, Table 2] The QA setup is internally contradictory. The text states, "For this dataset, we use the provided entity embeddings rather than generating them with Text2Graph," yet Table 2 lists both "predicted KG embs" and "real KG embs" for the QA task, and the Discussion claims that predicted embeddings perform slightly better than real ones. This ambiguity matters because the QA row labeled "real KG embs" is an oracle condition and does not test the full pipeline. Please clarify which condition was used for each row, and if the predicted-KG-emb QA row is part of the official evaluation, describe exactly how Text2Graph was applied to the questions.
- [Tables 2–5] No error bars, confidence intervals, or significance tests are reported. Several improvements are small (e.g., Table 3, LLaMA 3 True-False average: 0.96 to 0.97; Table 5, Mistral TruthfulQA drops from 0.422 to 0.414), and it is impossible to judge whether the reported differences are meaningful without variance estimates. Please provide repeated-run results, bootstrap confidence intervals, or a paired significance test (e.g., McNemar's test for classification tasks).
- [§3.1, §4.1–4.3] All evaluation benchmarks (HaluEval, True-False, FEVER) are derived from Wikipedia, the entity embeddings come from Wikidata (the same knowledge source), and the adapter is trained on ground-truth Wikipedia entity annotations. This creates a circularity risk: the model may be exploiting benchmark-specific correlations between entity embeddings and answer patterns rather than learning to reason with KG embeddings in general. Please include an out-of-domain factual evaluation (e.g., questions from non-Wikipedia sources such as TriviaQA or a manually curated set of non-Wikipedia facts) and a control with shuffled or random entity embeddings to demonstrate that the KG content, rather than the injection mechanism, is responsible for the gains.
- [§3.2–3.3] The input format for the KG modality is underspecified, which affects reproducibility. Concretely, the paper does not state how many KG embeddings are produced for a text containing multiple entities, how these embeddings are ordered relative to the <GRAPH_START>/<GRAPH_END> tokens, or what happens when no entities are detected. In addition, the Text2Graph mapper is described only at a high level (span size of 20 tokens, MSE loss, AdamW, 1 epoch), with no details on embedding dimension, number of training samples, or mapper accuracy. Please provide these specifications, and consider releasing the exact preprocessing code.
minor comments (6)
- [§4.1] Typo: "dialoque history" should be "dialogue history".
- [Table 3] The column heading "Cieacf" is unclear and likely a typo (perhaps intended as "Companies" or another topic). Please correct.
- [§4.5] Typo: "LLaMA3-8B suceeds" should be "succeeds".
- [§4.1] The phrase "a model harnessing approach" is awkward; consider rephrasing to "an approach that harnesses the model" or similar.
- [Figure 3] The qualitative examples are illustrative, but they should be accompanied by a small quantitative evaluation on a set of similar questions to demonstrate that the improvement is systematic rather than cherry-picked.
- [§5] The availability statement refers to an anonymized repository; the final version should include a permanent DOI or link, and ideally the dataset itself should be released with a license and versioned.
Circularity Check
No significant circularity: the KG-injection pipeline is trained on WikiEntities and evaluated on external benchmarks; the oracle-embedding condition is explicitly labeled rather than disguised as a prediction.
full rationale
The paper's derivation chain is not circular. Text2Graph and the linear Adapter are trained on WikiEntities (Wikipedia text with Wikidata entities and PyTorch-BigGraph embeddings), and the models are then evaluated on the external HaluEval, True-False, and FEVER benchmarks. No equation in the paper defines the predicted KG embedding in terms of the benchmark label, and no fitted parameter is renamed as a prediction. The QA condition labeled 'real KG embs' is explicitly disclosed: 'For this dataset, we use the provided entity embeddings rather than generating them with Text2Graph.' This is an oracle condition, presented side-by-side with 'predicted KG embs' from the actual Text2Graph pipeline, so it is not a disguised fit. The fact that Wikidata/Wikipedia is also the source of many benchmark facts is the intended mechanism of KG-grounded factual reasoning, not a reduction of the method to its inputs. The unmeasured gap between oracle embeddings used during adapter training and noisy Text2Graph predictions at inference is a genuine empirical risk, but it is missing evidence rather than circularity. The one self-citation (Ref. [12], OmniFusion) appears only in related work on multimodal adapters and carries no load in the claimed derivation. Therefore the circularity score is 0.
Assumptions & free parameters
free parameters (5)
- Span size (20 tokens) =
20
- Adapter learning rate =
5e-3
- Text2Graph learning rate =
1e-4
- Number of KG embeddings per text (adapter training) =
Unspecified
- Chunk length for summarization =
RoBERTa context length
assumptions (4)
- domain assumption TransE embeddings from PyTorch-BigGraph encode facts about entities in a way accessible via a linear layer.
- domain assumption Adapters trained with language modeling loss on Wikipedia transfer to hallucination detection tasks.
- ad hoc to paper The evaluation benchmarks (HaluEval, True-False, FEVER) are drawn from Wikipedia, so the KG embeddings of their entities are in the training distribution.
- standard math RoBERTa-large, the Text2Graph encoder, is a standard architecture.
Cite this review
Pith. "Pith review of Addressing Hallucinations in Language Models with Knowledge Graph Embeddings as an Additional Modality." pith.science (2026). https://pith.science/paper/DGGEF34D
@misc{pith2026241111531,
author = {Pith},
title = {Pith review of: Addressing Hallucinations in Language Models with Knowledge Graph Embeddings as an Additional Modality},
year = {2026},
howpublished = {\url{https://pith.science/paper/DGGEF34D}},
note = {Machine review of arXiv:2411.11531}
}
read the original abstract
In this paper we present an approach to reduce hallucinations in Large Language Models (LLMs) by incorporating Knowledge Graphs (KGs) as an additional modality. Our method involves transforming input text into a set of KG embeddings and using an adapter to integrate these embeddings into the language model space, without relying on external retrieval processes. To facilitate this, we created WikiEntities, a dataset containing over 3 million Wikipedia texts annotated with entities from Wikidata and their corresponding embeddings from PyTorch-BigGraph. This dataset serves as a valuable resource for training Entity Linking models and adapting the described method to various LLMs using specialized adapters. Our method does not require fine-tuning of the language models themselves; instead, we only train the adapter. This ensures that the model's performance on other tasks is not affected. We trained an adapter for the Mistral 7B, LLaMA 2-7B (chat), and LLaMA 3-8B (instruct) models using this dataset and demonstrated that our approach improves performance on the HaluEval, True-False benchmarks and FEVER dataset. The results indicate that incorporating KGs as a new modality can effectively reduce hallucinations and improve the factual accuracy of language models, all without the need for external retrieval.
Figures
Reference graph
Works this paper leans on
-
[1]
AI@Meta. 2024. Llama 3 Model Card. (2024). https://github.com/meta-llama/ llama3/blob/main/MODEL_CARD.md
2024
-
[2]
Amos Azaria and Tom Mitchell. 2023. The Internal State of an LLM Knows When It’s Lying. arXiv:2304.13734 [cs.CL]
arXiv 2023
-
[3]
Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Ok- sana Yakhnenko. 2013. Translating Embeddings for Modeling Multi-relational Data. In Advances in Neural Information Processing Systems , C.J. Burges, L. Bot- tou, M. Welling, Z. Ghahramani, and K.Q. Weinberger (Eds.), Vol. 26. Cur- ran Associates, Inc. https://proceedings.neurips....
work page 2013
-
[4]
Patrice Béchard and Orlando Marquez Ayala. 2024. Reducing hal- lucination in structured outputs via Retrieval-Augmented Generation. arXiv:2404.08189 [cs.LG]
arXiv 2024
-
[5]
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. 2023. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks. CoRR abs/2312.14238 (2023). https://doi.org/10.48550/ARXIV.2312.14238 arXiv:2312.14238
-
[6]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have Solved Question An- swering? Try ARC, the AI2 Reasoning Challenge. CoRR abs/1803.05457 (2018). arXiv:1803.05457 http://arxiv.org/abs/1803.05457
arXiv 2018
-
[7]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training Verifiers to Solve Math Word Problems. CoRR abs/2110.14168 (2021). arXiv:2110.14168 https://arxiv.org/ abs/2110.14168
arXiv 2021
-
[8]
Mohnish Dubey, Debayan Banerjee, Abdelrahman Abdelkawi, and Jens Lehmann
Show all 48 references
-
[9]
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Ke Li, Junteng Jia, Yuan Shangguan, Jay Mahadeokar, Ozlem Kalinli, Christian Fuegen, and Mike Seltzer
-
[10]
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...
2023 doi
-
[11]
Olga Golovneva, Moya Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023. ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali,...
2023
-
[12]
Elizaveta Goncharova, Anton Razzhigaev, Matvey Mikhalchuk, Maxim Kurkin, Irina Abdullaeva, Matvey Skripkin, Ivan Oseledets, Denis Dimitrov, and Andrey Kuznetsov. 2024. OmniFusion Technical Report. arXiv:2404.06212 [cs.CV] https://arxiv.org/abs/2404.06212
2024 arXiv
-
[13]
Johannes Hoffart, Mohamed Amir Yosef, Ilaria Bordino, Hagen Fürstenau, Man- fred Pinkal, Marc Spaniol, Bilyana Taneva, Stefan Thater, and Gerhard Weikum
-
[14]
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. arXiv:2311.05232 ...
2023 arXiv
-
[15]
Siqing Huo, Negar Arabzadeh, and Charles Clarke. 2023. Retrieving Support- ing Evidence for Generative Question Answering. In Proceedings of the An- nual International ACM SIGIR Conference on Research and Development in In- formation Retrieval in the Asia Pacific Region (SIGIR...
2023
-
[16]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thoma...
2023 arXiv
-
[17]
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Re...
2017 doi
-
[18]
Kaveri Kale, Pushpak Bhattacharyya, Milind Gune, Aditya Shetty, and Rustom Lawyer. 2023. KGVL-BART: Knowledge Graph Augmented Visual Language BART for Radiology Report Generation. In Proceedings of the 17th Conference of the European Chapter of the Association for Computationa...
2023 doi
-
[19]
Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier. 2024. A Survey of Reinforcement Learning from Human Feedback. arXiv:2312.14925 [cs.LG] https://arxiv.org/abs/2312.14925
2024
-
[20]
Jing Yu Koh, Ruslan Salakhutdinov, and Daniel Fried. 2023. Grounding Language Models to Images for Multimodal Inputs and Outputs. InInternational Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceed- ings of Machine Learning Research, Vol....
2023
-
[21]
Adam Lerer, Ledell Wu, Jiajun Shen, Timothee Lacroix, Luca Wehrstedt, Abhijit Bose, and Alex Peysakhovich. 2019. PyTorch-BigGraph: A Large-scale Graph Embedding System. In Proceedings of the 2nd SysML Conference . Palo Alto, CA, USA
2019
-
[22]
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Gener- ation for Knowledge-Intensive NLP Tasks. In Ad...
2020
-
[23]
Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen
-
[24]
Wei Li, Hehe Fan, Yongkang Wong, Mohan Kankanhalli, and Yi Yang. 2024. TOPA: Extend Large Language Models for Video Understanding via Text-Only Pre-Alignment. arXiv:2405.13911 [cs.CV] https://arxiv.org/abs/2405.13911
2024 arXiv
- [25]
-
[26]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Smaranda Mure- san, Preslav Nakov, and Aline Villavic...
2022 doi
-
[27]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved Baselines with Visual Instruction Tuning. arXiv:2310.03744 [cs.CV] https://arxiv.org/abs/ 2310.03744
2024 arXiv
-
[28]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Vi- sual Instruction Tuning. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Al...
2023
-
[29]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)
2019 arXiv
-
[30]
Seungwhan Moon, Pararth Shah, Anuj Kumar, and Rajen Subba. 2019. OpenDi- alKG: Explainable Conversational Reasoning with Attention-based Walks over Knowledge Graphs. In Proceedings of the 57th Annual Meeting of the Associa- tion for Computational Linguistics , Anna Korhonen, D...
2019
-
[31]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schul- man, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Pe- ter Welinder, Paul Christiano, Jan Leike,...
2022 arXiv
-
[32]
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Commun. ACM 64, 9 (2021), 99–106
2021
-
[33]
Liu, and Christopher D
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get To The Point: Summarization with Pointer-Generator Networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Regina Barzilay and Min-Yen Kan (Ed...
2017 doi
-
[34]
Bin Sun, Yitong Li, Fei Mi, Fanhu Bie, Yiwei Li, and Kan Li. 2023. Towards Fewer Hallucinations in Knowledge-Grounded Dialogue Generation via Aug- mentative and Contrastive Knowledge-Dialogue. In Proceedings of the 61st An- nual Meeting of the Association for Computational Lin...
2023 doi
- [35]
-
[36]
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal
-
[37]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guil- laume Lample. 2023. LLaMA: Open and Efficient Foundation ...
2023 arXiv
-
[38]
Denny Vrandečić and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Commun. ACM 57, 10 (sep 2014), 78–85. https://doi.org/10. 1145/2629489
2014
-
[39]
Silei Xu, Shicheng Liu, Theo Culhane, Elizaveta Pertseva, Meng-Hsi Wu, Sina Semnani, and Monica Lam. 2023. Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over Wikidata. In Proceedings of the 2023 Conference on Empirical Methods ...
2023 doi
-
[40]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language...
2018
-
[41]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a Machine Really Finish Your Sentence? CoRR abs/1905.07830 (2019). arXiv:1905.07830 http://arxiv.org/abs/1905.07830
2019 arXiv
-
[42]
True" if the given statement is true and
Wenliang Zhong, Wenyi Wu, Qi Li, Rob Barton, Boxin Du, Shioulin Sam, Karim Bouyarmane, Ismail Tutar, and Junzhou Huang. 2024. Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment. arXiv:2406.02987 [cs.CV] h...
2024 arXiv
-
[854]
https://doi.org/10.18653/v1/P19-1081
-
[2011]
In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing , Regina Barzilay and Mark Johnson (Eds.)
Robust Disambiguation of Named Entities in Text. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing , Regina Barzilay and Mark Johnson (Eds.). Association for Computational Linguistics, Addressing Hallucinations in Language Models with Kn...
2011
-
[2018]
In NAACL-HLT
FEVER: a Large-scale Dataset for Fact Extraction and VERification. In NAACL-HLT
-
[2019]
In Proceedings of the 18th International Semantic Web Conference (ISWC)
LC-QuAD 2.0: A Large Dataset for Complex Question Answering over Wikidata and DBpedia. In Proceedings of the 18th International Semantic Web Conference (ISWC). Springer
-
[2023]
https://arxiv.org/abs/2305.11747
HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. https://arxiv.org/abs/2305.11747
-
[2024]
arXiv:2311.06753 [cs.CL] https://arxiv.org/abs/2311.06753
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs. arXiv:2311.06753 [cs.CL] https://arxiv.org/abs/2311.06753
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.