Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Large Language Models, Knowledge Graphs and Search Engines: A Crossroads for Answering Users' Questions

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper argues that search engines, knowledge graphs, and large language models answer different pieces of the user's question, and that combining them—not choosing among them—is the research direction that will serve users best.

desk verdict A well-executed position paper that gives the community a useful taxonomy and roadmap for combining SEs, KGs, and LLMs; the capability matrix is asserted rather than measured, but that fits the genre. read the letter →

arxiv 2501.06699 v1 pith:V4ZLGAZH submitted 2025-01-12 cs.AI cs.IRcs.SC

classification cs.AIcs.IRcs.SC
keywords knowledgegraphslargelanguagemodelssearchenginesquestionansweringuserinformationneedstaxonomyretrieval-augmentedgenerationhybridsystems
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that search engines, knowledge graphs, and large language models are not competing ways to answer questions but complementary ones: each is strong exactly where the others are weak. To support this, it builds a taxonomy of user information needs—facts, explanations, planning, and advice, each with subtypes—and scores every technology against every need. It finds that no single technology covers the full spectrum: knowledge graphs excel at precise multi-hop and analytical factual queries, search engines bring broad, fresh, documented coverage, and language models synthesize, explain, and converse. The paper concludes that research should aim at hybrid systems, moving from simple augmentation to full amalgamation of the three. If the argument holds, betting the future on a single dominant technology would be a mistake.

What carries the argument

The engine of the paper is the paired analytical apparatus of Table 1 and Table 2. Table 1 compares search engines, knowledge graphs, and LLMs along a dozen capability dimensions (correctness, coverage, completeness, freshness, generation, synthesis, transparency, coherency, refinability, fairness, usability, expressivity, efficiency, multilingualism, personalization). Table 2 is the user-needs taxonomy: a structured partition of questions into four main classes and twelve subclasses, each annotated with example questions and a strength/weakness assessment for each technology. The argument runs through the mapping between these two tables—each technology's capability profile predicts which classes of user questions it can answer well—and the roadmap in Section 4 then proposes how to combine the technologies so that the mapping covers more of the taxonomy.

What would settle it

Collect a sample of real user queries, label each with one of the taxonomy's twelve subclasses, and have a strong search engine, a knowledge-graph question-answering system, and a retrieval-augmented LLM answer each; then compare which system produces the highest-quality answer per subclass. If the technology that the paper's Table 2 predicts as best is not the actual best for a majority of subclasses, the central complementarity mapping fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that the technologies are complementary rather than competitive, with the complementarity shown by a systematic comparison across dimensions such as correctness, coverage, completeness, freshness, generation, synthesis, transparency, and expressivity. Its key move is to view these dimensions through the lens of the user's information need: a taxonomy that divides questions into facts (popular, long-tail, dynamic, multi-hop, analytical), explanations (commonsense, causal, exploratory), planning (instructive, recommendation, spatio-temporal), and advice (lifestyle, cultural, philosophical). Against this taxonomy the paper assigns each technology clear strengths and weaknesses—KGs excel at complex factual queries but lack nuance and are hard to use; SEs are broad and fresh but cannot synthesize across documents; LLMs are flexible and generative but hallucinate, lag on long-tail facts, and are opaque. The paper argues that where one is weak, another is often strong, and it lays out a four-phase roadmap—augmentation, ensemble, federation, amalgamation—for combining all three to emphasize the pros and suppress the cons.

Load-bearing premise

The paper assumes that its taxonomy of user information needs matches how real questions actually divide, and that the capability scores it assigns to each technology still hold once the technologies are combined.

Editorial extensions

If this is right

  • If the paper is right, question-answering systems should be designed as hybrids that route each query to the technology best suited to its category, rather than relying on a single model or engine.
  • Knowledge graphs will remain load-bearing for multi-hop and analytical factual queries even as LLMs improve, because no amount of latent text statistics substitutes for explicit join and aggregation operators.
  • Retrieval augmentation (SE for LLM) is a necessary but not sufficient fix: the long-tail failure persists with RAG, so structured knowledge must be part of the answer.
  • Benchmarks for question answering should be built from the taxonomy's categories, so that progress is measured across the full spectrum of user needs, not just popular factual questions.
  • The eventual amalgamation goal implies that research on combined representations and aligned tokens (tying the LLM's textual 'Turing award' to the KG's entity and the SE's index term) is a concrete, high-value direction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy could double as an evaluation harness: a benchmark that samples each of the twelve subclasses from real query logs would test whether the predicted best technology for each category actually wins, something the paper does not do.
  • Real user questions are often hybrids of the categories (e.g., dynamic plus multi-hop, or analytical plus advice), so a practical system will need to decompose a query into sub-tasks rather than assign it to a single box.
  • The capability scores are a snapshot: as LLMs gain tool use, longer context, and live retrieval, some assignments (notably multi-hop and analytical) may shift, so the durable contribution is the framework, not the current cell values.
  • A transparency-aware design would follow from the authors' own emphasis: for high-stakes factual answers, a hybrid that surfaces KG provenance or SE sources alongside LLM synthesis may earn more user trust than the most fluent single model.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This position paper argues that Search Engines (SEs), Knowledge Graphs (KGs), and Large Language Models (LLMs) are complementary technologies for answering users' questions, and that future work should combine them according to a taxonomy of user information needs. The paper introduces a capability comparison across many dimensions (correctness, coverage, completeness, freshness, generation, etc.), proposes a taxonomy of user information needs (Table 2) with subcategories such as Popular, Long-Tail, Multi-hop, Analytical, Explanations, Planning, and Advice, and derives a four-phase roadmap for combining the three technologies (augmentation, ensemble, federation, amalgamation). The central claim is that no single technology is sufficient across the spectrum of user needs, and that hybrid systems can leverage the strengths of each while mitigating their weaknesses.

Significance. If the paper's central thesis holds, it provides a valuable organizing framework for the rapidly growing body of work on LLMs, KGs, and SEs. The explicit focus on end-user information needs is a useful corrective to technology-centric research, and the proposed taxonomy and roadmap could guide research agendas and benchmark design. The paper also includes concrete, well-chosen illustrative examples (Figures 1 and 2) that effectively convey the failure modes of LLMs on multi-hop and long-tail factual queries. As a position/vision piece, it does not claim to offer new empirical results; its value lies in synthesis and agenda-setting. However, the paper's load-bearing capability assignments in Table 2 are asserted rather than systematically validated, and this limits the evidential strength of the complementarity premise.

major comments (3)
  1. [Section 3, Table 2] The capability matrix in Table 2 is the empirical foundation for the paper's central claim that 'where one is weak, often another is strong' (Section 4) and for the delegation strategy in Section 4.4, but the entries are based on the authors' qualitative judgment and a small number of anecdotal examples rather than a systematic survey or benchmark evaluation. The paper should either ground these assignments in a broader set of existing results (e.g., citing benchmark studies for each cell) or, more realistically for a position paper, explicitly reframe the matrix as a falsifiable research hypothesis and add a validation protocol to the roadmap. Without such a protocol, the roadmap's key premise is untestable as stated.
  2. [Section 3, Table 2 (Analytical and Multi-hop rows)] Several Table 2 entries for LLMs appear time-sensitive and understate current capabilities, despite the Section 2 caveat that LLMs are considered 'in isolation.' For instance, the Analytical row states that LLMs have 'no datatypes, no aggregation,' but contemporary LLMs can generate and execute code or call calculators, and the Multi-hop row labels LLM reasoning as 'latent' without acknowledging chain-of-thought and other inference-time techniques. If the matrix is meant to describe isolated LLMs, this caveat should be repeated in Section 3 and in the table; if it is meant to describe deployed systems, the assignments need to be revised for current models. Either way, the paper should discuss how these assignments change as LLMs acquire tool use and RAG, since the roadmap itself proposes such combinations.
  3. [Section 3, Taxonomy and Table 2] The taxonomy is presented as a categorization of user information needs, but several categories are not mutually exclusive: for example, 'Recommendation' and 'Spatio-temporal' under Planning blend factual and subjective criteria, and 'Exploratory' describes a user behavior (recognizing an answer when seen) rather than a question type. Since Section 4.4 proposes delegating query types to specific technologies (e.g., multi-hop queries to KGs), the paper should clarify whether the categories are intended as a partition, a multi-label scheme, or a set of prototypical examples, and how a hybrid system should handle queries that mix categories. This clarification is needed to make the roadmap actionable.
minor comments (4)
  1. [Figures 1 and 2] The examples would be more informative if the model names, versions, and retrieval configurations (e.g., whether RAG was enabled and which search backend was used) were reported, since LLM behavior varies widely across versions and settings.
  2. [Section 2, Coherency] The claim that 'we can prove query equivalence' for SEs is too strong because ranking is typically non-deterministic and the paper itself notes that SE ranking shows variance; consider restricting the claim to the set of results rather than their ordering.
  3. [Section 4.3, LLM for KG] The figures '100 million entities, 1 billion facts in Wikidata' are likely outdated for 2025; please update to current counts or state that they are approximate as of the time of writing.
  4. [Section 4.4] The distinction between the 'ensemble' and 'federation' phases is not crisp, since both involve delegating queries or sub-tasks to different components; a sentence clarifying the difference (e.g., federation involves recursive sub-task delegation with feedback) would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the complementarity thesis is an expert synthesis supported by original examples and external benchmarks, not derived from or equivalent to its own inputs.

full rationale

This paper is a position/roadmap essay rather than a derivation: it contains no equations, no fitted parameters, no benchmark that is then used to predict itself, and no uniqueness theorem invoked to rule out alternatives. Its central claim—that SEs, KGs and LLMs are complementary—is supported by original anecdotal failure/success examples (Figures 1–2), expert-judgment capability matrices (Tables 1–2), and a body of cited literature including independent benchmark studies. The author self-citations that appear (KG survey [17], Wikidata [47], Head-to-Tail [42], Wikifunctions [46]) are used as background references or as pointers to externally evaluable resources and benchmarks; the complementarity claim is not justified exclusively or primarily by those self-citations, so none is load-bearing. Even if Table 2's capability assignments are contestable or in need of empirical validation, that is a correctness/evidence concern, not circularity: the conclusions rest on the table by assertion, not by construction. No step reduces to its own input, and no fitted quantity is renamed as a prediction.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no free parameters or invented entities; it is a qualitative comparison. The axioms it rests on are standard definitions of the three technologies (drawn from cited surveys), the empirical claim that LLMs hallucinate (illustrated anecdotally), and the paper's own taxonomy, which is unvalidated.

assumptions (4)
  • domain assumption LLMs can hallucinate and produce incomplete answers, as illustrated in Figures 1 and 2.
    The paper's motivation rests on these observed failure modes; they are presented anecdotally, not with a systematic evaluation.
  • domain assumption KGs can return complete answer sets with respect to their content via structured queries.
    Used to argue KGs excel at multi-hop/analytical queries (Section 3, Table 2); true by definition of query semantics, but completeness relative to user intent is not addressed.
  • domain assumption SEs index broad, fresh Web content and return documents, but cannot synthesize.
    Standard IR assumption (Appendix A.2) drawn from cited textbooks; also that SE coverage is broad.
  • ad hoc to paper The user-information-need taxonomy (Table 2) partitions real-world questions into stable categories.
    The taxonomy is introduced in this paper and is not empirically validated; the roadmap depends on it.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large Language Models, Knowledge Graphs and Search Engines: A Crossroads for Answering Users' Questions." pith.science (2026). https://pith.science/paper/V4ZLGAZH

@misc{pith2026250106699,
  author       = {Pith},
  title        = {Pith review of: Large Language Models, Knowledge Graphs and Search Engines: A Crossroads for Answering Users' Questions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V4ZLGAZH}},
  note         = {Machine review of arXiv:2501.06699}
}
read the original abstract

Much has been discussed about how Large Language Models, Knowledge Graphs and Search Engines can be combined in a synergistic manner. A dimension largely absent from current academic discourse is the user perspective. In particular, there remain many open questions regarding how best to address the diverse information needs of users, incorporating varying facets and levels of difficulty. This paper introduces a taxonomy of user information needs, which guides us to study the pros, cons and possible synergies of Large Language Models, Knowledge Graphs and Search Engines. From this study, we derive a roadmap for future research.

Figures

Figures reproduced from arXiv: 2501.06699 by the authors.

Figure 1
Figure 1. Replies from three LLMs to a user question [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. LLM reply to a long-tail question Rather the growing consensus – and a growing body of research – indicates that LLMs and KGs complement each other [31, 32, 42, 52]. This paper offers an exploration of the strengths, limitations and potential synergies of SEs, KGs and LLMs from the perspective of an information-seeking user. Specifically, we first present an analysis of the fundamental strengths and limitations of t… view at source ↗
Figure 3
Figure 3. Potential combinations of SEs, KGs and LLMs [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Knowledge graph describing Manuel Blum through a lightweight LLM, with the resulting embedding vector used as a search-intent representation for matching and ranking. Matching involves locating occurrences of the user’s keywords and short phrases as terms in the index,…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synergizing LLMs and Knowledge Graphs: A Novel Approach to Software Repository-Related Question Answering

    cs.SE 2024-12 conditional novelty 4.0 of 10

    An LLM plus a repository knowledge graph answers software repository questions with 84% accuracy when few-shot chain-of-thought prompting is added, outperforming an intent-based bot and web-search GPT-4o.

Reference graph

Works this paper leans on

62 extracted references · 20 canonical work pages · cited by 1 Pith paper

  1. [1]

    Reutter, and Domagoj Vrgoc

    Renzo Angles, Marcelo Arenas, Pablo Barceló, Aidan Hogan, Juan L. Reutter, and Domagoj Vrgoc. 2017. Foundations of Modern Query Languages for Graph Databases. ACM Comput. Surv. 50, 5 (2017), 68:1–68:40. https://doi.org/10.1145/ 3104031

  2. [2]

    Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary G. Ives. 2007. DBpedia: A Nucleus for a Web of Open Data. In The Semantic Web, 6th International Semantic Web Conference, 2nd Asian Semantic Web Conference, ISWC 2007 + ASWC 2007, Busan, Korea, November 11-15, 2007 (Lecture Notes in Computer Science, Vol. 4825), Kar...

  3. [3]

    Ribeiro-Neto

    Ricardo Baeza-Yates and Berthier A. Ribeiro-Neto. 2011. Modern Information Retrieval - the concepts and technology behind search, Second edition . Pearson Education Ltd., Harlow, England. http://www.mir2ed.org/

  4. [4]

    Hudson, Ehsan Adeli, Russ B

    Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ B. Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen Creel, Jared Quincy Davis, Dorottya Demszky, Chris Donahue, Moussa Doumbouya, Esin...

  5. [5]

    Yi Chang and Hongbo Deng. 2020. Query understanding for search engines . Springer

  6. [6]

    Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking Large Language Models in Retrieval-Augmented Generation. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Inno- vative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intellig...

  7. [7]

    Ernest Davis and Gary Marcus. 2015. Commonsense reasoning and commonsense knowledge in artificial intelligence. Commun. ACM 58, 9 (2015), 92–103. https: //doi.org/10.1145/2701413

  8. [8]

    Gianluca Demartini. 2019. Implicit Bias in Crowdsourced Knowledge Graphs. In Companion of The 2019 World Wide Web Conference, WWW 2019, San Francisco, CA, USA, May 13-17, 2019, Sihem Amer-Yahia, Mohammad Mahdian, Ashish Goel, Geert-Jan Houben, Kristina Lerman, Julian J. McAuley, Ricardo Baeza-Yates, and Leila Zia (Eds.). ACM, 624–630. https://doi.org/10.1...

Show all 62 references
  1. [9]

    Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui. 2023. A Survey for In-context Learn- ing. CoRR abs/2301.00234 (2023). https://doi.org/10.48550/ARXIV.2301.00234 arXiv:2301.00234

  2. [10]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ah- mad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sra- vankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston...

  3. [11]

    Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. A Survey on RAG Meeting LLMs: To- wards Retrieval-Augmented Large Language Models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Minin...

  4. [12]

    Yixing Fan, Xiaohui Xie, Yinqiong Cai, Jia Chen, Xinyu Ma, Xiangsheng Li, Ruqing Zhang, and Jiafeng Guo. 2022. Pre-training Methods in Information Retrieval. Found. Trends Inf. Retr. 16, 3 (2022), 178–317. https://doi.org/10.1561/ 1500000100

  5. [13]

    José Emilio Labra Gayo, Eric Prud’hommeaux, Iovka Boneva, and Dimitris Kontokostas. 2017. Validating RDF Data . Morgan & Claypool Publishers. https://doi.org/10.2200/S00786ED1V01Y201707WBE016

  6. [14]

    Claudio Gutierrez and Juan F. Sequeda. 2021. Knowledge graphs. Commun. ACM 64, 3 (2021), 96–104. https://doi.org/10.1145/3418294

  7. [15]

    Xu Han, Tianyu Gao, Yankai Lin, Hao Peng, Yaoliang Yang, Chaojun Xiao, Zhiyuan Liu, Peng Li, Jie Zhou, and Maosong Sun. 2020. More Data, More Relations, More Context and More Openness: A Review and Outlook for Relation Extraction. In Proceedings of the 1st Conference of the As...

  8. [16]

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring Mathematical Problem Solving With the MATH Dataset. InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks ...

  9. [17]

    Rashid, Anisa Rula, Lukas Schmelzeisen, Juan F

    Aidan Hogan, Eva Blomqvist, Michael Cochez, Claudia d’Amato, Gerard de Melo, Claudio Gutierrez, Sabrina Kirrane, José Emilio Labra Gayo, Roberto Navigli, Se- bastian Neumaier, Axel-Cyrille Ngonga Ngomo, Axel Polleres, Sabbir M. Rashid, Anisa Rula, Lukas Schmelzeisen, Juan F. S...

  10. [18]

    Xu, Jun Araki, and Graham Neubig

    Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. How Can We Know What Language Models Know. Trans. Assoc. Comput. Linguistics 8 (2020), 423–438. https://doi.org/10.1162/TACL_A_00324

  11. [19]

    Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig

    Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active Retrieval Aug- mented Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Si...

  12. [20]

    Maria Keet, Edlira Vakaj, and Gerard de Melo

    Lucie-Aimée Kaffee, Russa Biswas, C. Maria Keet, Edlira Vakaj, and Gerard de Melo. 2023. Multilingual Knowledge Graphs and Low-Resource Languages: A Review. TGDK 1, 1 (2023), 10:1–10:19. https://doi.org/10.4230/TGDK.1.1.10

  13. [21]

    Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raf- fel. 2023. Large Language Models Struggle to Learn Long-Tail Knowledge. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Hon- olulu, Hawaii, USA (Proceedings of Machine Learnin...

  14. [22]

    Hadas Kotek, Rikker Dockum, and David Q. Sun. 2023. Gender bias and stereotypes in Large Language Models. In Proceedings of The ACM Collec- tive Intelligence Conference, CI 2023, Delft, Netherlands, November 6-9, 2023 , Michael S. Bernstein, Saiph Savage, and Alessandro Bozzon...

  15. [23]

    Dirk Lewandowski. 2023. Understanding Search Engines . Springer. https: //doi.org/10.1007/978-3-031-22789-9

  16. [24]

    Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advanc...

  17. [25]

    Matteo Lissandrini, Torben Bach Pedersen, Katja Hose, and Davide Mottin. 2020. Knowledge graph exploration: where are we and where are we going? SIGWEB Newsl. 2020, Summer (2020), 4:1–4:8. https://doi.org/10.1145/3409481.3409485

  18. [26]

    Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Comput. Surv. 55, 9 (2023), 195:1–195:35. https://doi.org/10.1145/3560815

  19. [27]

    Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Han- naneh Hajishirzi. 2023. When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. In Proceedings of the 61st Annual Meeting of the Association for Com...

  20. [28]

    Donald Metzler, Yi Tay, Dara Bahri, and Marc Najork. 2021. Rethinking search: making domain experts out of dilettantes. SIGIR Forum 55, 1 (2021), 13:1–13:27. https://doi.org/10.1145/3476415.3476428

  21. [29]

    Roberto Navigli and Simone Paolo Ponzetto. 2012. BabelNet: The automatic con- struction, evaluation and application of a wide-coverage multilingual semantic network. Artif. Intell. 193 (2012), 217–250. https://doi.org/10.1016/J.ARTINT. 2012.07.001

  22. [30]

    Pandu Nayak. 2019. Understanding searches better than ever before. Google Blog. https://blog.google/products/search/search-language-understanding-bert/

  23. [31]

    Jeff Z. Pan, Simon Razniewski, Jan-Christoph Kalo, Sneha Singhania, Jiaoyan Chen, Stefan Dietze, Hajira Jabeen, Janna Omeliyanenko, Wen Zhang, Matteo Lissandrini, Russa Biswas, Gerard de Melo, Angela Bonifati, Edlira Vakaj, Mauro Dragoni, and Damien Graux. 2023. Large Language...

  24. [32]

    Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu

  25. [33]

    Patil, Tianjun Zhang, Xin Wang, and Joseph E

    Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. 2023. Gorilla: Large Language Model Connected with Massive APIs. CoRR abs/2305.15334 (2023). https://doi.org/10.48550/ARXIV.2305.15334 arXiv:2305.15334

  26. [34]

    Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller. 2019. Language Models as Knowledge Bases?. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Internati...

  27. [35]

    Manning, Ste- fano Ermon, and Chelsea Finn

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Ste- fano Ermon, and Chelsea Finn. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In Advances in Neural Infor- mation Processing Systems 36: Annual Conference on Neura...

  28. [36]

    Liz Reid. 2024. Generative AI in Search: Let Google do the searching for you. Google Blog. https://blog.google/products/search/generative-ai-google-search- may-2024/

  29. [37]

    Siddharth Samsi, Dan Zhao, Joseph McDonald, Baolin Li, Adam Michaleas, Michael Jones, William Bergeron, Jeremy Kepner, Devesh Tiwari, and Vijay Gadepally. 2023. From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference. In IEEE High Performance Extre...

  30. [38]

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language Models Can Teach Themselves to Use Tools. In Advances in Neural Information Processing Systems 36: Annual...

  31. [39]

    Wei Shen, Yuhan Li, Yinan Liu, Jiawei Han, Jianyong Wang, and Xiaojie Yuan

  32. [40]

    Amit Singhal. 2012. Introducing the Knowledge Graph: things, not strings. Google Blog. https://www.blog.google/products/search/introducing-knowledge- graph-things-not/

  33. [41]

    Suchanek, Gjergji Kasneci, and Gerhard Weikum

    Fabian M. Suchanek, Gjergji Kasneci, and Gerhard Weikum. 2007. Yago: a core of semantic knowledge. In Proceedings of the 16th International Conference on World Wide Web, WWW 2007, Banff, Alberta, Canada, May 8-12, 2007 , Carey L. Hogan et al. Williamson, Mary Ellen Zurko, Pete...

  34. [42]

    Kai Sun, Yifan Ethan Xu, Hanwen Zha, Yue Liu, and Xin Luna Dong. 2023. Head-to-Tail: How Knowledgeable are Large Language Models (LLM)? A.K.A. Will LLMs Replace Knowledge Graphs? CoRR abs/2308.10168 (2023). https: //doi.org/10.48550/ARXIV.2308.10168 arXiv:2308.10168

  35. [43]

    Kunal Suri, Atul Singh, Prakhar Mishra, Swapna Sourav Rout, and Rajesh Sabapathy. 2023. Language Models sounds the Death Knell of Knowledge Graphs. CoRR abs/2301.03980 (2023). https://doi.org/10.48550/ARXIV.2301.03980 arXiv:2301.03980

  36. [44]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems 30: An- nual Conference on Neural Information Processing Systems ...

  37. [45]

    Blerta Veseli, Simon Razniewski, Jan-Christoph Kalo, and Gerhard Weikum. 2023. Evaluating the Knowledge Base Completion Potential of GPT. In Proceedings of the Conference on Empirical Methods in Natural Language Processing: Findings of EMNLP, Singapore, 2023. Association for C...

  38. [46]

    Denny Vrandecic. 2021. Building a multilingual Wikipedia. Commun. ACM 64, 4 (2021), 38–41. https://doi.org/10.1145/3425778

  39. [47]

    Denny Vrandecic and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Commun. ACM 57, 10 (2014), 78–85. https://doi.org/10.1145/ 2629489

  40. [48]

    Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc V

    Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry W. Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc V. Le, and Thang Luong. 2023. FreshLLMs: Refreshing Large Language Models with Search Engine Augmenta- tion. CoRR abs/2310.03214 (2023). https://doi.org/10.4855...

  41. [49]

    Quan Wang, Zhendong Mao, Bin Wang, and Li Guo. 2017. Knowledge Graph Embedding: A Survey of Approaches and Applications. IEEE Trans. Knowl. Data Eng. 29, 12 (2017), 2724–2743. https://doi.org/10.1109/TKDE.2017.2754499

  42. [50]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompt- ing Elicits Reasoning in Large Language Models. In Advances in Neural Infor- mation Processing Systems 35: Annual Conference on ...

  43. [51]

    Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and Philip S. Yu. 2021. A Comprehensive Survey on Graph Neural Networks. IEEE Trans. Neural Networks Learn. Syst. 32, 1 (2021), 4–24. https://doi.org/10.1109/ TNNLS.2020.2978386

  44. [52]

    Linyao Yang, Hongyang Chen, Zhao Li, Xiao Ding, and Xindong Wu. 2023. ChatGPT is not Enough: Enhancing Large Language Models with Knowledge Graphs for Fact-aware Language Modeling. CoRR abs/2306.11489 (2023). https: //doi.org/10.48550/ARXIV.2306.11489 arXiv:2306.11489

  45. [53]

    Linyao Yang, Hongyang Chen, Zhao Li, Xiao Ding, and Xindong Wu. 2024. Give us the Facts: Enhancing Large Language Models With Knowledge Graphs for Fact-Aware Language Modeling. IEEE Trans. Knowl. Data Eng. 36, 7 (2024), 3091–3110. https://doi.org/10.1109/TKDE.2024.3360454

  46. [54]

    Xiao Yang, Kai Sun, Hao Xin, Yushi Sun, Nikita Bhalla, Xiangsen Chen, Sajal Choudhary, Rongze Daniel Gui, Ziran Will Jiang, Ziyu Jiang, Lingkun Kong, Brian Moran, Jiaqi Wang, Yifan Ethan Xu, An Yan, Chenyu Yang, Eting Yuan, Hanwen Zha, Nan Tang, Lei Chen, Nicolas Scheffer, Yue...

  47. [55]

    Manning, Percy Liang, and Jure Leskovec

    Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christo- pher D. Manning, Percy Liang, and Jure Leskovec. 2022. Deep Bidirec- tional Language-Knowledge Graph Pretraining. In Advances in Neural Infor- mation Processing Systems 35: Annual Conference on Neural Info...

  48. [56]

    Trippas, Jeff Dalton, and Filip Radlinski

    Hamed Zamani, Johanne R. Trippas, Jeff Dalton, and Filip Radlinski. 2023. Con- versational Information Seeking. Found. Trends Inf. Retr. 17, 3-4 (2023), 244–456. https://doi.org/10.1561/1500000081

  49. [57]

    Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and Guoyin Wang. 2023. Instruction Tuning for Large Language Models: A Survey. CoRR abs/2308.10792 (2023). https://doi.org/10.48550/ARXIV.2308.10792 arXiv:2308.10792

  50. [58]

    Xiang Zhang, Senyu Li, Bradley Hauer, Ning Shi, and Grzegorz Kondrak. 2023. Don’t Trust ChatGPT when your Question is not in English: A Study of Multilin- gual Abilities and Types of LLMs. In Proceedings of the 2023 Conference on Empir- ical Methods in Natural Language Process...

  51. [59]

    Danna Zheng, Mirella Lapata, and Jeff Z. Pan. 2024. Large Language Models as Reliable Knowledge Bases? CoRR abs/2407.13578 (2024). https://doi.org/10. 48550/ARXIV.2407.13578 arXiv:2407.13578

  52. [60]

    things, not strings

    Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji-Rong Wen. 2023. Large Language Models for Information Retrieval: A Survey. CoRR abs/2308.07107 (2023). https://doi. org/10.48550/ARXIV.2308.07107 arXiv:2308.07107 A BACKGROUND ...

  53. [2023]

    IEEE Trans

    Entity Linking Meets Deep Learning: Techniques and Solutions. IEEE Trans. Knowl. Data Eng. 35, 3 (2023), 2556–2578. https://doi.org/10.1109/TKDE. 2021.3117715

  54. [2024]

    IEEE Transactions on Knowledge and Data Engineering (2024)

    Unifying Large Language Models and Knowledge Graphs: A Roadmap. IEEE Transactions on Knowledge and Data Engineering (2024). https://doi.org/10. 1109/TKDE.2024.3352100 (to appear)

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.