{"id":"ff83598e-7117-4cde-8f00-5078a916ba41","arxiv_id":"2501.06699","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Search engines, knowledge graphs, and LLMs each suit different user question types, so future systems should combine them according to a taxonomy of information needs.","lead":"This paper compares the strengths and weaknesses of search engines, knowledge graphs, and large language models, and proposes a taxonomy of user information needs to guide when each technology, or a combination, should answer a question. It is a roadmap essay, not a new experiment, aimed at steering future research on hybrid question-answering systems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's capability matrix is asserted, not measured; if current LLM/SE strengths differ, the complementarity premise weakens.","rationale":"The reader's weakest_assumption is exactly the unvalidated taxonomy and capability matrix; I agree that this is the most load-bearing element. The paper's argument is a position/roadmap rather than an empirical study, so the lack of systematic validation is a limitation but not a fatal flaw: the qualitative claims are plausible, well-argued, and consistent with a substantial body of existing work on RAG, KG QA, and hybrid systems. My proposed test would operationalize the concern, but even if the matrix were found to be partly outdated, the paper's core message—that hybrid systems should be explored around user information needs—would likely survive in adapted form. Thus the reader's ACCEPT verdict needs no change.","tokens_in":24486,"tokens_out":3623,"duration_ms":37685,"concrete_test":"Build a benchmark with, say, 100 queries per Table 2 subcategory, sampled with clear inclusion criteria. Evaluate a current SE (e.g., Google/Bing), a KG QA system (e.g., over Wikidata), and a state-of-the-art LLM with RAG and tool use on each query, measuring accuracy, completeness, and provenance. Then implement a simple ensemble that routes by predicted category and compare its overall performance against each single technology. If the hybrid does not outperform the best single technology by a meaningful margin, or if the Table 2 assignments fail to predict which technology wins per category, the complementarity premise would not be supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that SEs, KGs, and LLMs are complementary and should be combined according to user information needs (Section 4)—depends on the capability assignments in Table 2 and the taxonomy in Section 3. These assignments are based on the authors' judgment and a few anecdotal examples (Figures 1–2) with unspecified model versions, not on a systematic evaluation. In particular, Table 2 labels LLMs as only having 'latent reasoning' with 'hallucinations' for multi-hop queries, while KGs get 'formal reasoning'; yet modern LLMs with chain-of-thought, tool use, and code execution can answer many aggregation and multi-hop queries, while KGs are often incomplete for long-tail multi-hop (as the paper itself notes under Completeness in Section 2). Similarly, SEs are marked as 'no datatypes, no aggregation' for analytical queries, but current SEs already return direct answers, knowledge panels, and structured data snippets for many such queries. If the true strengths differ for contemporary systems, the premise 'where one is weak, often another is strong' may not hold sufficiently to justify the proposed hybrid roadmap. The paper provides no procedure for validating or updating the matrix, so the roadmap's empirical foundation is untestable as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that Search Engines (SEs), Knowledge Graphs (KGs), and Large Language Models (LLMs) are complementary technologies for answering users' questions, and that future work should combine them according to a taxonomy of user information needs. The paper introduces a capability comparison across many dimensions (correctness, coverage, completeness, freshness, generation, etc.), proposes a taxonomy of user information needs (Table 2) with subcategories such as Popular, Long-Tail, Multi-hop, Analytical, Explanations, Planning, and Advice, and derives a four-phase roadmap for combining the three technologies (augmentation, ensemble, federation, amalgamation). The central claim is that no single technology is sufficient across the spectrum of user needs, and that hybrid systems can leverage the strengths of each while mitigating their weaknesses.","tokens_in":24727,"tokens_out":4281,"duration_ms":43762,"significance":"If the paper's central thesis holds, it provides a valuable organizing framework for the rapidly growing body of work on LLMs, KGs, and SEs. The explicit focus on end-user information needs is a useful corrective to technology-centric research, and the proposed taxonomy and roadmap could guide research agendas and benchmark design. The paper also includes concrete, well-chosen illustrative examples (Figures 1 and 2) that effectively convey the failure modes of LLMs on multi-hop and long-tail factual queries. As a position/vision piece, it does not claim to offer new empirical results; its value lies in synthesis and agenda-setting. However, the paper's load-bearing capability assignments in Table 2 are asserted rather than systematically validated, and this limits the evidential strength of the complementarity premise.","major_comments":[{"comment":"The capability matrix in Table 2 is the empirical foundation for the paper's central claim that 'where one is weak, often another is strong' (Section 4) and for the delegation strategy in Section 4.4, but the entries are based on the authors' qualitative judgment and a small number of anecdotal examples rather than a systematic survey or benchmark evaluation. The paper should either ground these assignments in a broader set of existing results (e.g., citing benchmark studies for each cell) or, more realistically for a position paper, explicitly reframe the matrix as a falsifiable research hypothesis and add a validation protocol to the roadmap. Without such a protocol, the roadmap's key premise is untestable as stated.","section":"Section 3, Table 2"},{"comment":"Several Table 2 entries for LLMs appear time-sensitive and understate current capabilities, despite the Section 2 caveat that LLMs are considered 'in isolation.' For instance, the Analytical row states that LLMs have 'no datatypes, no aggregation,' but contemporary LLMs can generate and execute code or call calculators, and the Multi-hop row labels LLM reasoning as 'latent' without acknowledging chain-of-thought and other inference-time techniques. If the matrix is meant to describe isolated LLMs, this caveat should be repeated in Section 3 and in the table; if it is meant to describe deployed systems, the assignments need to be revised for current models. Either way, the paper should discuss how these assignments change as LLMs acquire tool use and RAG, since the roadmap itself proposes such combinations.","section":"Section 3, Table 2 (Analytical and Multi-hop rows)"},{"comment":"The taxonomy is presented as a categorization of user information needs, but several categories are not mutually exclusive: for example, 'Recommendation' and 'Spatio-temporal' under Planning blend factual and subjective criteria, and 'Exploratory' describes a user behavior (recognizing an answer when seen) rather than a question type. Since Section 4.4 proposes delegating query types to specific technologies (e.g., multi-hop queries to KGs), the paper should clarify whether the categories are intended as a partition, a multi-label scheme, or a set of prototypical examples, and how a hybrid system should handle queries that mix categories. This clarification is needed to make the roadmap actionable.","section":"Section 3, Taxonomy and Table 2"}],"minor_comments":[{"comment":"The examples would be more informative if the model names, versions, and retrieval configurations (e.g., whether RAG was enabled and which search backend was used) were reported, since LLM behavior varies widely across versions and settings.","section":"Figures 1 and 2"},{"comment":"The claim that 'we can prove query equivalence' for SEs is too strong because ranking is typically non-deterministic and the paper itself notes that SE ranking shows variance; consider restricting the claim to the set of results rather than their ordering.","section":"Section 2, Coherency"},{"comment":"The figures '100 million entities, 1 billion facts in Wikidata' are likely outdated for 2025; please update to current counts or state that they are approximate as of the time of writing.","section":"Section 4.3, LLM for KG"},{"comment":"The distinction between the 'ensemble' and 'federation' phases is not crisp, since both involve delegating queries or sub-tasks to different components; a sentence clarifying the difference (e.g., federation involves recursive sub-task delegation with feedback) would improve readability.","section":"Section 4.4"}],"recommendation":"major_revision","confidential_remarks":"This is a position/vision paper rather than an empirical study, and the referee's recommendation is based on the quality of the argumentation and roadmap rather than on novel experimental results. The main revision request concerns the empirical status of Table 2 and the taxonomy; if the authors reframe these as explicitly testable hypotheses and add a lightweight validation or evaluation methodology to the roadmap, the paper would be a strong fit for the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid roadmap paper, worth a serious look if you work on question answering or hybrid systems. The genuinely new bit is the user-need taxonomy (Table 2) and the four-phase roadmap (augmentation → ensemble → federation → amalgamation) that bring search engines into the KG/LLM comparison. The 13-dimension comparison in Table 1 is a useful extension of the existing surveys [31,32], and the paper says so. That is the right scope: it is a review/position contribution, not a new result.\n\nWhat it does well: the illustrative examples are carefully chosen and correct. Only Manuel Blum satisfies the Turing/Latin America query; the ACM Fellows long-tail example is real and verified against the ACM site. The qualitative claims line up with the literature. The roadmap sections are concrete, saying what research in each pairwise combination would look like, and the paper is honest about what is known and what is open.\n\nSoft spots: Table 2's capability matrix is asserted, not measured. The stress-test note has a point: modern LLMs with tool use and chain-of-thought can handle more aggregation and multi-hop queries than the 'latent reasoning, hallucinations' cell implies, and SEs already return direct answers for many analytical queries. So some cells are dated or wrong for current systems. But this is not a fatal flaw. The paper's argument is structural—the technologies have complementary strengths—and the taxonomy is a conceptual starting point, not a benchmark result. The authors don't claim to have measured anything, and they cite the supporting literature for the broad strokes. I'd have liked a sentence admitting that the matrix is a snapshot that needs periodic updating, but it's not the kind of omission that undermines the thesis.\n\nOne more thing: the paper self-cites heavily, but the self-citations are the standard works (Hogan's KG survey, Wikidata, Head-to-Tail, CRAG), and they back the claims. No problem there.\n\nWho it's for: researchers building hybrid QA systems, benchmark designers, and anyone thinking about user information needs. It's a good reading-group discussion piece. I'd support citing it as a reference for the taxonomy.\n\nRecommendation: yes, send it to peer review. A serious referee can help tighten the framing and maybe ask for a note on the matrix's validity, but the paper deserves referee time rather than a desk reject.","headline":"A well-executed position paper that gives the community a useful taxonomy and roadmap for combining SEs, KGs, and LLMs; the capability matrix is asserted rather than measured, but that fits the genre.","tokens_in":25260,"tokens_out":2926,"would_cite":true,"duration_ms":28215,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that search engines, knowledge graphs, and large language models answer different pieces of the user's question, and that combining them—not choosing among them—is the research direction that will serve users best.","keywords":["knowledge graphs","large language models","search engines","question answering","user information needs","taxonomy","retrieval-augmented generation","hybrid systems"],"falsifier":"Collect a sample of real user queries, label each with one of the taxonomy's twelve subclasses, and have a strong search engine, a knowledge-graph question-answering system, and a retrieval-augmented LLM answer each; then compare which system produces the highest-quality answer per subclass. If the technology that the paper's Table 2 predicts as best is not the actual best for a majority of subclasses, the central complementarity mapping fails.","tokens_in":24300,"feed_emoji":"🧩","tokens_out":6607,"duration_ms":56110,"temperature":0.7,"pith_summary":"This paper argues that search engines, knowledge graphs, and large language models are not competing ways to answer questions but complementary ones: each is strong exactly where the others are weak. To support this, it builds a taxonomy of user information needs—facts, explanations, planning, and advice, each with subtypes—and scores every technology against every need. It finds that no single technology covers the full spectrum: knowledge graphs excel at precise multi-hop and analytical factual queries, search engines bring broad, fresh, documented coverage, and language models synthesize, explain, and converse. The paper concludes that research should aim at hybrid systems, moving from simple augmentation to full amalgamation of the three. If the argument holds, betting the future on a single dominant technology would be a mistake.","feed_headline":"Search engines, knowledge graphs, and LLMs are complements, not rivals","feed_subtitle":"A taxonomy of user information needs shows each technology fails alone and hybrids are the path forward.","key_machinery":"The engine of the paper is the paired analytical apparatus of Table 1 and Table 2. Table 1 compares search engines, knowledge graphs, and LLMs along a dozen capability dimensions (correctness, coverage, completeness, freshness, generation, synthesis, transparency, coherency, refinability, fairness, usability, expressivity, efficiency, multilingualism, personalization). Table 2 is the user-needs taxonomy: a structured partition of questions into four main classes and twelve subclasses, each annotated with example questions and a strength/weakness assessment for each technology. The argument runs through the mapping between these two tables—each technology's capability profile predicts which classes of user questions it can answer well—and the roadmap in Section 4 then proposes how to combine the technologies so that the mapping covers more of the taxonomy.","core_discovery":"The paper's central claim is that the technologies are complementary rather than competitive, with the complementarity shown by a systematic comparison across dimensions such as correctness, coverage, completeness, freshness, generation, synthesis, transparency, and expressivity. Its key move is to view these dimensions through the lens of the user's information need: a taxonomy that divides questions into facts (popular, long-tail, dynamic, multi-hop, analytical), explanations (commonsense, causal, exploratory), planning (instructive, recommendation, spatio-temporal), and advice (lifestyle, cultural, philosophical). Against this taxonomy the paper assigns each technology clear strengths and weaknesses—KGs excel at complex factual queries but lack nuance and are hard to use; SEs are broad and fresh but cannot synthesize across documents; LLMs are flexible and generative but hallucinate, lag on long-tail facts, and are opaque. The paper argues that where one is weak, another is often strong, and it lays out a four-phase roadmap—augmentation, ensemble, federation, amalgamation—for combining all three to emphasize the pros and suppress the cons.","pith_inferences":["The taxonomy could double as an evaluation harness: a benchmark that samples each of the twelve subclasses from real query logs would test whether the predicted best technology for each category actually wins, something the paper does not do.","Real user questions are often hybrids of the categories (e.g., dynamic plus multi-hop, or analytical plus advice), so a practical system will need to decompose a query into sub-tasks rather than assign it to a single box.","The capability scores are a snapshot: as LLMs gain tool use, longer context, and live retrieval, some assignments (notably multi-hop and analytical) may shift, so the durable contribution is the framework, not the current cell values.","A transparency-aware design would follow from the authors' own emphasis: for high-stakes factual answers, a hybrid that surfaces KG provenance or SE sources alongside LLM synthesis may earn more user trust than the most fluent single model."],"forward_implications":["If the paper is right, question-answering systems should be designed as hybrids that route each query to the technology best suited to its category, rather than relying on a single model or engine.","Knowledge graphs will remain load-bearing for multi-hop and analytical factual queries even as LLMs improve, because no amount of latent text statistics substitutes for explicit join and aggregation operators.","Retrieval augmentation (SE for LLM) is a necessary but not sufficient fix: the long-tail failure persists with RAG, so structured knowledge must be part of the answer.","Benchmarks for question answering should be built from the taxonomy's categories, so that progress is measured across the full spectrum of user needs, not just popular factual questions.","The eventual amalgamation goal implies that research on combined representations and aligned tokens (tying the LLM's textual 'Turing award' to the KG's entity and the SE's index term) is a concrete, high-value direction."],"supporting_citations":[{"why":"The foundational retrieval-augmented generation method that the paper builds its 'SE for LLM' discussion on.","marker":"[24]"},{"why":"A prior comparison of LLMs and knowledge graphs that this paper extends by adding search engines and the user perspective.","marker":"[31]"},{"why":"The work that initiated the 'LLM for KG' knowledge-generation direction and the critique of language models as knowledge bases.","marker":"[34]"},{"why":"Evidence that LLM parametric recall decays on rare entities, supporting the need for KGs and SEs for long-tail queries.","marker":"[21]"},{"why":"Evidence on when LLM parametric memory fails, motivating retrieval augmentation and hybrid design.","marker":"[27]"},{"why":"The opposing position that language models make knowledge graphs obsolete, which the paper argues against.","marker":"[43]"},{"why":"Evidence of LLM staleness, grounding the freshness limitation in the comparison.","marker":"[48]"},{"why":"An existing benchmark that categorizes user questions, which the taxonomy builds on and contrasts with.","marker":"[54]"}],"fun_headline_variants":["LLMs, KGs, and search engines: complementary by design","Why your questions need all three: LLMs, KGs, and search engines","Combining LLMs, knowledge graphs, and search: the user-first path","A roadmap to answer users with LLMs, knowledge graphs, and search"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that its taxonomy of user information needs matches how real questions actually divide, and that the capability scores it assigns to each technology still hold once the technologies are combined.","fun_headline_variants_meta":{"raw":{"variants":["LLMs, KGs, and search engines: complementary by design","Why your questions need all three: LLMs, KGs, and search engines","Combining LLMs, knowledge graphs, and search: the user-first path","A roadmap to answer users with LLMs, knowledge graphs, and search"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000279,"raw_usage":{"total_tokens":1609,"prompt_tokens":850,"completion_tokens":759,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":678}},"tokens_in":466,"tokens_out":759,"duration_ms":7061,"temperature":1.0,"reasoning_tokens":678,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:54:12.528343+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a sample of real user queries, label each with one of the taxonomy's twelve subclasses, and have a strong search engine, a knowledge-graph question-answering system, and a retrieval-augmented LLM answer each; then compare which system produces the highest-quality answer per subclass. If the technology that the paper's Table 2 predicts as best is not the actual best for a majority of subclasses, the central complementarity mapping fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Evidence that LLM parametric recall decays on rare entities, supporting the need for KGs and SEs for long-tail queries."},{"cited_title":"Language Models sounds the Death Knell of Knowledge Graphs","cited_arxiv_id":"2301.03980","evidence_quote":"The opposing position that language models make knowledge graphs obsolete, which the paper argues against."}],"review_version":1}