Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Answering Complex Questions by Joining Multi-Document Evidence with Quasi Knowledge Graphs

T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read QUEST builds a noisy graph from facts extracted across web pages and uses group Steiner trees to join evidence from different documents into direct answers for complex questions.

desk verdict Neat unsupervised QA method with a real benchmark-independence problem and a fixable weight-normalization gap; worth peer review if the authors document the cost conversion. read the letter →

arxiv 1908.00469 v4 pith:P5RVOYW4 submitted 2019-08-01 cs.IR

classification cs.IR
keywords questionsanswerscomplexquestansweringdocumentsevidencegraphs
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

QUEST tackles questions that can only be answered by combining clues from several web pages, such as asking which footballer played in two different finals and was born in Africa. Many QA systems look for one passage that contains all the clues, but such a passage often does not exist. QUEST instead extracts simple subject-predicate-object facts from each retrieved page, such as "Umtiti played for France" or "Matuidi is a footballer". These facts become nodes and edges in a noisy graph, where nodes are entity names, relation phrases, and type labels, and edges connect pieces that belong together.

The question is converted into a set of "cornerstone" nodes, one group per important word or phrase from the question. QUEST then looks for the cheapest tree in the graph that touches at least one node from every group. This is a group Steiner tree problem. The non-cornerstone nodes inside these trees are candidate answers. Because the answer has to connect every part of the question, the tree acts as a joint check that the pieces are about the same person or thing. Finally, candidates are filtered by expected answer type, duplicate surface forms are merged, and the answers are ranked mainly by how often they appear in low-cost trees.

The key design choice is to accept noise in the extracted facts instead of trying to clean them up, and to let the Steiner tree computation sort out which connections are consistent. The authors show that this unsupervised pipeline beats a reading-comprehension baseline, DrQA, on two complex-question benchmarks, but the absolute accuracy is modest and the ranking stage is the biggest source of errors.

Extended reading notes

Core claim

The central claim is stated in the abstract and Section 6: QUEST "can answer complex questions directly from textual sources on-the-fly, by computing similarity joins over partial results from different documents" and "significantly and consistently outperforms the neural baseline DrQA, and other graph-based methods" on the CQ-W and CQ-T benchmarks (best MRR 0.355 and 0.467 on top-10 corpora). If true, an unsupervised system can answer multi-document complex questions without training data by connecting question terms to a noisy quasi KG and computing top-k group Steiner trees.

Load-bearing premise

The load-bearing premise is that the similarity threshold and embedding matching used to pick cornerstones and alignment edges preserve the semantic structure of the question, so that a low-cost group Steiner tree connecting one node per question group will pass through a correct answer (Section 4.1). This premise is not derived from theory; it is an empirical modeling choice. The error analysis in Section 7 shows wrong cornerstone senses and spurious alignments are among the leading failure causes, confirming the premise is fragile.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents QUEST, an unsupervised system for answering multi-document fact-centric complex questions. QUEST retrieves a small document pool for each question, extracts SPO triples with a custom proximity-scored Open IE extractor, and assembles a noisy 'quasi KG' of entity, relation, and type nodes with alignment and type edges. Cornerstones are nodes whose labels match question words; candidate answers are non-terminal nodes of top-k Group Steiner Trees connecting at least one cornerstone from each question-token group. Candidates are filtered by lexical answer type, aggregated by string and alignment redundancy, and ranked by a tree-cost-weighted sum over the trees in which they occur. Experiments on two complex-question collections (CQ-W and CQ-T) compare QUEST with DrQA, BFS, and ShortestPaths under top-10 and stratified corpora, and report consistent MRR gains with ablations, error analyses, and parameter sensitivity.

Significance. QUEST is a coherent and clearly explained unsupervised alternative to neural reading-comprehension QA for a genuinely difficult setting: combining evidence across documents for questions with multiple entities and relations. Its main strengths are the fully unsupervised pipeline, the explicit modeling of multi-document evidence via the quasi KG, the use of Group Steiner Trees for joint disambiguation, and the unusually complete empirical package: two benchmarks, multiple corpus-quality strata, ablations, ranking variants, error analysis, and public data/code/demo. If the comparative results are reproduced, the system is a useful contribution to text-based QA. The main reservations are that both benchmarks were created by the authors' group, the neural baseline set is narrow, and the written description of the edge-weight normalization contains a gap that affects the theoretical claim of optimality.

major comments (3)
  1. [Sec. 3.1, Sec. 3.2, Sec. 4.1] Section 3.1 defines the triple-edge confidence as a sum over sentences of 1/d(S_i) with no stated upper bound, and Section 3.2 says these confidence scores are used as triple-edge weights. Section 4.1 then assumes '[0,1]-normalized similarity-score weights' and defines edge cost as 1 - score, while explicitly requiring w_ij >= 0. If raw triple scores are used, costs can be negative, violating the non-negative-weight condition of the GST dynamic program and invalidating the claim that the reported trees are minimum-cost trees. If a normalization step exists in the implementation, it is not documented. Please add the exact normalization formula (or define triple-edge scores with a bounded range) and verify that the GST implementation never receives negative costs; without this, the central claim that QUEST's advantage comes from cost-optimal Group Steiner Trees is not fully supported by the paper as written.
  2. [Sec. 5.2 and Table 2] CQ-W is adopted from a prior paper by the authors' group and CQ-T is constructed by the authors for this paper; the only neural baseline is DrQA. The abstract and Section 6 say QUEST 'substantially outperforms state-of-the-art baselines,' but the evidence supports a narrower claim: superiority over DrQA and two graph heuristics on two self-created collections. Please either add an independent evaluation (for example, on a third-party complex multi-document QA benchmark) or rewrite the abstract and Section 6 to state the comparison precisely. This distinction is important because both the multi-document setting and the benchmark construction are controlled by the authors.
  3. [Sec. 5.3 and Sec. 7 (Fig. 2)] Section 5.3 states that the three thresholds are set to 0.5 and that 'no tuning is involved,' yet Section 7 and Figure 2 report a sensitivity analysis over precisely these thresholds on CQ-W. If the default 0.5 was selected after inspecting results on CQ-W, the comparison in Table 2 is not tuning-free; if it was a priori, the 'no tuning' claim needs an explicit statement. Please specify how the default thresholds were chosen and, ideally, evaluate threshold choices on a development split distinct from the test questions, because the main comparative results depend on these free parameters.
minor comments (6)
  1. [Sec. 3.1] Please present the proximity-based scoring formula as an equation rather than prose; the current text should also state whether distances are normalized by sentence length and how repeated co-occurrences in the same document are aggregated.
  2. [Sec. 4.1] The text alternates between 'word or phrase' and 'per token of the question' for cornerstones; please clarify how multi-word question phrases map to terminal groups in the GST formulation.
  3. [Sec. 5.4] Please specify DrQA's exact configuration, including the retriever index type, the reader checkpoint used, the number of retrieved passages, and the maximum passage length, since DrQA's performance depends strongly on these choices.
  4. [Table 3] Please define the notation 'A' and explain the binning in the caption; the upper and lower halves of the table are hard to parse without a precise description of how edge contributions from distinct documents are counted.
  5. [Table 5] The error-scenario percentages sum to 101% for CQ-W and 99% for CQ-T; please specify whether the categories are mutually exclusive, clarify the denominator, and account for the rounding.
  6. [Sec. 6 and Table 2] Statistical significance is reported only for QUEST over DrQA; please state whether the comparisons against BFS and ShortestPaths were tested, and if so, report those p-values or note that no significance claim is made for them.
Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claim rests on a small set of modeling assumptions and hand-set parameters. The free parameters are thresholds and the top-k count; they are not learned from a fitting procedure, but k=50 is chosen by observing benchmark performance. The most important axiom is the GST connection premise: answers are assumed to lie on low-cost trees connecting question terms. The extractor-and-retrieval axiom is empirically bounded by the paper's own coverage analysis.

free parameters (5)
  • Cornerstone selection threshold = 0.5
    Similarity threshold for matching question tokens to nodes in the quasi KG; declared fixed with no tuning, but it controls which terminals enter the GST and directly affects whether answers can be found.
  • Alignment edge insertion threshold = 0.5
    Jaccard or cosine similarity threshold for adding entity and relation alignment edges; lower thresholds add noisy synonym edges and higher thresholds fragment the graph.
  • Answer merging threshold = 0.5
    Similarity threshold for aggregating duplicate surface-form answers after extraction.
  • Number of top-k GSTs = 50
    Number of Steiner trees used for candidate answers; chosen because results plateau beyond 50 on CQ-W, so it is effectively a tuned parameter.
  • Answer type pruning threshold = not reported
    Cosine similarity threshold above which candidates survive lexical type checking; the value is not reported in the paper.
assumptions (5)
  • standard math Top-k Group Steiner Trees can be computed optimally by the Ding et al. dynamic programming algorithm with O(n log n) complexity in graph size.
    Section 4.1 relies on this external algorithmic result for exactness and tractability.
  • ad hoc to paper A correct answer is a non-terminal node lying on a low-cost tree that connects at least one node from each group of question-cornerstones.
    Section 4.1 defines the GST objective; this is the core modeling premise of the method and is not derived from any prior theory.
  • domain assumption Token-level similarity from word2vec and AIDA dictionaries correctly identifies cornerstones, alignment edges, and answer types despite no disambiguation.
    Sections 3.2 and 4.2 assume these soft matches are sufficiently reliable; the paper's error analysis shows wrong cornerstones cause failures.
  • domain assumption The custom pattern-based Open IE extractor yields SPO triples that preserve the facts needed to answer the question (85.2% on CQ-W, 82.3% on CQ-T).
    Section 7 reports this upper bound; if facts are missing from the quasi KG, no GST can retrieve them.
  • domain assumption The corpus retrieved by the search engine contains the answer and the connecting evidence across documents.
    Section 5.3 assumes top-10 and stratified search results suffice; the paper verifies this indirectly but answer-not-in-corpus errors are 1% on CQ-W and 7% on CQ-T.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Answering Complex Questions by Joining Multi-Document Evidence with Quasi Knowledge Graphs." pith.science (2026). https://pith.science/paper/P5RVOYW4

@misc{pith2026190800469,
  author       = {Pith},
  title        = {Pith review of: Answering Complex Questions by Joining Multi-Document Evidence with Quasi Knowledge Graphs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P5RVOYW4}},
  note         = {Machine review of arXiv:1908.00469}
}
read the original abstract

Direct answering of questions that involve multiple entities and relations is a challenge for text-based QA. This problem is most pronounced when answers can be found only by joining evidence from multiple documents. Curated knowledge graphs (KGs) may yield good answers, but are limited by their inherent incompleteness and potential staleness. This paper presents QUEST, a method that can answer complex questions directly from textual sources on-the-fly, by computing similarity joins over partial results from different documents. Our method is completely unsupervised, avoiding training-data bottlenecks and being able to cope with rapidly evolving ad hoc topics and formulation style in user questions. QUEST builds a noisy quasi KG with node and edge weights, consisting of dynamically retrieved entity names and relational phrases. It augments this graph with types and semantic alignments, and computes the best answers by an algorithm for Group Steiner Trees. We evaluate QUEST on benchmarks of complex questions, and show that it substantially outperforms state-of-the-art baselines.

Figures

Figures reproduced from arXiv: 1908.00469 by the authors.

Figure 1
Figure 1. A quasi KG for our running example question. into simpler sub-questions requires syntactic decomposition pat￾terns that break with ungrammatical constructs, and also a way of stitching sub-results together. Other than for special cases like temporal modifiers in questions, this is beyond the scope of today’s systems. (iii) Modern approaches that leverage deep learning criti￾cally rely on training data, which is not … view at source ↗
Figure 2
Figure 2. Robustness of QUEST to various system configuration parameters, on CQ-W with top-10 corpora (similar trends on CQ-T). 1999 to 2007, recently revived it as the LiveQA track. IBM Wat￾son [26] extended this paradigm by combining it with learned models for special question types. QA over KGs. The advent of large knowledge graphs like Free￾base [9], YAGO [54], DBpedia [4] and Wikidata [64] has given rise to QA over KGs (… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Semantic Document Retrieval- Employing Group Steiner Tree Algorithm with Domain Knowledge Enrichment

    cs.IR 2025-08 reject novelty 4.0 of 10

    SemDR applies a Group Steiner Tree heuristic to a domain-enriched concept graph for agriculture document retrieval and reports large gains over Lucene, ElasticSearch, and Doc2Vec.

Reference graph

Works this paper leans on

78 extracted references · 75 canonical work pages · cited by 1 Pith paper

  1. [1]

    Abujabal, M

    A. Abujabal, M. Yahya, M. Riedewald, and G. Weikum. 2017. Automated template generation for question answering over knowledge graphs. In WWW

  2. [2]

    Agichtein, D

    E. Agichtein, D. Carmel, D. Pelleg, Y. Pinter, and D. Harman. 2015. Overview of the TREC 2015 LiveQA Track. In TREC

  3. [3]

    Angeli, M

    G. Angeli, M. J. J. Premkumar, and C. D. Manning. 2015. Leveraging linguistic structure for open domain information extraction. In ACL

  4. [4]

    S. Auer, C. Bizer, G. Kobilarov, J. Lehmann, R. Cyganiak, and Z. Ives. 2007. DBpedia: A nucleus for a Web of open data. In The Semantic Web. Springer

  5. [5]

    J. Bao, N. Duan, Z. Yan, M. Zhou, and T. Zhao. 2016. Constraint-Based Question Answering with Knowledge Graph. In COLING

  6. [6]

    Bast and E

    H. Bast and E. Haussmann. 2015. More accurate question answering on Freebase. In CIKM

  7. [7]

    Berant, A

    J. Berant, A. Chou, R. Frostig, and P. Liang. 2013. Semantic Parsing on Freebase from Question-Answer Pairs. In ACL

  8. [8]

    Bhalotia, A

    G. Bhalotia, A. Hulgeri, C. Nakhe, S. Chakrabarti, and S. Sudarshan. 2002. Keyword searching and browsing in databases using BANKS. In ICDE

Show all 78 references
  1. [9]

    Bollacker, C

    K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor. 2008. Freebase: A collaboratively created graph database for structuring human knowledge. In SIGMOD

  2. [10]

    Bordes, N

    A. Bordes, N. Usunier, S. Chopra, and J. Weston. 2015. Large-scale simple question answering with memory networks. In arXiv

  3. [11]

    Cai and A

    Q. Cai and A. Yates. 2013. Large-scale Semantic Parsing via Schema Matching and Lexicon Extension. In ACL

  4. [12]

    Camille Chanial, Rédouane Dziri, Helena Galhardas, Julien Leblay, Minh- Huong Le Nguyen, and Ioana Manolescu. 2018. Connectionlens: finding connec- tions across heterogeneous data sources. In VLDB

  5. [13]

    D. Chen, A. Fisch, J. Weston, and A. Bordes. 2017. Reading Wikipedia to Answer Open-Domain Questions. In ACL

  6. [14]

    Coffman and A

    J. Coffman and A. C. Weaver. 2014. An empirical performance evaluation of relational keyword search techniques. In TKDE

  7. [15]

    Del Corro and R

    L. Del Corro and R. Gemulla. 2013. ClausIE: Clause-based open information extraction. In WWW

  8. [16]

    R. Das, M. Zaheer, S. Reddy, and A. McCallum. 2017. Question Answering on Knowledge Bases and Text using Universal Schema and Memory Networks. In ACL

  9. [17]

    Dehghani, H

    M. Dehghani, H. Azarbonyad, J. Kamps, and M. de Rijke. 2019. Learning to Transform, Combine, and Reason in Open-Domain Question Answering. In WSDM

  10. [18]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv (2018)

  11. [19]

    Diefenbach, V

    D. Diefenbach, V. López, K. Deep Singh, and P. Maret. 2018. Core techniques of question answering systems over knowledge bases: A survey. Knowl. Inf. Syst. 55, 3 (2018)

  12. [20]

    Dietz and B

    L. Dietz and B. Gamari. 2017. TREC CAR: A Data Set for Complex Answer Retrieval. In TREC

  13. [21]

    B. Ding, J. X. Yu, S. Wang, L. Qin, X. Zhang, and X. Lin. 2007. Finding top- 𝑘 min-cost connected trees in databases. In ICDE

  14. [22]

    L. Dong, F. Wei, M. Zhou, and K. Xu. 2015. Question Answering over Freebase with Multi-Column Convolutional Neural Networks. In ACL

  15. [23]

    R. G. Downey and M. R. Fellows. 2013.Fundamentals of Parameterized Complexity. Springer

  16. [24]

    Fader, L

    A. Fader, L. Zettlemoyer, and O. Etzioni. 2013. Paraphrase-driven learning for open question answering. In ACL

  17. [25]

    Fader, L

    A. Fader, L. Zettlemoyer, and O. Etzioni. 2014. Open question answering over curated and extracted knowledge bases. In KDD

  18. [26]

    Ferrucci et al

    D. Ferrucci et al. 2012. This is Watson. IBM Journal Special Issue 56, 3 (2012)

  19. [27]

    N. Garg, G. Konjevod, and R. Ravi. 2000. A Polylogarithmic Approximation Algorithm for the Group Steiner Tree Problem. J. Algorithms 37, 1 (2000)

  20. [28]

    Gashteovski, R

    K. Gashteovski, R. Gemulla, and L. Del Corro. 2017. MinIE: Minimizing Facts in Open Information Extraction. In EMNLP

  21. [29]

    Grycner and G

    A. Grycner and G. Weikum. 2016. POLY: Mining Relational Paraphrases from Multilingual Sentences. In EMNLP

  22. [30]

    Ido Guy. 2018. The Characteristics of Voice Search: Comparing Spoken with Typed-in Mobile Web Search Queries. ACM Trans. Inf. Syst. (2018)

  23. [31]

    M. A. Hearst. 1992. Automatic Acquisition of Hyponyms from Large Text Corpora. In COLING

  24. [32]

    Hoffart, M

    J. Hoffart, M. A. Yosef, I. Bordino, H. Fürstenau, M. Pinkal, M. Spaniol, B. Taneva, S. Thater, and G. Weikum. 2011. Robust disambiguation of named entities in text. In EMNLP

  25. [33]

    S. Hu, L. Zou, J. X. Yu, H. Wang, and D. Zhao. 2018. Answering Natural Language Questions by Subgraph Matching over Knowledge Graphs. Trans. on Know. and Data Eng. 30, 5 (2018)

  26. [34]

    Iyyer, J

    M. Iyyer, J. L. Boyd-Graber, L. M. B. Claudino, R. Socher, and H. Daumé III. 2014. A Neural Network for Factoid Question Answering over Paragraphs. In EMNLP

  27. [35]

    Joshi, E

    M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In ACL

  28. [36]

    Joshi, U

    M. Joshi, U. Sawant, and S. Chakrabarti. 2014. Knowledge graph and corpus driven segmentation and answer inference for telegraphic entity-seeking queries. In EMNLP

  29. [37]

    Kacholia, S

    V. Kacholia, S. Pandit, S. Chakrabarti, S. Sudarshan, R. Desai, and H. Karambelkar

  30. [38]

    Kadry and L

    A. Kadry and L. Dietz. 2017. Open Relation Extraction for Support Passage Retrieval: Merit and Open Issues. In SIGIR

  31. [39]

    Kasneci, M

    G. Kasneci, M. Ramanath, M. Sozio, F. M. Suchanek, and G. Weikum. 2009. STAR: Steiner-tree approximation in relationship graphs. In ICDE

  32. [40]

    T. Khot, A. Sabharwal, and P. Clark. 2017. Answering complex questions using open information extraction. In ACL

  33. [41]

    C. Kwok, O. Etzioni, and D. S. Weld. 2001. Scaling Question Answering to the Web. ACM Trans. Inf. Syst. (2001)

  34. [42]

    R. Li, L. Qin, J. X. Yu, and R. Mao. 2016. Efficient and progressive group steiner tree search. In SIGMOD

  35. [43]

    Y. Lin, H. Ji, Z. Liu, and M. Sun. 2018. Denoising distantly supervised open-domain question answering. In ACL

  36. [44]

    Manning, M

    C. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. Bethard, and D. McClosky. 2014. The Stanford CoreNLP natural language processing toolkit. In ACL

  37. [45]

    Mausam. 2016. Open information extraction systems and downstream applica- tions. In IJCAI

  38. [46]

    Mikolov, I

    T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. 2013. Distributed Representations of Words and Phrases and their Compositionality. In NIPS

  39. [47]

    Pavlick, P

    E. Pavlick, P. Rastogi, J. Ganitkevitch, B. Van Durme, and C. Callison-Burch. 2015. PPDB 2.0: Better paraphrase ranking, fine-grained entailment relations, word embeddings, and style classification. In ACL

  40. [48]

    Pennington, R

    J. Pennington, R. Socher, and C. D. Manning. 2014. GloVe: Global Vectors for Word Representation. In EMNLP

  41. [49]

    Rajpurkar, J

    P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. In EMNLP

  42. [50]

    Ravichandran and E

    D. Ravichandran and E. Hovy. 2002. Learning Surface Text Patterns for a Question Answering System. In ACL

  43. [51]

    Savenkov and E

    D. Savenkov and E. Agichtein. 2016. When a knowledge base is not enough: Question answering over knowledge bases with external text data. In SIGIR

  44. [52]

    Sawant and S

    U. Sawant and S. Chakrabarti. 2013. Learning joint query interpretation and response ranking. In WWW

  45. [53]

    V. I. Spitkovsky and A. X. Chang. 2012. A Cross-Lingual Dictionary for English Wikipedia Concepts. In LREC. 3168–3175

  46. [54]

    F. M. Suchanek, G. Kasneci, and G. Weikum. 2007. YAGO: A core of semantic knowledge. In WWW

  47. [55]

    H. Sun, B. Dhingra, M. Zaheer, K. Mazaitis, R. Salakhutdinov, and W. W. Cohen

  48. [56]

    H. Sun, H. Ma, W. Yih, C. Tsai, J. Liu, and M. Chang. 2015. Open domain question answering via semantic enrichment. In WWW

  49. [57]

    Talmor and J

    A. Talmor and J. Berant. 2018. The Web as a Knowledge-base for Answering Complex Questions. In NAACL-HLT

  50. [58]

    C. Tan, F. Wei, Q. Zhou, N. Yang, B. Du, W. Lv, and M. Zhou. 2018. Context-Aware Answer Sentence Selection With Hierarchical Gated Recurrent Neural Networks. IEEE/ACM Trans. Audio, Speech & Language Processing 26, 3 (2018)

  51. [59]

    Unger, L

    C. Unger, L. Bühmann, J. Lehmann, A. N. Ngomo, D. Gerber, and P. Cimiano

  52. [60]

    Unger, A

    C. Unger, A. Freitas, and P. Cimiano. 2014. An Introduction to Question Answer- ing over Linked Data. In Reasoning Web

  53. [61]

    Usbeck, A

    R. Usbeck, A. N. Ngomo, B. Haarmann, A. Krithara, M. Röder, and G. Napolitano

  54. [62]

    E. M. Voorhees. 2014. The Effect of Sampling Strategy on Inferred Measures. In SIGIR

  55. [63]

    E. M. Voorhees and D. K. Harman. 2005. TREC: Experiment and evaluation in information retrieval. MIT press Cambridge

  56. [64]

    Vrandečić and M

    D. Vrandečić and M. Krötzsch. 2014. Wikidata: A free collaborative knowledge base. Commun. ACM 57, 10 (2014)

  57. [65]

    R. W. White, M. Richardson, and W. T. Yih. 2015. Questions vs. queries in informational search tasks. In WWW. 135–136

  58. [66]

    Wieting, M

    J. Wieting, M. Bansal, K. Gimpel, and K. Livescu. 2016. Towards universal para- phrastic sentence embeddings. In ICLR

  59. [67]

    K. Xu, S. Reddy, Y. Feng, S. Huang, and D. Zhao. 2016. Question answering on freebase via relation extraction and textual evidence. In ACL

  60. [68]

    Yahya, K

    M. Yahya, K. Berberich, S. Elbassuoni, M. Ramanath, V. Tresp, and G. Weikum

  61. [69]

    Yahya, S

    M. Yahya, S. Whang, R. Gupta, and A. Halevy. 2014. ReNoun: Fact extraction for nominal attributes. In EMNLP

  62. [70]

    W. Yih, M. Chang, X. He, and J. Gao. 2015. Semantic Parsing via Staged Query Graph Generation: Question Answering with Knowledge Base. In ACL

  63. [71]

    P. Yin, N. Duan, B. Kao, J. Bao, and M. Zhou. 2015. Answering Questions with Complex Semantic Constraints on Open Knowledge Bases. In CIKM

  64. [72]

    J. X. Yu, L. Qin, and L. Chang. 2009. Keyword Search in Databases. M & C

  65. [73]

    In EMNLP

    Natural Language Questions for the Web of Data. In EMNLP

  66. [78]

    Ziegler, A

    D. Ziegler, A. Abujabal, R. Saha Roy, and G. Weikum. 2017. Efficiency-aware An- swering of Compositional Questions using Answer Type Prediction. In IJCNLP.10

  67. [2005]

    Bidirectional expansion for keyword search on graph databases. In VLDB

  68. [2012]

    Template-based question answering over RDF data. In WWW

  69. [2017]

    7th Open Challenge on Question Answering over Linked Data. In QALD

  70. [2018]

    In EMNLP

    Open domain question answering using early fusion of knowledge bases and text. In EMNLP

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.