REVIEW 3 major objections 6 minor 1 cited by
Answering Complex Questions by Joining Multi-Document Evidence with Quasi Knowledge Graphs
T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read QUEST builds a noisy graph from facts extracted across web pages and uses group Steiner trees to join evidence from different documents into direct answers for complex questions.
desk verdict Neat unsupervised QA method with a real benchmark-independence problem and a fixable weight-normalization gap; worth peer review if the authors document the cost conversion. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
The question is converted into a set of "cornerstone" nodes, one group per important word or phrase from the question. QUEST then looks for the cheapest tree in the graph that touches at least one node from every group. This is a group Steiner tree problem. The non-cornerstone nodes inside these trees are candidate answers. Because the answer has to connect every part of the question, the tree acts as a joint check that the pieces are about the same person or thing. Finally, candidates are filtered by expected answer type, duplicate surface forms are merged, and the answers are ranked mainly by how often they appear in low-cost trees.
The key design choice is to accept noise in the extracted facts instead of trying to clean them up, and to let the Steiner tree computation sort out which connections are consistent. The authors show that this unsupervised pipeline beats a reading-comprehension baseline, DrQA, on two complex-question benchmarks, but the absolute accuracy is modest and the ranking stage is the biggest source of errors.
Extended reading notes
Core claim
The central claim is stated in the abstract and Section 6: QUEST "can answer complex questions directly from textual sources on-the-fly, by computing similarity joins over partial results from different documents" and "significantly and consistently outperforms the neural baseline DrQA, and other graph-based methods" on the CQ-W and CQ-T benchmarks (best MRR 0.355 and 0.467 on top-10 corpora). If true, an unsupervised system can answer multi-document complex questions without training data by connecting question terms to a noisy quasi KG and computing top-k group Steiner trees.
Load-bearing premise
The load-bearing premise is that the similarity threshold and embedding matching used to pick cornerstones and alignment edges preserve the semantic structure of the question, so that a low-cost group Steiner tree connecting one node per question group will pass through a correct answer (Section 4.1). This premise is not derived from theory; it is an empirical modeling choice. The error analysis in Section 7 shows wrong cornerstone senses and spurious alignments are among the leading failure causes, confirming the premise is fragile.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents QUEST, an unsupervised system for answering multi-document fact-centric complex questions. QUEST retrieves a small document pool for each question, extracts SPO triples with a custom proximity-scored Open IE extractor, and assembles a noisy 'quasi KG' of entity, relation, and type nodes with alignment and type edges. Cornerstones are nodes whose labels match question words; candidate answers are non-terminal nodes of top-k Group Steiner Trees connecting at least one cornerstone from each question-token group. Candidates are filtered by lexical answer type, aggregated by string and alignment redundancy, and ranked by a tree-cost-weighted sum over the trees in which they occur. Experiments on two complex-question collections (CQ-W and CQ-T) compare QUEST with DrQA, BFS, and ShortestPaths under top-10 and stratified corpora, and report consistent MRR gains with ablations, error analyses, and parameter sensitivity.
Significance. QUEST is a coherent and clearly explained unsupervised alternative to neural reading-comprehension QA for a genuinely difficult setting: combining evidence across documents for questions with multiple entities and relations. Its main strengths are the fully unsupervised pipeline, the explicit modeling of multi-document evidence via the quasi KG, the use of Group Steiner Trees for joint disambiguation, and the unusually complete empirical package: two benchmarks, multiple corpus-quality strata, ablations, ranking variants, error analysis, and public data/code/demo. If the comparative results are reproduced, the system is a useful contribution to text-based QA. The main reservations are that both benchmarks were created by the authors' group, the neural baseline set is narrow, and the written description of the edge-weight normalization contains a gap that affects the theoretical claim of optimality.
major comments (3)
- [Sec. 3.1, Sec. 3.2, Sec. 4.1] Section 3.1 defines the triple-edge confidence as a sum over sentences of 1/d(S_i) with no stated upper bound, and Section 3.2 says these confidence scores are used as triple-edge weights. Section 4.1 then assumes '[0,1]-normalized similarity-score weights' and defines edge cost as 1 - score, while explicitly requiring w_ij >= 0. If raw triple scores are used, costs can be negative, violating the non-negative-weight condition of the GST dynamic program and invalidating the claim that the reported trees are minimum-cost trees. If a normalization step exists in the implementation, it is not documented. Please add the exact normalization formula (or define triple-edge scores with a bounded range) and verify that the GST implementation never receives negative costs; without this, the central claim that QUEST's advantage comes from cost-optimal Group Steiner Trees is not fully supported by the paper as written.
- [Sec. 5.2 and Table 2] CQ-W is adopted from a prior paper by the authors' group and CQ-T is constructed by the authors for this paper; the only neural baseline is DrQA. The abstract and Section 6 say QUEST 'substantially outperforms state-of-the-art baselines,' but the evidence supports a narrower claim: superiority over DrQA and two graph heuristics on two self-created collections. Please either add an independent evaluation (for example, on a third-party complex multi-document QA benchmark) or rewrite the abstract and Section 6 to state the comparison precisely. This distinction is important because both the multi-document setting and the benchmark construction are controlled by the authors.
- [Sec. 5.3 and Sec. 7 (Fig. 2)] Section 5.3 states that the three thresholds are set to 0.5 and that 'no tuning is involved,' yet Section 7 and Figure 2 report a sensitivity analysis over precisely these thresholds on CQ-W. If the default 0.5 was selected after inspecting results on CQ-W, the comparison in Table 2 is not tuning-free; if it was a priori, the 'no tuning' claim needs an explicit statement. Please specify how the default thresholds were chosen and, ideally, evaluate threshold choices on a development split distinct from the test questions, because the main comparative results depend on these free parameters.
minor comments (6)
- [Sec. 3.1] Please present the proximity-based scoring formula as an equation rather than prose; the current text should also state whether distances are normalized by sentence length and how repeated co-occurrences in the same document are aggregated.
- [Sec. 4.1] The text alternates between 'word or phrase' and 'per token of the question' for cornerstones; please clarify how multi-word question phrases map to terminal groups in the GST formulation.
- [Sec. 5.4] Please specify DrQA's exact configuration, including the retriever index type, the reader checkpoint used, the number of retrieved passages, and the maximum passage length, since DrQA's performance depends strongly on these choices.
- [Table 3] Please define the notation 'A' and explain the binning in the caption; the upper and lower halves of the table are hard to parse without a precise description of how edge contributions from distinct documents are counted.
- [Table 5] The error-scenario percentages sum to 101% for CQ-W and 99% for CQ-T; please specify whether the categories are mutually exclusive, clarify the denominator, and account for the rounding.
- [Sec. 6 and Table 2] Statistical significance is reported only for QUEST over DrQA; please state whether the comparisons against BFS and ShortestPaths were tested, and if so, report those p-values or note that no significance claim is made for them.
Assumptions & free parameters
free parameters (5)
- Cornerstone selection threshold =
0.5
- Alignment edge insertion threshold =
0.5
- Answer merging threshold =
0.5
- Number of top-k GSTs =
50
- Answer type pruning threshold =
not reported
assumptions (5)
- standard math Top-k Group Steiner Trees can be computed optimally by the Ding et al. dynamic programming algorithm with O(n log n) complexity in graph size.
- ad hoc to paper A correct answer is a non-terminal node lying on a low-cost tree that connects at least one node from each group of question-cornerstones.
- domain assumption Token-level similarity from word2vec and AIDA dictionaries correctly identifies cornerstones, alignment edges, and answer types despite no disambiguation.
- domain assumption The custom pattern-based Open IE extractor yields SPO triples that preserve the facts needed to answer the question (85.2% on CQ-W, 82.3% on CQ-T).
- domain assumption The corpus retrieved by the search engine contains the answer and the connecting evidence across documents.
Cite this review
Pith. "Pith review of Answering Complex Questions by Joining Multi-Document Evidence with Quasi Knowledge Graphs." pith.science (2026). https://pith.science/paper/P5RVOYW4
@misc{pith2026190800469,
author = {Pith},
title = {Pith review of: Answering Complex Questions by Joining Multi-Document Evidence with Quasi Knowledge Graphs},
year = {2026},
howpublished = {\url{https://pith.science/paper/P5RVOYW4}},
note = {Machine review of arXiv:1908.00469}
}
read the original abstract
Direct answering of questions that involve multiple entities and relations is a challenge for text-based QA. This problem is most pronounced when answers can be found only by joining evidence from multiple documents. Curated knowledge graphs (KGs) may yield good answers, but are limited by their inherent incompleteness and potential staleness. This paper presents QUEST, a method that can answer complex questions directly from textual sources on-the-fly, by computing similarity joins over partial results from different documents. Our method is completely unsupervised, avoiding training-data bottlenecks and being able to cope with rapidly evolving ad hoc topics and formulation style in user questions. QUEST builds a noisy quasi KG with node and edge weights, consisting of dynamically retrieved entity names and relational phrases. It augments this graph with types and semantic alignments, and computes the best answers by an algorithm for Group Steiner Trees. We evaluate QUEST on benchmarks of complex questions, and show that it substantially outperforms state-of-the-art baselines.
Figures
Forward citations
Cited by 1 Pith paper
-
Enhancing Semantic Document Retrieval- Employing Group Steiner Tree Algorithm with Domain Knowledge Enrichment
SemDR applies a Group Steiner Tree heuristic to a domain-enriched concept graph for agriculture document retrieval and reports large gains over Lucene, ElasticSearch, and Doc2Vec.
Reference graph
Works this paper leans on
-
[1]
A. Abujabal, M. Yahya, M. Riedewald, and G. Weikum. 2017. Automated template generation for question answering over knowledge graphs. In WWW
work page 2017
-
[2]
E. Agichtein, D. Carmel, D. Pelleg, Y. Pinter, and D. Harman. 2015. Overview of the TREC 2015 LiveQA Track. In TREC
work page 2015
- [3]
-
[4]
S. Auer, C. Bizer, G. Kobilarov, J. Lehmann, R. Cyganiak, and Z. Ives. 2007. DBpedia: A nucleus for a Web of open data. In The Semantic Web. Springer
work page 2007
-
[5]
J. Bao, N. Duan, Z. Yan, M. Zhou, and T. Zhao. 2016. Constraint-Based Question Answering with Knowledge Graph. In COLING
work page 2016
-
[6]
H. Bast and E. Haussmann. 2015. More accurate question answering on Freebase. In CIKM
work page 2015
- [7]
-
[8]
G. Bhalotia, A. Hulgeri, C. Nakhe, S. Chakrabarti, and S. Sudarshan. 2002. Keyword searching and browsing in databases using BANKS. In ICDE
work page 2002
Show all 78 references
-
[9]
Bollacker, C
K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor. 2008. Freebase: A collaboratively created graph database for structuring human knowledge. In SIGMOD
2008
-
[10]
Bordes, N
A. Bordes, N. Usunier, S. Chopra, and J. Weston. 2015. Large-scale simple question answering with memory networks. In arXiv
2015
-
[11]
Cai and A
Q. Cai and A. Yates. 2013. Large-scale Semantic Parsing via Schema Matching and Lexicon Extension. In ACL
2013
-
[12]
Camille Chanial, Rédouane Dziri, Helena Galhardas, Julien Leblay, Minh- Huong Le Nguyen, and Ioana Manolescu. 2018. Connectionlens: finding connec- tions across heterogeneous data sources. In VLDB
2018
-
[13]
D. Chen, A. Fisch, J. Weston, and A. Bordes. 2017. Reading Wikipedia to Answer Open-Domain Questions. In ACL
2017
-
[14]
Coffman and A
J. Coffman and A. C. Weaver. 2014. An empirical performance evaluation of relational keyword search techniques. In TKDE
2014
-
[15]
Del Corro and R
L. Del Corro and R. Gemulla. 2013. ClausIE: Clause-based open information extraction. In WWW
2013
-
[16]
R. Das, M. Zaheer, S. Reddy, and A. McCallum. 2017. Question Answering on Knowledge Bases and Text using Universal Schema and Memory Networks. In ACL
2017
-
[17]
Dehghani, H
M. Dehghani, H. Azarbonyad, J. Kamps, and M. de Rijke. 2019. Learning to Transform, Combine, and Reason in Open-Domain Question Answering. In WSDM
2019
-
[18]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv (2018)
2018
-
[19]
Diefenbach, V
D. Diefenbach, V. López, K. Deep Singh, and P. Maret. 2018. Core techniques of question answering systems over knowledge bases: A survey. Knowl. Inf. Syst. 55, 3 (2018)
2018
-
[20]
Dietz and B
L. Dietz and B. Gamari. 2017. TREC CAR: A Data Set for Complex Answer Retrieval. In TREC
2017
-
[21]
B. Ding, J. X. Yu, S. Wang, L. Qin, X. Zhang, and X. Lin. 2007. Finding top- 𝑘 min-cost connected trees in databases. In ICDE
2007
-
[22]
L. Dong, F. Wei, M. Zhou, and K. Xu. 2015. Question Answering over Freebase with Multi-Column Convolutional Neural Networks. In ACL
2015
-
[23]
R. G. Downey and M. R. Fellows. 2013.Fundamentals of Parameterized Complexity. Springer
2013
-
[24]
Fader, L
A. Fader, L. Zettlemoyer, and O. Etzioni. 2013. Paraphrase-driven learning for open question answering. In ACL
2013
-
[25]
Fader, L
A. Fader, L. Zettlemoyer, and O. Etzioni. 2014. Open question answering over curated and extracted knowledge bases. In KDD
2014
-
[26]
Ferrucci et al
D. Ferrucci et al. 2012. This is Watson. IBM Journal Special Issue 56, 3 (2012)
2012
-
[27]
N. Garg, G. Konjevod, and R. Ravi. 2000. A Polylogarithmic Approximation Algorithm for the Group Steiner Tree Problem. J. Algorithms 37, 1 (2000)
2000
-
[28]
Gashteovski, R
K. Gashteovski, R. Gemulla, and L. Del Corro. 2017. MinIE: Minimizing Facts in Open Information Extraction. In EMNLP
2017
-
[29]
Grycner and G
A. Grycner and G. Weikum. 2016. POLY: Mining Relational Paraphrases from Multilingual Sentences. In EMNLP
2016
-
[30]
Ido Guy. 2018. The Characteristics of Voice Search: Comparing Spoken with Typed-in Mobile Web Search Queries. ACM Trans. Inf. Syst. (2018)
2018
-
[31]
M. A. Hearst. 1992. Automatic Acquisition of Hyponyms from Large Text Corpora. In COLING
1992
-
[32]
Hoffart, M
J. Hoffart, M. A. Yosef, I. Bordino, H. Fürstenau, M. Pinkal, M. Spaniol, B. Taneva, S. Thater, and G. Weikum. 2011. Robust disambiguation of named entities in text. In EMNLP
2011
-
[33]
S. Hu, L. Zou, J. X. Yu, H. Wang, and D. Zhao. 2018. Answering Natural Language Questions by Subgraph Matching over Knowledge Graphs. Trans. on Know. and Data Eng. 30, 5 (2018)
2018
-
[34]
Iyyer, J
M. Iyyer, J. L. Boyd-Graber, L. M. B. Claudino, R. Socher, and H. Daumé III. 2014. A Neural Network for Factoid Question Answering over Paragraphs. In EMNLP
2014
-
[35]
Joshi, E
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In ACL
2017
-
[36]
Joshi, U
M. Joshi, U. Sawant, and S. Chakrabarti. 2014. Knowledge graph and corpus driven segmentation and answer inference for telegraphic entity-seeking queries. In EMNLP
2014
-
[37]
Kacholia, S
V. Kacholia, S. Pandit, S. Chakrabarti, S. Sudarshan, R. Desai, and H. Karambelkar
-
[38]
Kadry and L
A. Kadry and L. Dietz. 2017. Open Relation Extraction for Support Passage Retrieval: Merit and Open Issues. In SIGIR
2017
-
[39]
Kasneci, M
G. Kasneci, M. Ramanath, M. Sozio, F. M. Suchanek, and G. Weikum. 2009. STAR: Steiner-tree approximation in relationship graphs. In ICDE
2009
-
[40]
T. Khot, A. Sabharwal, and P. Clark. 2017. Answering complex questions using open information extraction. In ACL
2017
-
[41]
C. Kwok, O. Etzioni, and D. S. Weld. 2001. Scaling Question Answering to the Web. ACM Trans. Inf. Syst. (2001)
2001
-
[42]
R. Li, L. Qin, J. X. Yu, and R. Mao. 2016. Efficient and progressive group steiner tree search. In SIGMOD
2016
-
[43]
Y. Lin, H. Ji, Z. Liu, and M. Sun. 2018. Denoising distantly supervised open-domain question answering. In ACL
2018
-
[44]
Manning, M
C. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. Bethard, and D. McClosky. 2014. The Stanford CoreNLP natural language processing toolkit. In ACL
2014
-
[45]
Mausam. 2016. Open information extraction systems and downstream applica- tions. In IJCAI
2016
-
[46]
Mikolov, I
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. 2013. Distributed Representations of Words and Phrases and their Compositionality. In NIPS
2013
-
[47]
Pavlick, P
E. Pavlick, P. Rastogi, J. Ganitkevitch, B. Van Durme, and C. Callison-Burch. 2015. PPDB 2.0: Better paraphrase ranking, fine-grained entailment relations, word embeddings, and style classification. In ACL
2015
-
[48]
Pennington, R
J. Pennington, R. Socher, and C. D. Manning. 2014. GloVe: Global Vectors for Word Representation. In EMNLP
2014
-
[49]
Rajpurkar, J
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. In EMNLP
2016
-
[50]
Ravichandran and E
D. Ravichandran and E. Hovy. 2002. Learning Surface Text Patterns for a Question Answering System. In ACL
2002
-
[51]
Savenkov and E
D. Savenkov and E. Agichtein. 2016. When a knowledge base is not enough: Question answering over knowledge bases with external text data. In SIGIR
2016
-
[52]
Sawant and S
U. Sawant and S. Chakrabarti. 2013. Learning joint query interpretation and response ranking. In WWW
2013
-
[53]
V. I. Spitkovsky and A. X. Chang. 2012. A Cross-Lingual Dictionary for English Wikipedia Concepts. In LREC. 3168–3175
2012
-
[54]
F. M. Suchanek, G. Kasneci, and G. Weikum. 2007. YAGO: A core of semantic knowledge. In WWW
2007
-
[55]
H. Sun, B. Dhingra, M. Zaheer, K. Mazaitis, R. Salakhutdinov, and W. W. Cohen
-
[56]
H. Sun, H. Ma, W. Yih, C. Tsai, J. Liu, and M. Chang. 2015. Open domain question answering via semantic enrichment. In WWW
2015
-
[57]
Talmor and J
A. Talmor and J. Berant. 2018. The Web as a Knowledge-base for Answering Complex Questions. In NAACL-HLT
2018
-
[58]
C. Tan, F. Wei, Q. Zhou, N. Yang, B. Du, W. Lv, and M. Zhou. 2018. Context-Aware Answer Sentence Selection With Hierarchical Gated Recurrent Neural Networks. IEEE/ACM Trans. Audio, Speech & Language Processing 26, 3 (2018)
2018
-
[59]
Unger, L
C. Unger, L. Bühmann, J. Lehmann, A. N. Ngomo, D. Gerber, and P. Cimiano
-
[60]
Unger, A
C. Unger, A. Freitas, and P. Cimiano. 2014. An Introduction to Question Answer- ing over Linked Data. In Reasoning Web
2014
-
[61]
Usbeck, A
R. Usbeck, A. N. Ngomo, B. Haarmann, A. Krithara, M. Röder, and G. Napolitano
-
[62]
E. M. Voorhees. 2014. The Effect of Sampling Strategy on Inferred Measures. In SIGIR
2014
-
[63]
E. M. Voorhees and D. K. Harman. 2005. TREC: Experiment and evaluation in information retrieval. MIT press Cambridge
2005
-
[64]
Vrandečić and M
D. Vrandečić and M. Krötzsch. 2014. Wikidata: A free collaborative knowledge base. Commun. ACM 57, 10 (2014)
2014
-
[65]
R. W. White, M. Richardson, and W. T. Yih. 2015. Questions vs. queries in informational search tasks. In WWW. 135–136
2015
-
[66]
Wieting, M
J. Wieting, M. Bansal, K. Gimpel, and K. Livescu. 2016. Towards universal para- phrastic sentence embeddings. In ICLR
2016
-
[67]
K. Xu, S. Reddy, Y. Feng, S. Huang, and D. Zhao. 2016. Question answering on freebase via relation extraction and textual evidence. In ACL
2016
-
[68]
Yahya, K
M. Yahya, K. Berberich, S. Elbassuoni, M. Ramanath, V. Tresp, and G. Weikum
-
[69]
Yahya, S
M. Yahya, S. Whang, R. Gupta, and A. Halevy. 2014. ReNoun: Fact extraction for nominal attributes. In EMNLP
2014
-
[70]
W. Yih, M. Chang, X. He, and J. Gao. 2015. Semantic Parsing via Staged Query Graph Generation: Question Answering with Knowledge Base. In ACL
2015
-
[71]
P. Yin, N. Duan, B. Kao, J. Bao, and M. Zhou. 2015. Answering Questions with Complex Semantic Constraints on Open Knowledge Bases. In CIKM
2015
-
[72]
J. X. Yu, L. Qin, and L. Chang. 2009. Keyword Search in Databases. M & C
2009
-
[73]
In EMNLP
Natural Language Questions for the Web of Data. In EMNLP
-
[78]
Ziegler, A
D. Ziegler, A. Abujabal, R. Saha Roy, and G. Weikum. 2017. Efficiency-aware An- swering of Compositional Questions using Answer Type Prediction. In IJCNLP.10
2017
-
[2005]
Bidirectional expansion for keyword search on graph databases. In VLDB
-
[2012]
Template-based question answering over RDF data. In WWW
-
[2017]
7th Open Challenge on Question Answering over Linked Data. In QALD
-
[2018]
In EMNLP
Open domain question answering using early fusion of knowledge bases and text. In EMNLP
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.