Pith. sign in

REVIEW 1 cited by

QUADRo: Dataset and Models for QUestion-Answer Database Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.01003 v1 pith:O4BKDDXL submitted 2023-03-30 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelsquestionpairsquestionsanswerapproachcompetitivedata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

An effective paradigm for building Automated Question Answering systems is the re-use of previously answered questions, e.g., for FAQs or forum applications. Given a database (DB) of question/answer (q/a) pairs, it is possible to answer a target question by scanning the DB for similar questions. In this paper, we scale this approach to open domain, making it competitive with other standard methods, e.g., unstructured document or graph based. For this purpose, we (i) build a large scale DB of 6.3M q/a pairs, using public questions, (ii) design a new system based on neural IR and a q/a pair reranker, and (iii) construct training and test data to perform comparative experiments with our models. We demonstrate that Transformer-based models using (q,a) pairs outperform models only based on question representation, for both neural search and reranking. Additionally, we show that our DB-based approach is competitive with Web-based methods, i.e., a QA system built on top the BING search engine, demonstrating the challenge of finding relevant information. Finally, we make our data and models available for future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. QuIM-RAG: Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance

    cs.CL 2025-01 conditional novelty 4.0 of 10

    QuIM-RAG retrieves chunks by matching a user question to LLM-generated questions from each chunk in a quantized embedding space, reporting higher QA scores than a traditional RAG baseline on an NDSU website corpus.

Pith tools