REVIEW 2 cited by
More Than Reading Comprehension: A Survey on Datasets and Metrics of Textual Question Answering
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Textual Question Answering (QA) aims to provide precise answers to user's questions in natural language using unstructured data. One of the most popular approaches to this goal is machine reading comprehension(MRC). In recent years, many novel datasets and evaluation metrics based on classical MRC tasks have been proposed for broader textual QA tasks. In this paper, we survey 47 recent textual QA benchmark datasets and propose a new taxonomy from an application point of view. In addition, We summarize 8 evaluation metrics of textual QA tasks. Finally, we discuss current trends in constructing textual QA benchmarks and suggest directions for future work.
Forward citations
Cited by 2 Pith papers
-
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
MTRAG is a human-generated multi-turn RAG benchmark (110 conversations, 842 tasks, four domains) on which state-of-the-art LLM RAG systems perform poorly.
-
QA-TOOLBOX: Conversational Question-Answering for process task guidance in manufacturing
An LLM-augmented dataset and baseline evaluation for manufacturing task guidance QA, using LLM-as-a-judge with expert validation.
Discussion (0). Continue with ORCID to comment.