ViQA-COVID is a new Vietnamese COVID-19 reading comprehension dataset with 6,444 question-answer pairs, the first for Vietnamese with multi-span answers, benchmarked at 85.97% F1 by XLM-R large.
Rapidly Bootstrapping a Question Answering Dataset for COVID-19
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We present CovidQA, the beginnings of a question answering dataset specifically designed for COVID-19, built by hand from knowledge gathered from Kaggle's COVID-19 Open Research Dataset Challenge. To our knowledge, this is the first publicly available resource of its type, and intended as a stopgap measure for guiding research until more substantial evaluation resources become available. While this dataset, comprising 124 question-article pairs as of the present version 0.1 release, does not have sufficient examples for supervised machine learning, we believe that it can be helpful for evaluating the zero-shot or transfer capabilities of existing models on topics specifically related to COVID-19. This paper describes our methodology for constructing the dataset and presents the effectiveness of a number of baselines, including term-based techniques and various transformer-based models. The dataset is available at http://covidqa.ai/
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
ViQA-COVID is a new Vietnamese COVID-19 reading comprehension dataset with 6,444 question-answer pairs, the first for Vietnamese with multi-span answers, benchmarked at 85.97% F1 by XLM-R large.