REVIEW 4 major objections 4 minor 1 cited by
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper introduces ViQA-COVID, a 6,444-question-answer Vietnamese dataset for COVID-19 machine reading comprehension, and reports that XLM-R large is the strongest tested model.
desk verdict A genuinely new Vietnamese multi-span COVID-19 MRC dataset, but the paper's own statistics don't add up and the benchmark numbers rest on unverified preprocessing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the multi-span extraction setup: instead of predicting a single start–end interval, the models are trained with a sequence-tagging head that labels each token B (begin), I (inside), or O (outside) an answer span, so multiple answers can be recovered from B and O tokens. Because most passages are longer than the 384/256-token model input limit, the input pipeline splits passages into overlapping features with a stride of 128, which is what the paper relies on to keep answer spans intact. This combination—B/I/O tagging plus sliding-window splitting—is what the paper uses to make a long-passage, multi-answer Vietnamese benchmark tractable for BERT-style encoders.
What would settle it
Re-annotate a random sample of the passages with independent annotators and compute inter-annotator agreement on the answer spans; if agreement is low (especially on multi-span answers), the reported benchmark numbers cannot be trusted as measures of model capability.
Extended reading notes
Core claim
The central claim is that ViQA-COVID is a valid, reusable benchmark for Vietnamese COVID-19 machine reading comprehension, and the first Vietnamese multi-span extraction MRC dataset. The dataset contains 6,444 question-answer pairs over 537 passages, with 21.0–21.3% of answers multi-span, 10–12% non-span (unanswerable), and the rest single-span. On this benchmark, XLM-R large outperforms the other four tested models, achieving 72.00% exact match and 85.97% F1 on the test set, confirming that cross-lingual pretraining transfers to Vietnamese health texts better than Vietnamese-only PhoBERT variants. The paper also analyzes error types, showing that multi-span questions and long sequences of dates, places, or people are the main sources of failure.
Load-bearing premise
The dataset's ground truth is taken as correct without a reported inter-annotator agreement measure, so if the manual answer spans are noisy or inconsistent, the model scores are not a valid measure of reading ability.
Editorial extensions
If this is right
- Vietnamese health-domain question-answering systems can now be trained and evaluated on a COVID-19-specific benchmark instead of relying only on general-domain Vietnamese datasets.
- Multi-span extraction becomes a measurable sub-task in Vietnamese NLP, with about 21% of answers requiring multiple spans, so progress on that specific challenge can be tracked.
- The reported model ranking (XLM-R large best, then XLM-R base, then PhoBERT variants, then mBERT) gives practitioners a clear baseline hierarchy for future work on Vietnamese MRC.
- Because the dataset includes unanswerable questions (about 10–12%), it also supports evaluation of a model's ability to abstain rather than hallucinate an answer.
Reading between the lines
- If ViQA-COVID is released and maintained, it could become a standard low-resource Vietnamese reading benchmark, and its 21% multi-span share might push Vietnamese models toward span-set decoding rather than single-interval predictions.
- The error analysis suggests that enumerations of dates, places, and people are the hardest cases; an extension would be a dataset split that deliberately oversamples those question types to stress-test models.
- The same annotation pipeline of public-health reports plus expert-advised questions could transfer to other Vietnamese health topics, such as dengue or influenza, without redesigning the dataset format.
- Because a non-Vietnamese-specific XLM-R large wins despite the competition, a testable follow-up is whether a Vietnamese-specific model trained with a span-set objective would close the multi-span gap.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ViQA-COVID, an extractive machine reading comprehension dataset for Vietnamese focused on COVID-19, constructed from CDC Vietnam and other reputable sources. The dataset contains 6,444 question-answer pairs over 537 passages, with roughly 20% of answers requiring multiple spans, and it is claimed to be the first multi-span extraction MRC dataset and the first COVID-19 MRC dataset for Vietnamese. The authors also report fine-tuning experiments with mBERT, PhoBERT-base/large, and XLM-R-base/large, concluding that XLM-R-large achieves the best test performance with 85.97% F1 and 72.00% EM. The paper includes dataset statistics, a description of the annotation process, an error analysis, and a discussion of the difficulty profile.
Significance. If the dataset and benchmark results are reliable, ViQA-COVID fills a real gap: there is no public COVID-19 MRC benchmark for Vietnamese, and a multi-span extraction dataset for Vietnamese would be a useful resource for low-resource NLP and health-domain QA. The paper provides a substantial annotation effort, documents question-type and answer-type distributions, reports experiments with four pretrained model families, and includes an error analysis that gives a concrete picture of the remaining challenges. The claimed contribution is therefore significant for the Vietnamese NLP community. The main value depends on the dataset being released and on the annotation quality and evaluation pipeline being verifiable, which the current manuscript does not fully establish.
major comments (4)
- [Section 4.2 and Section 4.4] The passage-splitting scheme does not guarantee that every gold answer span is contained in at least one input feature. With maximum feature length 384 (PhoBERT: 256) and stride 128, a span such as tokens 250 to 400 falls neither in window [0, 383] nor in window [256, 639]; the overlap region is only 128 tokens wide, so spans longer than the stride that straddle a split boundary are unreachable. The paper states that overlap handles answers at split positions, but it reports no check that every gold span is a substring of some feature, no count of skipped or truncated examples, and no special handling for multi-span answers whose individual spans cross boundaries. Since Table 3 indicates 475 passages in the "greater than or equal to 512 tokens" class, this is not a corner case; if any gold spans in the test set are unreachable, the EM/F1 values in Table 4 are not valid estimates of performance on ViQA-COVID as claimed.
- [Section 3.2, Tables 1 and 3] The passage totals in Tables 1 and 3 are inconsistent. Table 1 reports 537 passages (284 train + 139 dev + 114 test = 537), whereas Table 3's length distribution sums to 612 (335 train + 151 dev + 126 test = 612). Because Section 4.2 motivates the splitting procedure directly from Table 3, this discrepancy makes the preprocessing pipeline unverifiable. The authors should reconcile the two tables and state explicitly whether Table 3 counts original passages, split input features, or some other unit. Until this is resolved, the relationship between the reported dataset statistics and the actual experimental setup is unclear.
- [Section 3.1] No inter-annotator agreement is reported. The annotation process is described as creation and cross-checking by three CDC analysts with advice from two CDC experts, but no quantitative consistency measure, such as span-level agreement, Cohen's kappa, or a small-scale double-annotation study, is provided. Since the ground-truth spans are the basis for all reported benchmark scores, and since multi-span annotation is inherently more subjective than single-span annotation, the reliability of the labels is unverified. The authors should add an agreement statistic or a carefully described verification study; otherwise, the model scores in Table 4 cannot be distinguished from scores on noisy labels.
- [Section 4.3 and Section 5] The evaluation protocol for non-span and multi-span answers is underspecified. Table 1 shows that 10 to 12 percent of answers are non-span, and roughly 20 percent are multi-span, but the paper never states how EM and F1 are computed for these cases under the B/I/O tagging approach. For example, is a predicted span on an unanswerable question scored as 0, and how are partially overlapping sets of predicted spans aggregated into a single F1? Without this definition, the aggregate numbers in Table 4 cannot be reproduced or meaningfully compared with other MRC benchmarks. The authors should provide the exact scoring formula used for non-span and multi-span predictions.
minor comments (4)
- [Abstract and throughout] The manuscript contains numerous grammatical and typographical errors, such as "After two years of appearance," "a answer can include multi-span," and "Data was encrypted sensitive information." A thorough language edit would improve clarity.
- [Sections 4.2 and 4.4] Section 4.2 says the models' maximum input feature length is 512 tokens, while Section 4.4 states that the maximum feature length used is 384 (PhoBERT 256). These statements should be reconciled, and the choice of 384/256 with stride 128 should be justified in terms of answer length and model capacity.
- [Section 3.2 and Section 6] The paper says the dataset will be "publicly release[d] soon," but no URL or repository is provided. For a resource paper, a release link or a concrete availability statement is expected, and without public access the central contribution cannot yet be used or independently checked.
- [Section 5 and Table 4] The experimental results are reported from what appears to be a single run per model. No confidence intervals, multiple-seed standard deviations, or significance tests are given, so the smaller performance gaps, such as XLM-R-base versus XLM-R-large on multi-span F1 (77.83 vs. 79.10 on the test set), may not be statistically meaningful. Adding variance across seeds or a paired significance test would strengthen the comparative claim.
Circularity Check
No significant circularity; benchmark results are empirical measurements on a newly constructed dataset, and the only self-citations are non-load-bearing related-work mentions.
full rationale
The paper's central claims are the creation of the ViQA-COVID dataset and the measured baseline performances on it. Neither claim is derived from an input that already contains the output: the annotation process (Section 3.1) constructs questions and answer spans from CDC and online sources independently of any model, and Table 4 reports directly measured EM and F1 scores for standard pretrained models. There is no fitted parameter that is renamed as a prediction, and the sequence-tagging multi-span approach is imported from an external source [20]. The self-citations [8] and [9] appear only in the related-work discussion as Vietnamese COVID-19 NER datasets and are not load-bearing for the dataset's construction, its firstness claim, or the benchmark numbers. The claim that ViQA-COVID is the first multi-span extraction MRC dataset for Vietnamese is a literature assertion rather than a derivation, so it cannot reduce to the paper's own assumptions. Other concerns, such as the absence of inter-annotator agreement statistics, the discrepancy between 537 passages in Table 1 and 612 rows in Table 3, and whether the sliding-window preprocessing of Section 4.2 preserves every gold span, are validity and correctness risks rather than instances of circularity. Accordingly, no specific circular step can be exhibited from the text, and the appropriate finding is a low score of 1, reflecting only the presence of minor non-load-bearing self-citations.
Assumptions & free parameters
free parameters (3)
- learning rate =
5e-5
- max feature length / stride =
384 (PhoBERT 256) / 128
- training epochs / batch size / weight decay =
30 / 32 / 0.01
assumptions (3)
- domain assumption Manual annotation by three CDC analysts with two CDC expert advisors yields correct answer spans without a measured inter-annotator agreement metric.
- domain assumption Splitting passages into overlapping features with stride 128 preserves every answer span and scoring integrity.
- domain assumption RDRSegmenter word segmentation is correct for the passages and aligns with PhoBERT's tokenization.
Cite this review
Pith. "Pith review of ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese." pith.science (2026). https://pith.science/paper/LKSNO5EW
@misc{pith2026250421017,
author = {Pith},
title = {Pith review of: ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese},
year = {2026},
howpublished = {\url{https://pith.science/paper/LKSNO5EW}},
note = {Machine review of arXiv:2504.21017}
}
read the original abstract
After two years of appearance, COVID-19 has negatively affected people and normal life around the world. As in May 2022, there are more than 522 million cases and six million deaths worldwide (including nearly ten million cases and over forty-three thousand deaths in Vietnam). Economy and society are both severely affected. The variant of COVID-19, Omicron, has broken disease prevention measures of countries and rapidly increased number of infections. Resources overloading in treatment and epidemics prevention is happening all over the world. It can be seen that, application of artificial intelligence (AI) to support people at this time is extremely necessary. There have been many studies applying AI to prevent COVID-19 which are extremely useful, and studies on machine reading comprehension (MRC) are also in it. Realizing that, we created the first MRC dataset about COVID-19 for Vietnamese: ViQA-COVID and can be used to build models and systems, contributing to disease prevention. Besides, ViQA-COVID is also the first multi-span extraction MRC dataset for Vietnamese, we hope that it can contribute to promoting MRC studies in Vietnamese and multilingual.
Figures
Forward citations
Cited by 1 Pith paper
-
Nested Named-Entity Recognition on Vietnamese COVID-19: Dataset and Experiments
A manually annotated Vietnamese COVID-19 NER dataset with 11 entity types and up to four nesting levels, plus BiLSTM and PhoBERT baselines where PhoBERT-large-CRF with cross-sentence context achieved the highest F1.
Reference graph
Works this paper leans on
-
[1]
Adnane Cabani, Karim Hammoudi, Halim Benhab- iles, and Mahmoud Melkemi. 2020. Maskedface-net – a dataset of correctly/incorrectly masked face images in the context of covid-19.Smart Health
work page 2020
-
[2]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Fran- cisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsuper- vised cross-lingual representation learning at scale. InProceedings of the 58th Annual Meeting of the As- sociation for Computational Linguistics, pages 8440– 8451, Onlin...
work page 2020
-
[3]
Pradeep Dasigi, Nelson F. Liu, Ana Marasovi ´c, Noah A. Smith, and Matt Gardner. 2019. Quoref: A reading comprehension dataset with questions re- quiring coreferential reasoning. InProceedings of the 2019 Conference on Empirical Methods in Natu- ral Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP...
work page 2019
-
[4]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language under- standing. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, Volume 1 (Long and Short Papers), pages 4171–4186, Mi...
work page 2019
-
[5]
Phong Nguyen-Thuan Do, Nhat Duy Nguyen, Tin Van Huynh, Kiet Van Nguyen, Anh Gia-Tuan Nguyen, and Ngan Luu-Thuy Nguyen. 2021. Sen- tence extraction-based machine reading comprehen- sion for vietnamese.CoRR, abs/2105.09043
arXiv 2021
-
[6]
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requir- ing discrete reasoning over paragraphs. InProceed- ings of the 2019 Conference of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technologies, Volume 1 (Long and S...
work page 2019
-
[7]
Zeqian Ju, Subrato Chakravorty, Xuehai He, Shu Chen, Xingyi Yang, and Pengtao Xie. 2020. Covid- dialog: Medical dialogue datasets about covid-19. https://github.com/UCSD-AI4H/COVID-Dialogue
work page 2020
-
[8]
Ngoc C Lê, Hai-Chung Nguyen-Phung, Thu- Huong Pham Thi, Hue Vu, Phuong-Thao Nguyen Thi, Thu-Thuy Tran, Hong-Nhung Le Thi, Thuy- Duong Nguyen-Thi, and Thanh-Huy Nguyen. 2025. Nested named-entity recognition on vietnamese covid-19: Dataset and experiments.arXiv preprint arXiv:2504.21016
work page Pith review arXiv 2025
Show all 36 references
-
[9]
Lê, Hai-Chung Nguyen-Phung, Thuy Thu Tran, Ngoc-Uyen Thi Nguyen, Dang-Khoi Pham Nguyen, and Thanh-Huy Nguyen
Ngoc C. Lê, Hai-Chung Nguyen-Phung, Thuy Thu Tran, Ngoc-Uyen Thi Nguyen, Dang-Khoi Pham Nguyen, and Thanh-Huy Nguyen. 2023.On Natural Language Processing to Attack COVID-19 Pandemic: Experiences of Vietnam, pages 313–335. Springer International Publishing, Cham
2023
-
[10]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov
-
[11]
Ilya Loshchilov and Frank Hutter. 2019. Decou- pled weight decay regularization. InInternational Conference on Learning Representations
2019
-
[12]
Timo Möller, Anthony Reina, Raghavan Jayakumar, and Malte Pietsch. 2020. COVID-QA: A question answering dataset for COVID-19. InProceedings of the 1st Workshop on NLP for COVID-19 at ACL 2020, Online. Association for Computational Linguistics
2020
-
[13]
Dat Quoc Nguyen and Anh Tuan Nguyen. 2020. PhoBERT: Pre-trained language models for Viet- namese. InFindings of the Association for Computa- tional Linguistics: EMNLP 2020, pages 1037–1042
2020
-
[14]
Dat Quoc Nguyen, Dai Quoc Nguyen, Thanh Vu, Mark Dras, and Mark Johnson. 2018. A Fast and Ac- curate Vietnamese Word Segmenter. InProceedings of the 11th International Conference on Language Resources and Evaluation (LREC 2018), pages 2582– 2587
2018
-
[15]
Kiet Nguyen, Vu Nguyen, Anh Nguyen, and Ngan Nguyen. 2020. A Vietnamese dataset for evaluating machine reading comprehension. InProceedings of the 28th International Conference on Computational Linguistics, pages 2595–2605, Barcelona, Spain (On- line). International Committee ...
2020
-
[16]
Kiet Van Nguyen, Duc-Vu Nguyen, Anh Gia-Tuan Nguyen, and Ngan Luu-Thuy Nguyen. 2020. New vietnamese corpus for machine readingcomprehen- sion of health news articles.CoRR, abs/2006.11138
2020 arXiv
-
[17]
Luu, Anh Gia-Tuan Nguyen, and Ngan Luu-Thuy Nguyen
Kiet Van Nguyen, Khiem Vinh Tran, Son T. Luu, Anh Gia-Tuan Nguyen, and Ngan Luu-Thuy Nguyen
-
[18]
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable ques- tions for SQuAD. InProceedings of the 56th Annual Meeting of the Association for Computational Lin- guistics (Volume 2: Short Papers), pages 784–789, Melbourne, Australia. Associat...
2018
-
[19]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. InProceedings of the 2016 Conference on Empirical Methods in Nat- ural Language Processing, pages 2383–2392, Austin, Texas. Association for Com...
2016
-
[20]
Elad Segal, Avia Efrat, Mor Shoham, Amir Glober- son, and Jonathan Berant. 2020. A simple and effec- tive model for answering multi-span questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3074–3080, Online. Assoc...
2020
-
[21]
Raphael Tang, Rodrigo Nogueira, Edwin Zhang, Nikhil Gupta, Phuong Cam, Kyunghyun Cho, and Jimmy Lin. 2020. Rapidly bootstrapping a ques- tion answering dataset for COVID-19.CoRR, abs/2004.11339
2020 arXiv
-
[22]
Thinh Hung Truong, Mai Hoang Dao, and Dat Quoc Nguyen. 2021. COVID-19 Named Entity Recogni- tion for Vietnamese. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
2021
-
[23]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. InAdvances in Neural Information Pro- cessing Systems, volume 30. Curran Associates, Inc
2017
-
[24]
Thanh Vu, Dat Quoc Nguyen, Dai Quoc Nguyen, Mark Dras, and Mark Johnson. 2018. VnCoreNLP: A Vietnamese natural language processing toolkit. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Compu- tational Linguistics: Demonstrations, p...
2018
-
[25]
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. Superglue: A stickier benchmark for general-purpose language un- derstanding systems. InAdvances in Neural Infor- mation Processing Systems, volume 32...
2019
-
[26]
Alex Wang, Amanpreet Singh, Julian Michael, Fe- lix Hill, Omer Levy, and Samuel Bowman. 2018. GLUE: A multi-task benchmark and analysis plat- form for natural language understanding. InProceed- ings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Netw...
2018
-
[27]
Linda Wang, Zhong Qiu Lin, and Alexander Wong
-
[28]
Zhongyuan Wang, Guangcheng Wang, Baojin Huang, Zhangyang Xiong, Qi Hong, Hao Wu, Peng Yi, Kui Jiang, Nanxi Wang, Yingjiao Pei, Heling Chen, Yu Miao, Zhibing Huang, and Jinbi Liang
-
[29]
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2019. Ccnet: Extracting high quality monolingual datasets from web crawl data
2019
-
[30]
Covid-net: a tailored deep convolutional neural network design for detection of covid-19 cases from chest x-ray images.Scientific Reports, 10(1):19549
-
[31]
Xinliang Frederick Zhang, Heming Sun, Xiang Yue, Simon Lin, and Huan Sun. 2021. COUGH: A chal- lenge dataset and models for COVID-19 FAQ re- trieval. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, pages 3759–3769
2021
-
[32]
Masked face recognition dataset and applica- tion.CoRR, abs/2003.09093
2003 arXiv
-
[34]
Guangtao Zeng, Qingyang Wu, Yichen Zhang, Zhou Yu, Eric Xing, and Pengtao Xie. 2020. Develop medical dialogue systems for covid-19. https://github.com/UCSD-AI4H/COVID-Dialogue
2020
-
[36]
Ming Zhu, Aman Ahuja, Da-Cheng Juan, Wei Wei, and Chandan K. Reddy. 2020. Question answer- ing with long multiple-span answers. InFindings of the Association for Computational Linguistics: EMNLP 2020, pages 3840–3849, Online. Association for Computational Linguistics
2020
-
[2019]
Cite arxiv:1907.11692
Roberta: A robustly optimized bert pretraining approach. Cite arxiv:1907.11692
1907 arXiv
-
[2020]
Enhancing lexical-based approach with ex- ternal knowledge for vietnamese multiple-choice machine reading comprehension.IEEE Access, 8:201404–201417
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.