{"id":"b64c2af4-b6b1-40eb-bd26-0b5b250ae7ba","arxiv_id":"2504.21016","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A manually annotated Vietnamese COVID-19 NER dataset with 11 entity types and up to four nesting levels, plus BiLSTM and PhoBERT baselines where PhoBERT-large-CRF with cross-sentence context achieved the highest F1.","lead":"This paper introduces a manually annotated Vietnamese dataset for nested named-entity recognition on COVID-19 reports, with 11 entity types and up to four levels of nesting. It reports baselines where PhoBERT-large-CRF with cross-sentence context performed best, but the dataset is not publicly released and the entity count is reported inconsistently.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dataset description is internally inconsistent: abstract claims 11,128 entities while Table 2 sums to 17,818, and no data release link is provided, so the central resource claim is unverifiable.","rationale":"The reader correctly identified dataset reliability as the weak point, and I agree that the paper should not be fully accepted without stronger evidence about the annotations. However, my load-bearing concern is narrower and more concrete than the absence of inter-annotator agreement: the paper's own numbers are irreconcilable. Abstract and Section 3.2 report 11,128 entities, while Table 2 sums to 17,818. Since the dataset is not released, this inconsistency cannot be resolved by inspection, and the central contribution is therefore unverifiable. This is an internal inconsistency rather than a disagreement with external consensus, so it is a correctness risk, not a matter of taste. I give credit for the conventional experimental design and the credible use of PhoBERT baselines, but these do not repair the resource-level claim. The natural next step is a public release with a corrected count and, ideally, annotation reliability measures; conditional acceptance on that release is the appropriate verdict. I do not see grounds for rejection, because the mismatch could reflect a missing counting convention rather than a fabricated dataset, and the experimental findings may still hold once the statistics are fixed.","tokens_in":9629,"tokens_out":3319,"duration_ms":32923,"concrete_test":"Release the annotation files (or a redacted version) and run a deterministic tag-counting script that parses every ENAMEX tag, counts entities by type and nesting level, counts unique surface spans ignoring nesting, and re-counts sentences after the reported RDRsegmenter pipeline. Then check whether the abstract's 11,128, the Table 2 sums, and the level sums reconcile under any stated counting convention, and recompute the reported F1 scores on the same corrected split. If the counts do not reconcile, the paper must correct the statistics and rerun the experiments before the resource claim can be accepted.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is a reusable manually annotated dataset of 10,271 Vietnamese COVID-19 sentences with nested entities. The only quantitative characterization of that resource is inconsistent. The abstract and Section 3.2 say 11,128 entities, but Table 2 reports per-type counts that sum to 17,818, and the per-level counts 15,870 + 1,501 + 375 + 72 also sum to 17,818. This is not a one-row typo: it changes the reported dataset size by roughly 60%. Because the dataset is not released and no public link is given, a reader cannot determine whether the correct aggregate is 11,128 unique surface mentions, 17,818 ENAMEX tags, or some third counting convention, and cannot reproduce the train/valid/test split behind Tables 5–6. The resource is the contribution; if its size and contents cannot be pinned down, the benchmark numbers inherit an unverified base. The absence of inter-annotator agreement and an annotation guideline makes the annotation-quality question undecidable as well, but the count inconsistency is directly observable and is the more load-bearing internal defect.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a manually annotated Vietnamese COVID-19 named-entity recognition dataset with nested entities, 11 entity types, up to four nesting levels, and 10,271 sentences. It describes the entity types, the three-phase annotation process, and evaluates BiLSTM-CRF, PhoBERT base/large, and CRF variants with and without cross-sentence context. The authors report that PhoBERTlarge-CRF with cross-sentence context yields the best results, with test F1 of 82.34 for Level-1 and 67.23 for Level-2 entities. The paper does not provide a link to the dataset or code.","tokens_in":9832,"tokens_out":4432,"duration_ms":36519,"significance":"If the dataset were released and its description corrected, it could be a reusable resource for Vietnamese clinical NER, particularly for fine-grained date and contact entities relevant to epidemic tracing. The comparative experiments provide a useful set of baselines, and the cross-sentence context idea is a sensible adaptation of prior work. However, the contribution's significance is currently limited by an unresolved internal inconsistency in the reported dataset size, the absence of any annotation-quality metrics, and the lack of a public data or code release. The circularity concern raised in the stress test does not land: all results are measured on a held-out test split, so there is no circularity in the evaluation.","major_comments":[{"comment":"The abstract and Section 3.2 state that the dataset contains 10,271 sentences and 11,128 entities, but Table 2 reports per-type counts that sum to 17,818 and per-level counts (15,870 + 1,501 + 375 + 72) that also sum to 17,818. The intermediate Phase 1 and Phase 2 entity counts (6,481 and 7,069) do not sum to 11,128 either. Because no dataset URL or counting convention is specified, a reader cannot determine whether 11,128 refers to unique mentions, 17,818 to ENAMEX labels, or some other quantity. This internal inconsistency directly affects the central resource claim and must be resolved.","section":"Abstract and Section 3.2 vs. Table 2"},{"comment":"No inter-annotator agreement, annotation guideline, or adjudication procedure is reported. The paper states that the data was 're-checked manually' and 'reviewed' but gives no quantitative reliability measure such as Cohen's kappa. Since the models are trained and evaluated entirely against this manual annotation, label quality is an unverified load-bearing premise for the reported F1 scores.","section":"Section 3.2 (Annotation)"},{"comment":"Results are reported as means over five runs without standard deviations, confidence intervals, or significance tests. For example, in Table 5 the Level-1 F1 gap between PhoBERTlarge-CRF with cross-sentence context (82.34) and PhoBERTbase-CRF with cross-sentence context (81.89) is 0.45 points; without variance information it is impossible to judge whether this difference is meaningful. Please report per-run results or error bars and, if possible, significance tests.","section":"Sections 4.3–4.4, Tables 3–6"},{"comment":"The paper provides no link, DOI, or repository for the dataset or the code used in the experiments. For a dataset contribution, this is a major omission: the announced resource cannot be downloaded, and the benchmark numbers cannot be independently reproduced. A public release (or a clear statement of access conditions) is needed for the claims to be verifiable.","section":"Dataset availability"}],"minor_comments":[{"comment":"The numeric entries are malformed in several places, with numbers concatenated together, e.g., '85.7883.9876.29' and '81.1278.63'. These tables need to be regenerated so that each value is clearly separated.","section":"Tables 5 and 6"},{"comment":"The reference list appears in full twice, after Section 5 and again after the first copy of the references; the duplicate should be removed.","section":"References"},{"comment":"The manuscript contains numerous English typos and grammatical errors (e.g., 'which be defined', 'recognization', 'takes place' in the abstract). A thorough proofreading pass is needed.","section":"Language and typos"},{"comment":"The paper says sentences and words were segmented with both RDRsegmenter and Trankit, and then RDRsegmenter was chosen; please explain the selection criterion and why Trankit was discarded.","section":"Section 4.1"},{"comment":"It is unclear whether the per-entity results in Table 7 are for Level-1 entities only or include nested levels; please clarify the evaluation scope.","section":"Table 7"},{"comment":"The nested entity examples all use dates; no example is given for nested entities involving other types, and the statement that 'inside entity is more important than outside entity' is not operationalized in the evaluation.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is not yet ready for publication in its current form. The entity-count inconsistency is directly observable and must be fixed. I would also like to see the dataset made available and some form of inter-annotator agreement reported. The circularity concern raised in the stress test does not apply: the evaluations are conducted on a held-out test split, so no circularity is involved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a dataset paper, not a method paper. The genuinely new artifact is a manually annotated Vietnamese COVID-19 NER set with 11 entity types, five date subtypes, and up to four nesting levels. If the data is real and usable, that is a real resource for Vietnamese clinical NLP. But the paper cannot currently support that \"if\": the abstract and Section 3 say 11,128 entities, while Table 2 sums to 17,818 (Level-1 alone is 15,870), and no dataset link is provided. The contribution is unverifiable in its present form.\n\nWhat it does well: the entity schema is thoughtfully designed with medical input, with separate date types for symptom onset, hospitalization, isolation, positive test, and contact. That is a sensible fine-grained scheme for contact tracing. The annotation process is described in three phases, and the baseline experiments are standard and honestly reported as averages over five runs. The cross-sentence context idea is taken from Luoma and Pyysalo and applied consistently; it gives a small but consistent improvement, which is plausible. PhoBERT-large-CRF with cross-sentence context reaching 82.34 F1 at Level-1 and 67.23 at Level-2 is a believable result for this task.\n\nSoft spots, in order of severity. First, the entity count inconsistency is load-bearing because the dataset is the paper's only real contribution. It may be that 11,128 refers to something like non-nested mentions or a pre-merge count, but the paper never says so, and 15,870 Level-1 entities already exceed 11,128, so no obvious reading reconciles the numbers. Second, there is no release link and no inter-annotator agreement; label quality is asserted but not demonstrated. Third, Tables 3–6 give only means over five runs, with no standard deviations or significance tests, so the claimed best-model ordering is suggestive, not established. These are all fixable, and fixing the first two is necessary.\n\nWho should read it: Vietnamese NLP researchers looking for a COVID-19 NER benchmark, and people working on nested NER in lower-resource settings. The experiments section will not teach them much. If the authors release the data and correct the statistics, this becomes a modest but citable resource. I would not cite it now, and I would not bring it to a reading group yet. But it deserves peer review rather than desk rejection: the resource is potentially useful, and a good referee can force the release and the numbers to be fixed.","headline":"A potentially useful Vietnamese COVID-19 nested NER dataset is presented, but the entity counts do not add up and the data is not released, so the contribution is currently unverifiable.","tokens_in":10417,"tokens_out":3296,"would_cite":false,"duration_ms":30146,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"These experiments establish a manually annotated Vietnamese COVID-19 NER corpus of 10,271 sentences and 11,128 entities, with PhoBERT-large-CRF plus cross-sentence context scoring best at 82.34 F1 for top-level entities.","keywords":["nested named entity recognition","Vietnamese NLP","COVID-19 dataset","contact tracing","PhoBERT","clinical text mining","entity annotation","cross-sentence context"],"falsifier":"Have independent Vietnamese-speaking annotators re-annotate a random sample of the test set with the same eleven-type schema and measure their agreement with the released labels; if agreement is low, for example below κ≈0.6, the reported F1 scores do not measure a stable ground truth. A secondary quick check is to reconcile the entity count: the abstract says 11,128 entities while Table 2 sums to 17,818.","tokens_in":9445,"feed_emoji":"🦠","tokens_out":10495,"duration_ms":84288,"temperature":0.7,"pith_summary":"The paper aims to build an automatic information-extraction aid for Vietnam's COVID-19 contact-tracing effort, where patient interviews and movement reports are currently processed by hand. To that end, it introduces a manually annotated nested named-entity recognition dataset for Vietnamese: 10,271 sentences and 11,128 entities, labelled with eleven entity types covering patient demographics, five date types, contact persons, and locations, with nesting up to four levels. The paper then compares BiLSTM-CRF and PhoBERT-based models, adding a CRF layer and cross-sentence context, and reports that PhoBERT-large-CRF with cross-sentence context gives the best test F1 of 82.34 for outer-level entities and 67.23 for the second nesting level. A sympathetic reader would take away that automatic extraction of the dates and contacts needed for contact tracing is feasible in Vietnamese, at least for the high-frequency fields.","feed_headline":"Nested Vietnamese COVID-19 NER dataset lifts top F1 to 82.3","feed_subtitle":"With eleven entity types and four nesting levels, PhoBERT-large-CRF reaches a best F1 of 82.3.","key_machinery":"The load-bearing object is the dataset itself: nested XML-style annotations with eleven entity types and up to four levels, converted into CoNLL-style files for training. The main mechanism for handling nesting is the joint tag, which merges a token's level tags into one composite tag such as HOS DATE+ISO DATE, allowing a standard sequence tagger to predict nested structure. The other mechanism that moves the scores is cross-sentence context: each input example begins with the target sentence and then packs the following sentences up to the model's length limit, which lets the model resolve date entities that are only specified in later sentences, such as an isolation date referred to as 'same day'.","core_discovery":"The central claim is that a manually produced Vietnamese COVID-19 NER dataset, built from medical reports, news sites, and public sources and cleaned and reviewed by hand, supports an eleven-type nested entity schema, and that a fine-tuned PhoBERT-large model with a CRF head and cross-sentence context is the strongest of the tested recognizers. On the test split, this model reaches F1 82.34 for Level-1 entities and 67.23 for Level-2 entities; per-type results are near-ceiling for NAME, AGE, GENDER, and ADDRESS (F1 95.65 to 99.77), while the overlapping date types score lower but still usable (SYM DATE 74.85, POS DATE 76.18, HOS DATE 77.45, ISO DATE 79.85). The paper treats the nesting order as carrying medical importance, with inner entities more important than outer ones, and attributes the best results to the combination of the pretrained model, the CRF layer, and cross-sentence context rather than to any single component.","pith_inferences":["Editorial inference: if annotation guidelines and inter-annotator agreement are published, the same eleven-type schema could transfer to other Vietnamese clinical documents, not just COVID-19 reports.","Editorial inference: the gain from cross-sentence context suggests that explicitly resolving temporal references such as 'same day' would close much of the remaining error on ISO, HOS, and POS dates.","Editorial inference: because nesting order encodes medical priority, downstream applications could treat the hierarchy itself as structured output, generating for each patient a timeline of contact, quarantine, hospitalization, and confirmation dates.","Editorial inference: a public release of the corpus with its nesting levels would let other researchers test whether the reported F1 numbers hold under independent annotation audits."],"forward_implications":["A reliable automatic extractor at these F1 levels could cut the hours spent manually reading each patient report, because high-confidence fields like name, age, gender, and address can be auto-filled and only ambiguous dates routed to human review.","The dataset gives Vietnamese NLP a benchmark for nested NER with a clinical and epidemiological schema, complementing general-purpose Vietnamese NER resources.","The five-way split of date entities makes the task more demanding than ordinary NER but also more directly useful for reconstructing patient timelines.","The reported gap between Level-1 and Level-2 F1, about 15 points, quantifies how much harder nested recognition is on this schema.","The joint-tag method shows that nested recognition can be handled by a single sequence tagger without a separate nested-decoding architecture."],"supporting_citations":[{"why":"Supplies the prior Vietnamese NER dataset and annotation tradition that the new corpus extends.","marker":"[Nguyen and Vu, 2016]"},{"why":"VLSP 2018 shared task, the source of the nested-entity annotation style the paper says it follows.","marker":"[Nguyen et al., 2019]"},{"why":"PhoNER COVID19, the prior Vietnamese COVID-19 NER dataset whose entity types the paper contrasts with its eleven types.","marker":"[Truong et al., 2021]"},{"why":"PhoBERT, the pretrained Vietnamese language model used as the backbone of the best-performing systems.","marker":"[Nguyen and Nguyen, 2020]"},{"why":"Cross-sentence context construction that the paper credits with the largest gains.","marker":"[Luoma and Pyysalo, 2020]"},{"why":"BiLSTM-CRF architecture, the non-pretrained baseline and the CRF head used with PhoBERT.","marker":"[Huang et al., 2015]"},{"why":"Joint-tag representation for nested entities, the method used to fuse level tags into one tag.","marker":"[Minh, 2018]"},{"why":"CoNLL 2003 format into which the annotated data were converted for training.","marker":"[Sang and Meulder, 2003]"}],"fun_headline_variants":["Nested Vietnamese COVID-19 NER hits 82.3 F1 with PhoBERT-CRF","11-entity nested NER for Vietnam COVID-19 scores 82.3 F1","Manual nested NER dataset for Vietnamese COVID-19 reaches 82.3","PhoBERT-CRF lifts Vietnamese nested COVID-19 NER to 82.3 F1"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the manual annotations being correct and consistent, but the paper does not report any agreement measure between annotators or an independent check of the labels.","fun_headline_variants_meta":{"raw":{"variants":["Nested Vietnamese COVID-19 NER hits 82.3 F1 with PhoBERT-CRF","11-entity nested NER for Vietnam COVID-19 scores 82.3 F1","Manual nested NER dataset for Vietnamese COVID-19 reaches 82.3","PhoBERT-CRF lifts Vietnamese nested COVID-19 NER to 82.3 F1"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1264,"prompt_tokens":863,"completion_tokens":401,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":306}},"tokens_in":479,"tokens_out":401,"duration_ms":3747,"temperature":1.0,"reasoning_tokens":306,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:38:21.800589+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent Vietnamese-speaking annotators re-annotate a random sample of the test set with the same eleven-type schema and measure their agreement with the released labels; if agreement is low, for example below κ≈0.6, the reported F1 scores do not measure a stable ground truth. A secondary quick check is to reconcile the entity count: the abstract says 11,128 entities while Table 2 sums to 17,818.","supporting_citations":[{"cited_title":"Vlsp shared task: Named entity recognition.Journal of Computer Science and Cy- bernetics, 34(4):283–294,","cited_arxiv_id":null,"evidence_quote":"VLSP 2018 shared task, the source of the nested-entity annotation style the paper says it follows."},{"cited_title":"Vlsp 2016 shared task: Named entity recognition","cited_arxiv_id":null,"evidence_quote":"Supplies the prior Vietnamese NER dataset and annotation tradition that the new corpus extends."},{"cited_title":"COVID-19 Named Entity Recog- nition for Vietnamese","cited_arxiv_id":null,"evidence_quote":"PhoNER COVID19, the prior Vietnamese COVID-19 NER dataset whose entity types the paper contrasts with its eleven types."},{"cited_title":"PhoBERT: Pre-trained language models for Vietnamese","cited_arxiv_id":null,"evidence_quote":"PhoBERT, the pretrained Vietnamese language model used as the backbone of the best-performing systems."},{"cited_title":"Exploring cross-sentence contexts for named entity recognition with bert","cited_arxiv_id":null,"evidence_quote":"Cross-sentence context construction that the paper credits with the largest gains."},{"cited_title":"A Feature-Based Model for Nested Named-Entity Recognition at VLSP-2018 NER Evaluation Campaign","cited_arxiv_id":"1803.08463","evidence_quote":"Joint-tag representation for nested entities, the method used to fuse level tags into one tag."},{"cited_title":"Tjong Kim Sang and Fien De Meulder","cited_arxiv_id":null,"evidence_quote":"CoNLL 2003 format into which the annotated data were converted for training."}],"review_version":1}