A manually annotated Vietnamese COVID-19 NER dataset with 11 entity types and up to four nesting levels, plus BiLSTM and PhoBERT baselines where PhoBERT-large-CRF with cross-sentence context achieved the highest F1.
A Feature-Based Model for Nested Named-Entity Recognition at VLSP-2018 NER Evaluation Campaign
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this report, we describe our participant named-entity recognition system at VLSP 2018 evaluation campaign. We formalized the task as a sequence labeling problem using BIO encoding scheme. We applied a feature-based model which combines word, word-shape features, Brown-cluster-based features, and word-embedding-based features. We compare several methods to deal with nested entities in the dataset. We showed that combining tags of entities at all levels for training a sequence labeling model (joint-tag model) improved the accuracy of nested named-entity recognition.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Nested Named-Entity Recognition on Vietnamese COVID-19: Dataset and Experiments
A manually annotated Vietnamese COVID-19 NER dataset with 11 entity types and up to four nesting levels, plus BiLSTM and PhoBERT baselines where PhoBERT-large-CRF with cross-sentence context achieved the highest F1.