A new Vietnamese dataset for text segmentation and multiple-choice reading comprehension, with benchmarks showing multilingual BERT models lead on both tasks.
VnCoreNLP: A Vietnamese Natural Language Processing Toolkit
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We present an easy-to-use and fast toolkit, namely VnCoreNLP---a Java NLP annotation pipeline for Vietnamese. Our VnCoreNLP supports key natural language processing (NLP) tasks including word segmentation, part-of-speech (POS) tagging, named entity recognition (NER) and dependency parsing, and obtains state-of-the-art (SOTA) results for these tasks. We release VnCoreNLP to provide rich linguistic annotations to facilitate research work on Vietnamese NLP. Our VnCoreNLP is open-source and available at: https://github.com/vncorenlp/VnCoreNLP
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
A Vietnamese Dataset for Text Segmentation and Multiple Choices Reading Comprehension
A new Vietnamese dataset for text segmentation and multiple-choice reading comprehension, with benchmarks showing multilingual BERT models lead on both tasks.