A new dataset of 63,471 Sinhala YouTube comments on music videos and 964 derived stop-words is presented and compared against general Sinhala corpora.
NSINA: A News Corpus for Sinhala
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The introduction of large language models (LLMs) has advanced natural language processing (NLP), but their effectiveness is largely dependent on pre-training resources. This is especially evident in low-resource languages, such as Sinhala, which face two primary challenges: the lack of substantial training data and limited benchmarking datasets. In response, this study introduces NSINA, a comprehensive news corpus of over 500,000 articles from popular Sinhala news websites, along with three NLP tasks: news media identification, news category prediction, and news headline generation. The release of NSINA aims to provide a solution to challenges in adapting LLMs to Sinhala, offering valuable resources and benchmarks for improving NLP in the Sinhala language. NSINA is the largest news corpus for Sinhala, available up to date.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Linguistic Analysis of Sinhala YouTube Comments on Sinhala Music Videos: A Dataset Study
A new dataset of 63,471 Sinhala YouTube comments on music videos and 964 derived stop-words is presented and compared against general Sinhala corpora.