pith. sign in

arxiv: 2012.15621 · v3 · pith:XZMJEZOHnew · submitted 2020-12-31 · 💻 cs.CL

Open Korean Corpora: A Practical Report

classification 💻 cs.CL
keywords koreancorporalistopenresearchthenadvertisedavailability
0
0 comments X
read the original abstract

Korean is often referred to as a low-resource language in the research community. While this claim is partially true, it is also because the availability of resources is inadequately advertised and curated. This work curates and reviews a list of Korean corpora, first describing institution-level resource development, then further iterate through a list of current open datasets for different types of tasks. We then propose a direction on how open-source dataset construction and releases should be done for less-resourced languages to promote research.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.