Pith. sign in

REVIEW 1 cited by

KOBEST: Korean Balanced Evaluation of Significant Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.04541 v1 pith:FHFB6XWO submitted 2022-04-09 cs.CL

classification cs.CL
keywords koreantasksbenchmarksevaluationmodelsbalancedbenchmarkdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A well-formulated benchmark plays a critical role in spurring advancements in the natural language processing (NLP) field, as it allows objective and precise evaluation of diverse models. As modern language models (LMs) have become more elaborate and sophisticated, more difficult benchmarks that require linguistic knowledge and reasoning have been proposed. However, most of these benchmarks only support English, and great effort is necessary to construct benchmarks for other low resource languages. To this end, we propose a new benchmark named Korean balanced evaluation of significant tasks (KoBEST), which consists of five Korean-language downstream tasks. Professional Korean linguists designed the tasks that require advanced Korean linguistic knowledge. Moreover, our data is purely annotated by humans and thoroughly reviewed to guarantee high data quality. We also provide baseline models and human performance results. Our dataset is available on the Huggingface.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Opt.Gear Technical Report

    cs.CL 2026-08 conditional novelty 4.0 of 10

    Opt.Gear is a family of efficient on-device language models using a ConvKV-gated mixer with sparse attention, trained on 0.5T tokens without distillation, claiming up to 4.9x NPU speedups and 20 TPS on a Cortex-M7.

Pith tools