Pith. sign in

REVIEW 2 cited by

KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.07251 v1 pith:YBUSH5II submitted 2024-12-10 cs.CL

classification cs.CL
keywords modelsculturallanguagebenchkoreankultureassesscapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models have exhibited significant enhancements in performance across various tasks. However, the complexity of their evaluation increases as these models generate more fluent and coherent content. Current multilingual benchmarks often use translated English versions, which may incorporate Western cultural biases that do not accurately assess other languages and cultures. To address this research gap, we introduce KULTURE Bench, an evaluation framework specifically designed for Korean culture that features datasets of cultural news, idioms, and poetry. It is designed to assess language models' cultural comprehension and reasoning capabilities at the word, sentence, and paragraph levels. Using the KULTURE Bench, we assessed the capabilities of models trained with different language corpora and analyzed the results comprehensively. The results show that there is still significant room for improvement in the models' understanding of texts related to the deeper aspects of Korean culture.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MyCulture: Exploring Malaysia's Diverse Culture under Low-Resource Language Constraints

    cs.CL 2025-08 reject novelty 6.0 of 10

    MyCulture, a new Malay-language cultural benchmark, shows LLM accuracy drops by at least 17% when multiple-choice questions are converted to an open-ended format.

  2. BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

    cs.LG 2025-05 conditional novelty 5.0 of 10

    BenchHub is an automatically categorized, customizable LLM benchmark suite covering 303K questions across 38 benchmarks in English and Korean.

Pith tools