REVIEW 4 major objections 4 minor 21 references
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Language models know Korean superstition facts but fail to apply them in real-world advice, a new 247-question benchmark shows.
desk verdict Useful benchmark with a real gap: the headline claim conflates mentioning the superstition with applying it, but the resource itself deserves review and a validity pass. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is Nunchi-Bench itself, built from 31 superstition topics that survived a fill-in-the-blank quiz with 33 Korean participants. Each topic appears as a factual multiple-choice question and as open-ended Trap and Interpretation questions, and each open-ended item comes in four versions: Korean or English, with the people in the scenario either explicitly identified as Korean (Specified) or left unspecified (Neutral). The scoring machinery is a four-level rubric with scores of 2, 1, 0, and -1, applied by GPT-4 Turbo after iterative prompt refinement, with reported exact agreement with human raters of 90 percent on Trap questions and 88.3 percent on Interpretation questions; the rubric works by rewarding responses that name the specific superstition and penalizing fabricated cultural content.
What would settle it
Re-score the complete set of open-ended responses, or a large random sample rather than 30, with independent human raters using the published rubric and compare exact agreement with GPT-4 Turbo; if agreement falls far below the reported 88 to 90 percent, or if the ordering of Specified over Neutral versions reverses under human scoring, the paper's central claim about cultural framing is not settled.
Extended reading notes
Core claim
The paper reports a knowledge-application gap: on the factual multiple-choice items most models pick the right superstition, yet on the corresponding Trap questions, where a user asks whether a culturally charged action is okay (for example, writing a name in red or serving seaweed soup on exam day), many models give generic, culturally unaware advice. Scores from a four-level rubric, which gives two points for explicitly invoking the relevant superstition, one for general cultural awareness, zero for no cultural consideration, and minus one for hallucinated cultural content, show that adding the word 'Korean' to the scenario raises scores more than translating the prompt into Korean. The authors read this as evidence that cultural reasoning is not a by-product of language matching, and that Korean-focused training data, rather than mere bilingual ability, drives performance on Korean cultural tasks.
Load-bearing premise
The benchmark's central comparison depends on trusting the GPT-4 Turbo scoring rubric, whose human alignment was checked on only 30 responses per task in its final phase; if that judge is biased, the findings about context, language, and training data change.
Editorial extensions
If this is right
- Factual multiple-choice performance is not a reliable proxy for cultural competence; scenario-based tasks are needed to expose application failures.
- Explicitly framing a prompt with its cultural context, such as mentioning a Korean roommate, is a cheap and effective way to get culturally aware output, and it helps more than translating the prompt into Korean.
- Models trained primarily on Korean data, such as HyperClova-X and EXAONE, behave differently on Korean prompts, so training-data language composition shapes cultural reasoning.
- Newer models evaluated in the paper's appendix, including GPT-4.5 and Gemini 2.5 Pro, improve overall but still score far lower on Neutral than on Specified versions, indicating the context gap persists.
- Because the Trap and Interpretation formats are culturally agnostic, the same benchmark template can be transplanted to other superstition traditions beyond Korea.
Reading between the lines
- A test the paper does not run: holding response length and verbosity constant across Specified and Neutral prompts. Since the rubric rewards mentioning a superstition, long and elaborate answers could inflate Specified scores independently of genuine cultural reasoning.
- The paper's own correlation analysis shows that factual MCQ scores do not predict open-ended performance within the same superstition topic; I read this as an argument that future cultural benchmarks should include open-ended tasks, not just multiple-choice items.
- Because the paper assigns zero points to refusals, a model that politely declines to comment on a superstition is indistinguishable from one that gives generic advice; adding a separate refusal category would alter the model ranking.
- The evaluator is part of the measurement system: if GPT-4 Turbo shares the cultural blind spots of the models it scores, the reported 88 to 90 percent alignment on a 30-response sample may not hold at larger scale.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Nunchi-Bench, a benchmark of 247 questions on Korean superstitions spanning 31 topics, with three task formats: multiple-choice factual questions, open-ended Trap questions that ask for culturally sensitive advice, and Interpretation questions that require explaining culturally meaningful reactions. The benchmark is provided in Korean and English, with Specified and Neutral versions of the open-ended tasks. The authors evaluate 12 private and open-source LLMs using GPT-4 Turbo as an automated judge applying a 0/1/2/-1 rubric, and report four main findings: models generally know superstition facts but struggle to apply them in scenarios; explicit cultural framing improves performance; prompt language alone is less effective than cultural framing; and language-specific training matters. The benchmark and leaderboard are publicly released.
Significance. If the central claims are valid, Nunchi-Bench is a valuable resource for cultural-reasoning evaluation: it moves beyond factual MCQ-style benchmarks such as CLIcK and KorNAT to open-ended scenarios that test situated application, and it is publicly released with a leaderboard. The paper also introduces a multi-phase human-alignment procedure for an LLM-based judge, which is a useful methodological contribution. However, the validity of the headline 'knows facts but cannot apply them' claim hinges on whether the scoring rubric measures culturally appropriate behavior rather than merely explicit mention of a superstition; the current evidence for that construct validity is thin, and the paper's own Limitations section concedes the absence of gold reference responses. The benchmark artifact itself is a credible and useful contribution, but the empirical conclusions need strengthening.
major comments (4)
- [§3.2, Appendix C] The 0-point rubric criterion ('The response does not mention cultural differences') means that advice which is behaviorally appropriate but does not verbalize the superstition—e.g., suggesting a different color for the name without saying 'red ink means death in Korea'—receives the same score as advice that ignores the superstition entirely. The paper's headline claim that models 'recognize factual information but struggle to apply it' is computed from scores produced by this rubric, so it partly measures whether models explicitly state superstitions rather than whether they apply them. The Appendix C alignment study shows only that human raters can reproduce GPT-4 Turbo's application of this rubric on 30 responses per task; it does not establish that 0-point responses are culturally inappropriate. The Limitations section's own admission that 'the absence of gold reference responses may still affect reliability' underscores this construct-validity gap. I ask for a direct validity check: human judges should rate the cultural appropriateness of the actual advice/interpretation independently of whether the superstition is mentioned, and the rubric should be revised or supplemented accordingly.
- [Appendix C, Table 12] The calibration evidence for the automated judge rests on 30 responses per task in the final phase, with 27/30 (90.0%) agreement on Trap and 26/30 (88.3%) on Interpretation. For n=30, the 95% confidence interval for 90% agreement is roughly 73–98%, so the reported alignment is compatible with substantial disagreement. In addition, only exact per-response score agreement is reported; there is no chance-corrected measure (e.g., Cohen's kappa) and no report of human inter-annotator agreement, making it hard to know how much of the agreement is trivial (e.g., both raters assigning 0 to generic responses). Because all open-ended scores and all headline findings depend on this judge, I ask for a larger alignment sample, per-category agreement, and chance-corrected statistics.
- [§3.3, §4] The central claim that explicit cultural framing contributes more than prompt language (English+Specified > Korean+Neutral) is presented as a qualitative pattern over raw aggregate scores in Table 4 and Figure 2, without any paired significance test across models, questions, or topics. Likewise, the claim about language-specific training (HyperClova-X and EXAONE leading in Korean versions) rests on descriptive Table 4 and Figure 8 patterns. With 12 models and per-topic response samples, these comparisons need confidence intervals, paired tests (e.g., Wilcoxon or bootstrap by model/topic), or at least per-condition score distributions; without them, the paper's three main findings are not statistically supported.
- [Appendix D, Figures 6–7] The correlational analyses use n=12 model-level observations and run 16 MCQ-vs-Trap/Interpretation tests (Tables 13 and 14) without multiple-comparison control; the statement that 'no other significant correlations were found' is therefore weak evidence for the absence of relationships. The within-topic Spearman correlations in Figure 7 fluctuate widely and are based on very few observations per topic, so the conclusion that MCQs 'fail to capture deeper contextual understanding' is not established by this analysis. I recommend treating this part as exploratory or adding appropriate correction and power analysis.
minor comments (4)
- [Figure 2, Table 4] Several headers run words together (e.g., 'Gemini1.5 ProClaude 3Opus' and 'EXAON 3.0'), making the figure and table hard to read; please fix spacing and spelling.
- [Table 5] 'GPT3.5 Turbo' should be 'GPT-3.5 Turbo', and 'in your feature on light and nutritious foods' appears to be a typo (likely 'feast' or 'meal').
- [Throughout] The model name is written inconsistently as HyperCLOV A-X, HyperCLOVA-X, and HyperClova-X across the paper; please unify the spelling.
- [Appendix A, Table 8] The footnote explaining the asterisk markers is easy to miss because the asterisks appear inside the score column; please move the markers next to the topic IDs and clarify which IDs are excluded entirely versus excluded only from Trap questions.
Circularity Check
No circular derivation: benchmark results are empirical outcomes of an explicit rubric, with no fitted parameter renamed as prediction.
full rationale
Nunchi-Bench is an empirical benchmark paper rather than a formal derivation, so there is no derivation chain whose conclusion is equivalent to its inputs. The central findings—MCQ accuracy versus open-ended scores, the Specified/Neutral contrast, and prompt-language effects—are direct measurements under the explicitly stated rubric in Section 3.2. That rubric operationalizes cultural consideration as explicit mention of the relevant superstition (0 points: 'The response does not mention cultural differences'; 2 points: the response 'explicitly acknowledges and incorporates the specific Korean superstition'). This is a construct-validity choice, and the paper itself flags the absence of gold reference responses and possible evaluator bias in the Limitations section. A construct-validity concern is not circularity: the open-ended scores are not forced by construction, the Specified/Neutral comparison is an empirical contrast, and no fitted parameter is later renamed as a prediction. The paper contains no load-bearing self-citations, imported uniqueness theorems, or ansatz smuggled in via prior work. The Appendix C alignment study provides human-anchor evidence for the GPT-4 Turbo evaluator, albeit on a small final sample (30 responses per task), which is a reliability limitation rather than a circular step. Therefore no specific circular reduction can be exhibited, and the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (4)
- topic_selection_threshold =
50%
- question_relevance_threshold =
2 of 3 evaluators
- scoring_rubric_weights =
2, 1, 0, -1
- evaluator_alignment_sample_size =
30 responses (phase 3)
assumptions (4)
- domain assumption The 31 selected superstitions and the human panels (33 quiz takers, 12 relevance raters) are representative of Korean superstition culture.
- domain assumption GPT-4 Turbo with the authors' rubric produces scores that are a valid proxy for human cultural-sensitivity judgments.
- domain assumption Refusals or non-responses are failures (0 points), equivalent to responses lacking cultural consideration.
- domain assumption Weighted sums across questions and versions are comparable despite different maximum scores (MCQ max 31, Trap/Interpretation max roughly 184-189).
Cite this review
Pith. "Pith review of Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition." pith.science (2026). https://pith.science/paper/TO54FRNJ
@misc{pith2026250704014,
author = {Pith},
title = {Pith review of: Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition},
year = {2026},
howpublished = {\url{https://pith.science/paper/TO54FRNJ}},
note = {Machine review of arXiv:2507.04014}
}
read the original abstract
As large language models (LLMs) become key advisors in various domains, their cultural sensitivity and reasoning skills are crucial in multicultural environments. We introduce Nunchi-Bench, a benchmark designed to evaluate LLMs' cultural understanding, with a focus on Korean superstitions. The benchmark consists of 247 questions spanning 31 topics, assessing factual knowledge, culturally appropriate advice, and situational interpretation. We evaluate multilingual LLMs in both Korean and English to analyze their ability to reason about Korean cultural contexts and how language variations affect performance. To systematically assess cultural reasoning, we propose a novel evaluation strategy with customized scoring metrics that capture the extent to which models recognize cultural nuances and respond appropriately. Our findings highlight significant challenges in LLMs' cultural reasoning. While models generally recognize factual information, they struggle to apply it in practical scenarios. Furthermore, explicit cultural framing enhances performance more effectively than relying solely on the language of the prompt. To support further research, we publicly release Nunchi-Bench alongside a leaderboard.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Badr AlKhamissi, Muhammad ElNokrashy, Mai Alkhamissi, and Mona Diab. 2024. https://doi.org/10.18653/v1/2024.acl-long.671 Investigating cultural alignment of large language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12404--12422, Bangkok, Thailand. Association for Compu...
-
[2]
AI Anthropic. 2024. https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf The claude 3 model family: Opus, sonnet, haiku . In Anthropic Model Card, 1
work page 2024
-
[3]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...
arXiv 2020
-
[4]
Yu Ying Chiu, Liwei Jiang, Maria Antoniak, Chan Young Park, Shuyue Stella Li, Mehar Bhatia, Sahithya Ravi, Yulia Tsvetkov, Vered Shwartz, and Yejin Choi. 2024. http://arxiv.org/abs/2404.06664 Culturalteaming: Ai-assisted interactive red-teaming for challenging llms' (lack of) multicultural knowledge
arXiv 2024
-
[5]
Mohsen Fayyaz, Fan Yin, Jiao Sun, and Nanyun Peng. 2024. https://arxiv.org/abs/2407.00219 Evaluating human alignment and model faithfulness of llm rationale . ArXiv, abs/2407.00219
arXiv 2024
-
[6]
Gemini-Team-Google. 2024. http://arxiv.org/abs/2403.05530 Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
arXiv 2024
-
[7]
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders S gaard. 2022. https://doi.org/10.18653/v1/2022.acl-long.482 Challenges and strategies in cross-cultural NLP . In Pro...
-
[8]
Eunsu Kim, Juyoung Suk, Philhoon Oh, Haneul Yoo, James Thorne, and Alice Oh. 2024 a . Click: A benchmark dataset of cultural and linguistic intelligence in korean. arXiv preprint arXiv:2403.06412
arXiv 2024
Show all 21 references
-
[9]
Jeongwook Kim, Taemin Lee, Yoonna Jang, Hyeonseok Moon, Suhyune Son, Seungyoon Lee, and Dongjun Kim. 2024 b . Kullm3: Korea university large language model 3. https://github.com/nlpai-lab/kullm
2024
-
[10]
KT. 2023. https://huggingface.co/KT-AT/midm-bitext-S-7B-inst-v1 Mi:dm: Kt bilingual (korean,english) generative pre-trained transformer . https://genielabs.ai
2023
-
[11]
Jiyoung Lee, Minwoo Kim, Seungho Kim, Junghwan Kim, Seunghyun Won, Hwaran Lee, and Edward Choi. 2024. https://doi.org/10.18653/v1/2024.findings-acl.666 K or NAT : LLM alignment benchmark for K orean social values and common knowledge . In Findings of the Association for Comput...
2024 doi
-
[12]
LG-AI-Research. 2024. Exaone 3.0 7.8b instruction tuned language model. arXiv preprint arXiv:2408.03541
2024
-
[13]
Chen Liu, Fajri Koto, Timothy Baldwin, and Iryna Gurevych. 2024. https://doi.org/10.18653/v1/2024.naacl-long.112 Are multilingual LLM s culturally-diverse reasoners? an investigation into multicultural proverbs and sayings . In Proceedings of the 2024 Conference of the North A...
2024 doi
-
[14]
Llama-Team. 2024. http://arxiv.org/abs/2407.21783 The llama 3 herd of models
2024 arXiv
-
[15]
Junho Myung, Nayeon Lee, Yi Zhou, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Perez-Almendros, Abinew Ali Ayele, et al. 2024. Blend: A benchmark for llms on everyday knowledge in diverse cultures and languages. arXiv preprint arXiv:2406.09948
2024 arXiv
-
[16]
Qwen-Team. 2024. https://qwenlm.github.io/blog/qwen2.5/ Qwen2.5: A party of foundation models
2024
-
[17]
Naver HyperCLOVA X Team. 2024. http://arxiv.org/abs/2404.01954 Hyperclova x technical report
2024 arXiv
-
[18]
Robert Vacareanu, Anurag Pratik, Evangelia Spiliopoulou, Zheng Qi, Giovanni Paolini, Neha Anna John, Jie Ma, Yassine Benajiba, and Miguel Ballesteros. 2024. https://arxiv.org/abs/2405.00204 General purpose verification for chain of thought prompting . ArXiv, abs/2405.00204
2024 arXiv
-
[19]
Yuhang Wang, Yanxu Zhu, Chao Kong, Shuyu Wei, Xiaoyuan Yi, Xing Xie, and Jitao Sang. 2024. https://doi.org/10.18653/v1/2024.c3nlp-1.1 CDE val: A benchmark for measuring the cultural dimensions of large language models . In Proceedings of the 2nd Workshop on Cross-Cultural Cons...
2024 doi
-
[20]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...
-
[21]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.