REVIEW 6 major objections 5 minor 25 references
KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding
T0 review · 6 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read KFinEval-Pilot, a 1,145-question Korean financial benchmark, claims to reveal LLM gaps in knowledge, legal reasoning, and safety.
desk verdict A useful pilot benchmark for Korean financial LLM evaluation, but the model rankings rest on an unvalidated LLM judge and need major revision before the numbers can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the benchmark's three-part task design. Financial knowledge uses 377 multiple-choice questions with four options drawn from official Korean financial glossaries. Financial reasoning uses 209 open-ended questions that pair statutory excerpts with expert-written chain-of-thought rationales, forcing models to perform multi-step legal inference rather than numeric calculation. Financial toxicity uses 484 adversarial red-teaming prompts covering fraud, illicit flows, and privacy breaches; models are judged on whether they refuse or give harmful help. Items are generated by GPT-4o through staged prompts, checked for answerability and logical validity, then pass two rounds of expert review; open-ended outputs are scored by GPT-o1 as a judge using rubric-based prompts. This design converts 'Korean financial competence' into separate measurable claims about factual recall, procedural legal reasoning, and safety alignment.
What would settle it
Take a random sample of 100 reasoning and 100 toxicity responses from the evaluated models, have Korean financial experts score them with the same rubrics, and compare the scores with GPT-o1's judge scores; large disagreement or systematic favoritism toward same-family models would invalidate the reported rankings and the accuracy-safety trade-off.
Extended reading notes
Core claim
The central claim is that a Korean-specific benchmark can expose capabilities that generic financial benchmarks overlook, and that current LLMs are uneven across the three measured abilities. Quantitatively, GPT-o1 reaches 71.35% on financial knowledge, GPT-o3-mini scores 7.66 on legal reasoning, while Qwen2.5-7B-Instruct is the strongest open model on both (64.19% and 6.30). On toxicity, lower is safer, and scores span from 1.46 (GPT-o3-mini) to 9.56 (Qwen2.5-7B-Instruct), with the authors' finance-tuned 8B model in the middle at 6.96. The paper interprets this spread as evidence of an accuracy-safety trade-off across model families and as a demonstration that evaluation must be grounded in local language, regulation, and realistic abuse scenarios.
Load-bearing premise
The benchmark's reasoning and safety rankings depend on trusting GPT-o1 as the judge, with no human validation reported, even though that judge comes from the same model family as several systems it grades.
Editorial extensions
If this is right
- Korean financial institutions can use the benchmark to compare models on regulatory reasoning and refusal behavior before deployment, rather than relying on English financial QA scores.
- Model selection should weigh safety against accuracy: in this evaluation the strongest open-source knowledge model is also the least safe, and the best proprietary reasoning model is the safest.
- Domain-specific post-training helps on knowledge but does not automatically harden safety, since the authors' finance-tuned model scores competitively on knowledge yet only mid-pack on toxicity.
- The reasoning and toxicity task formats give other non-English financial sectors a template for building evaluation sets tied to their own laws and fraud patterns.
Reading between the lines
- Beyond the paper, the same construct-and-verify pipeline is portable: any country with a distinct regulatory code could build an analogous benchmark from its own statutes and fraud cases, enabling cross-country comparisons of financial LLM readiness.
- Because the 209-item reasoning set is small, rank differences of about one point between open models may not be stable; a larger item pool could reshuffle the open-source ordering.
- The use of a single judge from the same vendor as several evaluated models is a circularity risk; a multi-judge or human-panel scoring pass would be a natural follow-up that the paper does not provide.
- If the toxicity items are released, they could serve a second purpose as red-team training data for safety fine-tuning, not only as an evaluation set; the paper does not propose this use.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces KFinEval-Pilot, a Korean financial language understanding benchmark with 1,145 multiple-choice, reasoning, and toxicity items, constructed through GPT-4o-assisted generation and expert validation. The authors evaluate commercial and open-weight LLMs and report that proprietary models, especially GPT-o3-mini and GPT-o1, lead in reasoning and safety, while Qwen models lead among open models. The paper also introduces KFTC-8B-Finance-Instruct, a domain-adapted 8B model, and shows it is competitive on knowledge but not on reasoning.
Significance. If the evaluation methodology were properly validated, KFinEval-Pilot would be a valuable resource: it fills a real gap by targeting Korean financial language understanding, includes procedural reasoning grounded in legal statutes, covers a toxicity dimension under-explored in prior financial benchmarks, and is built on publicly documented Korean regulatory sources with expert review. The paper also usefully makes its generation and evaluation prompts explicit in Appendix A. However, the quantitative claims in Tables 8 and 9 currently rest on an unvalidated LLM judge and an undefined aggregation of scores, which substantially weakens the benchmark's diagnostic value as it stands. The dataset is not publicly released and the evaluation code is not provided, limiting reproducibility.
major comments (6)
- [Table 8, Appendix A (Table 10)] Table 8 reports a single 'Reasoning' score per model, but the Appendix A evaluation prompt (Table 10) defines six distinct 1–10 sub-scores (정합성, 일관성, 정확성, 완전성, 추론성, 전체품질). No aggregation rule is stated anywhere in the manuscript. Because these six dimensions are not necessarily of equal weight, the reported reasoning scores cannot be reproduced or interpreted. Please specify the aggregation formula (e.g., average, weighted average) and, ideally, report the sub-scores or their summary statistics.
- [§4.2, Tables 8–9] All open-ended reasoning and toxicity scores in Tables 8 and 9 are produced by a single LLM judge, GPT-o1-2025-04-04, with no human validation, no inter-annotator agreement, and no judge-consistency analysis. Since the paper's headline claims about performance gaps and accuracy–safety trade-offs rest entirely on these scores, the claims are unsupported absent evidence that the judge's scoring aligns with human judgments. The manuscript's own limitation section acknowledges the need for 'more rigorous human-in-the-loop validation,' which is consistent with this concern. Please add a human evaluation on a random sample of responses, report agreement metrics, and discuss potential judge bias toward or against particular model families.
- [§3.3 and §4.2] Section 3.3 states that financial toxicity outputs are 'evaluated through manual review,' whereas Section 4.2 states that all open-ended tasks use the LLM-as-a-judge methodology. These statements are directly contradictory. Please resolve the contradiction and clarify whether any toxicity or reasoning responses were actually human-evaluated, and if so, in which subsection.
- [§3.1.2, §4.1] The benchmark questions were partly generated by GPT-4o (Section 3.1.2), and GPT-4o is also among the evaluated models (Section 4.1). The paper does not report any contamination or leakage analysis, such as comparing GPT-4o's performance on generated versus expert-written items or checking for verbatim overlap with the training data. Without such an analysis, the relative performance of GPT-4o and other models on this benchmark is confounded by potential distributional similarity. Please add a leakage analysis or explicitly discuss this risk.
- [§4.3, Table 8] The performance differences in Table 8 are reported without confidence intervals or any significance testing. For example, the difference between GPT-4o-mini (60.74) and GPT-4o (59.68) on the 377 knowledge questions is within sampling noise, and the same applies to several open-model comparisons (e.g., Llama-3.1-8B 61.80 vs. Qwen2.5-3B 62.33). The claim of 'notable performance differences' requires statistical backing, such as bootstrap confidence intervals or pairwise significance tests, especially for the small reasoning and toxicity item counts.
- [§3.2, Table 5, §5] The paper states in the conclusion that the benchmark contains 1,145 instances, but Table 5's printed subcategory counts sum to 377 + 284 + 484 = 1,145 for the three main categories, while Section 3.2 says financial reasoning totals 209 questions. This is an internal numerical inconsistency that affects a headline claim: either the prose '209' is wrong or the table's reasoning subcategory counts are not additive as shown. Please correct the counts. Additionally, the dataset is not publicly released and is only accessible via the Datop platform, so readers cannot reproduce the benchmark without separate approval; please provide a publicly available or review-ready version with evaluation scripts.
minor comments (5)
- [§4.1] The text repeatedly uses 'LMM' (e.g., 'equitable comparison among LMMs') where 'LLM' is intended. Please correct this terminology.
- [§4.2] The sentence 'Consequently, GPT-4-o1 was excluded from evaluating reasoning scores' refers to a model named 'GPT-4-o1,' but the model list in §4.1 and Table 8 use 'GPT-o1.' Please standardize the model naming.
- [Introduction, references [4,5]] The introduction cites [4] (MMLU) and [5] (C-Eval) as examples of studies showing LLMs' influence on financial decision-making such as investment advisory; these are general benchmark papers, not financial decision-making studies. Please replace them with more appropriate references.
- [Table 7(b)] In the reasoning example, the question refers to 'Article 24' of the Certified Public Accountant Act while the provided context quotes 'Article 21(2).' The relationship between these provisions should be clarified or the citation corrected, since this appears in the motivating example for the benchmark.
- [Appendix A, Table 11] The toxicity evaluation prompt instructs the model to output only a numeric score, but the paper does not specify how missing, malformed, or non-numeric outputs from the judge were handled. This should be reported for reproducibility.
Circularity Check
No significant circularity: benchmark construction and evaluation are empirical, with validity caveats that do not reduce to inputs.
full rationale
I examined the paper's construction and evaluation chain. The benchmark items are produced by GPT-4o prompts followed by expert validation and selection (§3.1.2–§3.1.4), and the correct answers in the knowledge and reasoning tasks are grounded in source documents with expert-reviewed reference rationales (§3.3). Although GPT-4o is among the evaluated models, its scores are measurements on expert-filtered, externally sourced questions, not quantities defined by the model's own outputs. The open-ended reasoning and toxicity scores are produced by GPT-o1 as an LLM judge (§4.2), and the paper explicitly excludes GPT-o1 from reasoning evaluation to avoid self-assessment; the fact that GPT-o1-mini is evaluated in toxicity is a methodological concern about judge validity, not a proof that the reported scores equal the judge's inputs by construction. The inconsistency between §3.3's 'manual review' statement for toxicity and §4.2's LLM-as-a-judge adoption weakens evidential reliability but is not a circular derivation. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, no ansatz is smuggled in solely via self-citation, and no established result is merely renamed. The limitations section explicitly acknowledges the need for 'more rigorous human-in-the-loop validation' in future work, which further confirms that the current judge-based results are presented as provisional rather than as a forced consequence of the benchmark's definition. I therefore find no circular step meeting the quoted-evidence standard, and the appropriate score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The human expert validation ensures factual accuracy and domain relevance of the generated questions.
- domain assumption GPT-o1 as an LLM judge produces valid and unbiased scores for reasoning quality and toxicity.
- domain assumption The source documents from the Bank of Korea, Financial Services Commission, and other institutions are authoritative and correctly interpreted in the generated questions.
- ad hoc to paper The single reported reasoning score in Table 8 meaningfully aggregates the six sub-scores defined in the evaluation prompt.
Cite this review
Pith. "Pith review of KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding." pith.science (2026). https://pith.science/paper/DIAMFSQ2
@misc{pith2026250413216,
author = {Pith},
title = {Pith review of: KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding},
year = {2026},
howpublished = {\url{https://pith.science/paper/DIAMFSQ2}},
note = {Machine review of arXiv:2504.13216}
}
read the original abstract
We introduce KFinEval-Pilot, a benchmark suite specifically designed to evaluate large language models (LLMs) in the Korean financial domain. Addressing the limitations of existing English-centric benchmarks, KFinEval-Pilot comprises over 1,000 curated questions across three critical areas: financial knowledge, legal reasoning, and financial toxicity. The benchmark is constructed through a semi-automated pipeline that combines GPT-4-generated prompts with expert validation to ensure domain relevance and factual accuracy. We evaluate a range of representative LLMs and observe notable performance differences across models, with trade-offs between task accuracy and output safety across different model families. These results highlight persistent challenges in applying LLMs to high-stakes financial applications, particularly in reasoning and safety. Grounded in real-world financial use cases and aligned with the Korean regulatory and linguistic context, KFinEval-Pilot serves as an early diagnostic tool for developing safer and more reliable financial AI systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[2]
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, et al
-
[3]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schel- ten, Alex Vaughan, et al . 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)
arXiv 2024
-
[4]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language under- standing. arXiv preprint arXiv:2009.03300 (2020)
arXiv 2020
-
[5]
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Yao Fu, et al. 2023. C-eval: A multi- level multi-discipline chinese evaluation suite for foundation models. Advances in Neural Information Processing Systems 36 (2023), 62991–63010
2023
-
[6]
Yewon Hwang, Sungbum Jung, Hanwool Lee, and Sara Yu. 2025. TWICE: What Advantages Can Low-Resource Domain-Specific Embedding Model Bring?-A Case Study on Korea Financial Texts. arXiv preprint arXiv:2502.07131 (2025)
work page Pith review arXiv 2025
-
[7]
Pranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, and Bertie Vidgen. 2023. Financebench: A new benchmark for financial question answering. arXiv preprint arXiv:2311.11944 (2023)
arXiv 2023
-
[8]
Yeeun Kim, Young Rok Choi, Eunkyung Choi, Jinhwan Choi, Hai Jin Park, and Wonseok Hwang. 2024. Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models.arXiv preprint arXiv:2410.08731 (2024)
arXiv 2024
Show all 25 references
-
[9]
Yang Lei, Jiangtong Li, Dawei Cheng, Zhijun Ding, and Changjun Jiang. 2023. Cfbenchmark: Chinese financial assistant benchmark for large language model. arXiv preprint arXiv:2311.05812 (2023)
2023 arXiv
- [10]
-
[11]
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Cheng- gang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 (2024)
2024 arXiv
-
[12]
Chanjun Park, Hyeonwoo Kim, Dahyun Kim, Seonghwan Cho, Sanghoon Kim, Sukyung Lee, Yungi Kim, and Hwalsuk Lee. 2024. Open ko-llm leaderboard: Evaluating large language models in korean with ko-h5 benchmark. arXiv preprint arXiv:2405.20574 (2024)
2024 arXiv
-
[13]
Sungjoon Park, Jihyung Moon, Sungdong Kim, Won Ik Cho, Jiyoon Han, Jangwon Park, Chisung Song, Junseong Kim, Yongsook Song, Taehwan Oh, et al. 2021. Klue: Korean language understanding evaluation.arXiv preprint arXiv:2105.09680 (2021)
2021 arXiv
-
[14]
Aman Rangapur, Haoran Wang, Ling Jian, and Kai Shu. 2023. Fin-fact: A bench- mark dataset for multimodal financial fact checking and explanation generation. arXiv preprint arXiv:2309.08793 (2023)
2023 arXiv
-
[15]
Raj Sanjay Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah, Wendi Du, Sudheer Chava, Natraj Raman, Charese Smiley, Jiaao Chen, and Diyi Yang. 2022. When flue meets flang: Benchmarks and large pre-trained language model for financial domain. arXiv preprint arXiv:2211.00083 (2022)
2022 arXiv
-
[16]
Guijin Son, Hyunjun Jeon, Chami Hwang, and Hanearl Jung. 2024. KRX Bench: Automating Financial Benchmark Creation via Large Language Models. In Pro- ceedings of the Joint Workshop of the 7th Financial Technology and Natural Lan- guage Processing, the 5th Knowledge Discovery fr...
2024
-
[17]
Guijin Son, Hanwool Lee, Suwan Kim, Huiseo Kim, Jaecheol Lee, Je Won Yeom, Jihyu Jung, Jung Woo Kim, and Songseong Kim. 2023. Hae-rae bench: Evaluation of korean knowledge in language models. arXiv preprint arXiv:2309.02706 (2023)
2023 arXiv
-
[18]
Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. 2024. Kmmlu: Measuring massive multitask language understanding in korean. arXiv preprint arXiv:2402.11548 (2024)
2024 arXiv
-
[19]
Qianqian Xie, Weiguang Han, Zhengyu Chen, Ruoyu Xiang, Xiao Zhang, Yueru He, Mengxi Xiao, Dong Li, Yongfu Dai, Duanyu Feng, et al . 2024. Finben: A holistic financial benchmark for large language models. Advances in Neural Information Processing Systems 37 (2024), 95716–95743
2024
-
[20]
Qianqian Xie, Weiguang Han, Xiao Zhang, Yanzhao Lai, Min Peng, Alejandro Lopez-Lira, and Jimin Huang. 2023. Pixiu: A large language model, instruction data and evaluation benchmark for finance. arXiv preprint arXiv:2306.05443 (2023)
2023 arXiv
-
[21]
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al . 2024. Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115 (2024)
2024 arXiv
-
[22]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhang- hao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica. 2023. Judging LLM- as-a-Judge with MT-Bench and Chatbot Arena. In Advances in Neu- ral Information ...
2023
-
[23]
Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. TAT-QA: A question an- swering benchmark on a hybrid of tabular and textual content in finance. arXiv preprint arXiv:2105.07624 (2021)
2021 arXiv
-
[24]
Z/Yen Group and China Development Institute. 2025. The Global Financial Cen- tres Index 37 . Technical Report. Z/Yen Group. https://www.longfinance.net/ publications/long-finance-reports/the-global-financial-centres-index-37/ Ac- cessed: 2025-04-10. KFinEval-Pilot: A Comprehen...
2025
-
[2021]
arXiv preprint arXiv:2109.00122 (2021)
Finqa: A dataset of numerical reasoning over financial data. arXiv preprint arXiv:2109.00122 (2021)
2021 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.