REVIEW 1 major objections 1 minor 9 references
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
T0 review · 1 major / 1 minor · reviewed 2026-05-10 · grok-4.3
Pith's one-line read LLMs' primary limitation on Vietnamese legal texts is accurate reasoning rather than summarization or readability.
desk verdict The paper benchmarks four LLMs on Vietnamese legal simplification and uses error analysis on 60 articles to argue reasoning is the main gap, but the small curated set leaves the generalization under-supported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A dual-aspect framework that pairs quantitative benchmarks across accuracy, readability, and consistency with qualitative error analysis using an expert-validated typology applied to sixty complex Vietnamese legal articles.
What would settle it
A replication on a larger or differently sampled set of Vietnamese legal texts that instead finds summarization-level errors outnumbering reasoning errors, or that shows models without targeted reasoning training match the performance of the tested models.
Extended reading notes
Core claim
The evaluation shows that models such as Grok-1 achieve strong readability and consistency yet lose fine-grained legal accuracy, while Claude 3 Opus scores higher on accuracy metrics that still conceal subtle reasoning mistakes. Across the sixty articles, the most common errors fall into the categories of incorrect examples and misinterpretation of legal provisions. These patterns lead directly to the conclusion that current LLMs can produce fluent simplifications but lack reliable control over the precise legal reasoning required to preserve meaning without distortion.
Load-bearing premise
The sixty selected Vietnamese legal articles plus the error typology together capture the full range of challenges and failure modes that arise in the domain.
Editorial extensions
If this is right
- Grok-1 trades off accuracy to gain readability and consistency.
- Claude 3 Opus reaches higher accuracy scores while still producing critical but hidden reasoning mistakes.
- Incorrect examples and misinterpretations dominate the observed error distribution.
- Efforts to improve LLMs for legal use should target controlled reasoning rather than general simplification.
Reading between the lines
- The same dual-aspect method could be reused on legal texts from other languages or jurisdictions to test whether the reasoning bottleneck is language-specific.
- If the finding holds, training data and evaluation suites for legal LLMs should emphasize step-by-step legal inference chains over fluency alone.
- Standard single-metric benchmarks may systematically overstate model readiness for high-stakes text simplification tasks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a dual-aspect evaluation framework for LLMs on Vietnamese legal text simplification. It first benchmarks four models (GPT-4o, Claude 3 Opus, Gemini 1.5 Pro, Grok-1) on Accuracy, Readability, and Consistency. It then performs error analysis on a curated set of 60 complex Vietnamese legal articles using a novel expert-validated error typology, identifying Incorrect Example and Misinterpretation as the dominant failure modes and concluding that the primary LLM challenge is controlled, accurate legal reasoning rather than summarization.
Significance. If the error analysis methodology is strengthened, the work offers a useful template for moving beyond surface metrics in legal NLP evaluation, particularly for low-resource languages, and surfaces actionable trade-offs (e.g., readability vs. fine-grained accuracy) that could inform model development for justice-access applications.
major comments (1)
- [Error analysis and dataset curation] The headline conclusion—that Incorrect Example and Misinterpretation errors dominate and therefore 'the primary challenge ... is not summarization but controlled, accurate legal reasoning'—depends on the error analysis of the 60-article set. The manuscript provides no sampling frame, domain stratification, or justification for why these 60 complex articles are representative of Vietnamese legal texts overall (see the description of the curated dataset and the error typology). Without inter-annotator agreement statistics or explicit mapping showing how the typology separates summarization-style inaccuracies (e.g., loss of legal qualifiers) from deeper reasoning failures, the observed prevalence cannot support the domain-wide claim.
minor comments (1)
- [Results and error analysis] Clarify the exact scale of the 'large-scale error analysis' given that it covers only 60 articles; consider adding a table summarizing error frequencies per model and per error type for transparency.
Simulated Author's Rebuttal
We thank the referee for their constructive feedback, which identifies key areas where the error analysis methodology can be clarified. We respond point by point to the major comment below and are prepared to revise the manuscript accordingly.
read point-by-point responses
-
Referee: The headline conclusion—that Incorrect Example and Misinterpretation errors dominate and therefore 'the primary challenge ... is not summarization but controlled, accurate legal reasoning'—depends on the error analysis of the 60-article set. The manuscript provides no sampling frame, domain stratification, or justification for why these 60 complex articles are representative of Vietnamese legal texts overall (see the description of the curated dataset and the error typology). Without inter-annotator agreement statistics or explicit mapping showing how the typology separates summarization-style inaccuracies (e.g., loss of legal qualifiers) from deeper reasoning failures, the observed prevalence cannot support the domain-wide claim.
Authors: We agree that the manuscript would benefit from expanded details on curation and typology application. The 60 articles were deliberately selected as complex cases (high lexical density, interdependent clauses, and expert-flagged reasoning demands) from official Vietnamese legal repositories to focus the error analysis on challenging instances rather than a random or stratified sample of all legal texts; we will add an explicit description of the selection criteria and sources in the revised version. The error typology was expert-validated but annotations were performed by the research team under legal expert oversight, so inter-annotator agreement statistics were not computed; we will note this limitation transparently. We will also insert a mapping table that classifies each error type, e.g., 'Incorrect Example' as fabrication of non-existent legal illustrations (reasoning failure) versus 'Misinterpretation' as misreading of qualifiers or conditionals (distinct from mere summarization loss). These revisions will better ground the prevalence claims while preserving the observation that, within the evaluated complex texts, reasoning errors predominate over basic summarization shortfalls. We do not claim the 60 articles represent the full distribution of Vietnamese legal texts, only that they surface actionable failure modes for legal applications. revision: partial
Circularity Check
No circularity: pure empirical evaluation without derivations or self-referential claims
full rationale
The paper conducts a benchmark of four LLMs on Vietnamese legal simplification using accuracy/readability/consistency metrics, followed by error analysis on a fixed curated set of 60 articles with an expert-validated typology. No equations, fitted parameters, first-principles derivations, or predictions appear in the provided text. Conclusions about reasoning vs. summarization failures are direct counts from the error typology applied to the external human-validated dataset; they do not reduce to any input by construction. The study is self-contained against external benchmarks and human judgments, with no load-bearing self-citations or ansatzes.
Assumptions & free parameters
assumptions (1)
- domain assumption Accuracy, readability, and consistency metrics are valid and sufficient for assessing legal text simplification quality.
invented entities (1)
-
Expert-validated error typology
Cite this review
Pith. "Pith review of From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text." pith.science (2026). https://pith.science/paper/2604.16270
@misc{pith2026260416270,
author = {Pith},
title = {Pith review of: From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text},
year = {2026},
howpublished = {\url{https://pith.science/paper/2604.16270}},
note = {Machine review of arXiv:2604.16270}
}
read the original abstract
The complexity of Vietnam's legal texts presents a significant barrier to public access to justice. While Large Language Models offer a promising solution for legal text simplification, evaluating their true capabilities requires a multifaceted approach that goes beyond surface-level metrics. This paper introduces a comprehensive dual-aspect evaluation framework to address this need. First, we establish a performance benchmark for four state-of-the-art large language models (GPT-4o, Claude 3 Opus, Gemini 1.5 Pro, and Grok-1) across three key dimensions: Accuracy, Readability, and Consistency. Second, to understand the "why" behind these performance scores, we conduct a large-scale error analysis on a curated dataset of 60 complex Vietnamese legal articles, using a novel, expert-validated error typology. Our results reveal a crucial trade-off: models like Grok-1 excel in Readability and Consistency but compromise on fine-grained legal Accuracy, while models like Claude 3 Opus achieve high Accuracy scores that mask a significant number of subtle but critical reasoning errors. The error analysis pinpoints \textit{Incorrect Example} and \textit{Misinterpretation} as the most prevalent failures, confirming that the primary challenge for current LLMs is not summarization but controlled, accurate legal reasoning. By integrating a quantitative benchmark with a qualitative deep dive, our work provides a holistic and actionable assessment of LLMs for legal applications.
Figures
Reference graph
Works this paper leans on
-
[1]
Unsupervised simplification of legal texts.arXiv preprint arXiv:2209.00557, 2022
Mert Cemri, Tolga C ¸ ukur, and Aykut Koc ¸. Unsupervised simplification of legal texts.arXiv preprint arXiv:2209.00557, 2022
-
[2]
Legal-bert: The muppets straight out of law school.arXiv preprint arXiv:2010.02559, 2020
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. Legal-bert: The muppets straight out of law school.arXiv preprint arXiv:2010.02559, 2020
-
[3]
Large legal fictions: Profiling legal hallucinations in large language models
Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E Ho. Large legal fictions: Profiling legal hallucinations in large language models. Journal of Legal Analysis, 16(1):64–93, 2024
work page 2024
-
[4]
Text simplification for legal domain: Insights and challenges
Aparna Garimella, Abhilasha Sancheti, Vinay Aggarwal, Ananya Ganesh, Niyati Chhaya, and Nanda Kambhatla. Text simplification for legal domain: Insights and challenges. InProceedings of the Natural Legal Language Processing Workshop 2022, pages 296–304, 2022
work page 2022
-
[5]
Joshua Kelsall, Xingwei Tan, Aislinn Bergin, Jiahong Chen, Maria Waheed, Tom Sorell, Rob Procter, Maria Liakata, Jenny Chim, and Serene Chi. A rapid evidence review of evaluation techniques for large language models in legal use cases: trends, gaps, and recommendations for future research.AI & SOCIETY, pages 1–19, 2025
work page 2025
-
[6]
Large language models in law: A survey.AI Open, 5:181–196, 2024
Jinqi Lai, Wensheng Gan, Jiayang Wu, Zhenlian Qi, and Philip S Yu. Large language models in law: A survey.AI Open, 5:181–196, 2024
work page 2024
-
[7]
Tan-Minh Nguyen, Hoang-Trung Nguyen, Trong-Khoi Dao, Xuan-Hieu Phan, Ha-Thanh Nguyen, and Thi-Hai- Y en Vuong. Vlqa: The first com- prehensive, large, and high-quality vietnamese dataset for legal question answering.arXiv preprint arXiv:2507.19995, 2025
-
[8]
Access to justice in Vietnam: State supply–private distrust
Pip Nicholson. Access to justice in Vietnam: State supply–private distrust. InLegal Reforms in China and Vietnam, pages 188–215. Routledge, 2010
work page 2010
Show all 9 references
-
[9]
Top 2 at alqac 2024: Large language models (llms) for legal question answering.International Journal of Asian Language Processing, 35(01):2450010, 2025
Huy Quang Pham, Quan Van Nguyen, Dan Quang Tran, Thang Kien- Bao Nguyen, and Kiet Van Nguyen. Top 2 at alqac 2024: Large language models (llms) for legal question answering.International Journal of Asian Language Processing, 35(01):2450010, 2025
2024
Reviewed May 10, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.