REVIEW 3 major objections 6 minor 3 cited by
This survey provides a systematic catalog of 16 legal LLM series, 47 LLM-based frameworks, 15 benchmarks, and 29 datasets, plus a capability taxonomy that shows where legal-LLM evaluation is concentrated and where it is thin.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A structured review of legal LLMs, LLM-based frameworks, benchmarks, and datasets, with a taxonomy and future directions.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection Useful survey, but the 47-framework count is unverifiable and one SOTA entry is misattributed; revise before trusting as a catalog. the 3 major comments →
Large Language Models Meet Legal Artificial Intelligence: A Survey
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
On the paper's own terms, the discovery is the organization: to the authors' knowledge, this is the first systematic review that jointly covers legal datasets (pre-training, supervised fine-tuning, benchmark, and traditional), legal LLMs, and LLM-based frameworks. The central claim is that legal-LLM research can be structured by a taxonomy distinguishing fine-tuned legal LLMs from frameworks that leverage existing LLMs through prompting, retrieval, multi-agent simulation, or domain-model collaboration, and that benchmark tasks can be grouped into seven capabilities. The survey shows that Classification is the only capability represented in all ten general legal benchmarks it examines, while
What carries the argument
The central object is the survey's three-part organizational scheme: datasets (pre-training, SFT, benchmark, traditional), methods (legal LLMs versus LLM-based frameworks), and tasks. The load-bearing sub-mechanism is the seven-capability benchmark taxonomy (Arithmetic, Classification, Information Extraction, Knowledge Assessment, Question Answering, Reasoning, Retrieval), which lets the authors quantify which legal abilities are over-tested and which are neglected. The taxonomy does the argument's work by giving a coordinate system for placing any legal-LLM contribution and by exposing the field's evaluation gaps.
Load-bearing premise
The survey's comprehensiveness rests on an undocumented manual literature selection by two of the authors, so if a substantial number of relevant legal-LLM works published before the cutoff are missing, the central catalog claim fails.
What would settle it
A concrete check: count the actual entries in the paper's Appendix F tables and verify that at least 47 distinct LLM-based frameworks and 29 datasets are individually enumerated; then run a documented, reproducible search over arXiv, the ACL Anthology, and DBLP for legal-LLM papers up to the September 2025 cutoff and count how many relevant works are absent from the survey's 96 cited works. If the appendix lists fewer than 47 frameworks, or the search reveals a substantial number of missing works, the catalog claim collapses.
If this is right
- Newcomers can use the taxonomy to locate any legal-LLM approach as either a fine-tuned model or a framework, and any evaluation as one of seven capabilities.
- The capability analysis implies that legal-LLM evaluation is lopsided: classification-style tasks dominate every major benchmark, so models may be over-fit to labeling and under-tested on arithmetic or legal-knowledge recall.
- The state-of-the-art table (Appendix F) provides concrete reference points for legal information extraction, judgment prediction, question answering, reasoning, retrieval, and summarization across Chinese, English, French, and Burmese.
- The five challenges identified—hallucination, multimodality, multilinguality, interpretability, and data quality—constitute a concrete research agenda for the next phase of legal AI.
- The companion GitHub repository offers a continuously usable index alongside the paper, so the catalog can be extended after publication.
Where Pith is reading between the lines
- If the measured capability imbalance is real, benchmark builders could rebalance evaluation by adding arithmetic and knowledge-recall tasks; this is a plausible consequence the paper does not itself recommend.
- The paper's capability taxonomy could serve as the backbone of a living, continuously updated legal-LLM benchmark tracking models across all seven axes, since the paper's snapshot will age quickly.
- The reported pattern that 6B–7B legal LLMs beat larger models on several Chinese benchmarks suggests that domain-specific fine-tuning can partly compensate for scale; the survey's data point to this hypothesis but do not prove it.
- Because the survey's coverage rests on two authors' manual selection, an independent systematic review with a documented search protocol would either confirm the catalog's completeness or quantify what is missing; that verification is a natural next step the paper does not perform.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews large language model (LLM) methods for legal artificial intelligence. Its stated contributions are a catalog of 16 legal LLM series, 47 LLM-based frameworks, 15 benchmarks, and 29 traditional datasets; a taxonomy distinguishing fine-tuned legal LLMs from frameworks that employ LLMs through prompting or hybrid systems; and an analysis of challenges and future directions. Section 2 covers pre-training, SFT, benchmark, and traditional datasets and metrics; Section 3 discusses legal LLMs (3.1) and framework-based approaches across six legal tasks (3.2); Sections 4–5 present challenges and future directions; six appendices and a GitHub repository supply supporting detail. The numeric catalog is the paper's central value proposition, and its verifiability is the main issue under review.
Significance. If accurate, the catalog is a serviceable entry point for the legal-NLP community: the 16 LLM series are enumerated in Table 5 and Figure 4, the 15 benchmarks are the 11+4 listed in §2.3, and the 29 datasets sum from Table 3. The taxonomy (fine-tuned legal LLMs vs. LLM-based frameworks), the SOTA summary, the honest limitation statement, and the public GitHub resource list are genuine assets. However, the headline '47 LLM-based frameworks' is not enumerated anywhere in the manuscript, and Appendix F contains at least one demonstrable citation error. The survey's usefulness will hinge on making the counts reproducible and correcting the factual errors; as it stands, the central comprehensiveness claim cannot be checked from the manuscript.
major comments (3)
- [Abstract; §6; §3.2] The Abstract and §6 claim '47 LLM-based frameworks,' but no enumeration exists in the manuscript. §3.2 describes frameworks per task with no table, numbered list, or per-task subtotals; Appendix F's Table 7 lists SOTA systems for only six tasks and cannot serve as the enumeration. The counting rule is also undefined: works cited only as 'demonstrating potential' (e.g., Deroy et al., Pont et al., Godbole et al. in Legal Summarization) and works discussed in multiple task sections (e.g., Deng et al. 2024a) are ambiguous. A hand count of named methods in §3.2 yields roughly 40 works, not 47. Since this number is a headline claim, it must be made reproducible—e.g., a table of all 47 frameworks with task, method, citation, and explicit inclusion criteria.
- [Appendix F, Table 7] The Legal Retrieval row reports a 'Self-Construct Dataset,' Burmese, with 87.32 ROUGE-L, citing Nigam et al. (2025). That work is NyayaAnumana/InLegalLLaMA, an Indian English legal judgment prediction dataset and model; nothing in it concerns Burmese retrieval. The low-resource multilingual RAG retrieval work cited in §3.2 is Phyu et al. (2024), which is the reference that could plausibly support this row. This is a concrete factual misattribution in a table whose purpose is to establish SOTA credibly; all rows of Table 7 should be fact-checked against primary sources and corrected.
- [Limitations] The Limitations section states that only two authors performed retrieval, investigation, and categorization, with no documented protocol (databases searched, keywords, date range, inclusion/exclusion criteria, screening counts). For a survey whose contribution is 'comprehensive' and whose abstract foregrounds integer counts, the absence of a reproducible selection process means a reader cannot distinguish 'meticulous' coverage from coverage that missed a substantial body of work. Please add a search-protocol summary or point to a versioned artifact (e.g., the cited GitHub) that records the full screening list and the per-task subtotals underlying the 47-framework count.
minor comments (6)
- [§2.3] The text says 'We categorize the tasks from 10 benchmark datasets' but the same paragraph lists 11 general benchmarks and Figure 3 shows 11 names; reconcile the count.
- [Figure 3] 'SMILDE' should be 'SMILED' (SMILED is the benchmark referenced in §2.3). Check that the figure legend matches the capability assignments in the text.
- [Appendix F, Table 7] The Legal Reasoning SOTA row cites COLIEE 2021 (via Yu et al. 2022a) although Table 3 lists COLIEE 2022/2023 for reasoning; clarify dataset-version naming. Use consistent ROUGE notation (ROUGE-L vs ROUGE-Lsum) and state evaluation settings (zero-shot vs. fine-tuned) for each row.
- [§3.1] The claim that 'Fuzi-Mingcha (6B) demonstrates the best performance' is unsupported by the sparse Table 6, where most cells are missing and Fuzi-Mingcha has scores on only two of four benchmarks. Qualify the claim or compute an aggregate over shared benchmarks.
- [Figure 1 caption; Table 2; Appendix B] Presentation typos: 'a system jointed by LLM and Domain Model' should read 'composed of'; 'community formus' should be 'forums'; Appendix B says 'give an introduce of detailed description' (should be 'give an introduction').
- [Figure 4] The timeline axis (Apr–Oct 2023; Apr–Oct 2024) does not accommodate LLaMandement (released Jan 2024 per Table 5) and omits Lawyer LLaMA 2 (2024.04) despite its presence in Table 5. Add explicit date markers or annotate approximate placements to avoid implying incorrect dates.
Circularity Check
No significant circularity: survey content is an aggregation of external cited works with no derived predictions or fitted parameters.
full rationale
This paper is a survey/catalog of legal LLMs, frameworks, benchmarks, and datasets. It contains no equations, no fitted parameters, no predictions, and no derivation chain that could reduce to its inputs. The headline counts (16 legal LLMs, 47 frameworks, 15 benchmarks, 29 datasets) are aggregations of externally published works; the 47-framework count is asserted without an enumeration, which is a verifiability/completeness limitation rather than a circularity. The only self-referential element is the authors' own GitHub resource list, which is not load-bearing evidence for any claim. The comparative benchmark table in Appendix E relies on independently published benchmark evaluations. There is no instance of a result being defined in terms of another result, no fitted input renamed as a prediction, and no load-bearing self-citation chain. Therefore the paper is not circular.
Axiom & Free-Parameter Ledger
axioms (2)
- domain assumption The selection of 96 works and the 16/47/15/29 counts are complete and representative of the legal LLM literature up to 2025.
- domain assumption The authors' categorization of models, frameworks, datasets, and SOTA results accurately reflects the original papers.
Cite this review
Pith. "Pith review of Large Language Models Meet Legal Artificial Intelligence: A Survey." pith.science (2026). https://pith.science/paper/IV46SFRP
@misc{pith2026250909969,
author = {Pith},
title = {Pith review of: Large Language Models Meet Legal Artificial Intelligence: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/IV46SFRP}},
note = {Machine review of arXiv:2509.09969}
}
read the original abstract
Large Language Models (LLMs) have significantly advanced the development of Legal Artificial Intelligence (Legal AI) in recent years, enhancing the efficiency and accuracy of legal tasks. To advance research and applications of LLM-based approaches in legal domain, this paper provides a comprehensive review of 16 legal LLMs series and 47 LLM-based frameworks for legal tasks, and also gather 15 benchmarks and 29 datasets to evaluate different legal capabilities. Additionally, we analyse the challenges and discuss future directions for LLM-based approaches in the legal domain. We hope this paper provides a systematic introduction for beginners and encourages future research in this field. Resources are available at https://github.com/ZhitianHou/LLMs4LegalAI.
Figures
Forward citations
Cited by 3 Pith papers
-
LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
LegalCiteBench reveals that current LLMs achieve under 7% accuracy on closed-book legal citation retrieval and completion tasks, with misleading answer rates above 94% for nearly all tested models.
-
Exploring Lightweight Large Language Models for Court View Generation
Lightweight LLMs are benchmarked for court view generation and charge prediction across architectures, sizes, DNN comparisons, and task ordering on three datasets using the new CVGEvalKit framework.
-
Towards Intelligent Legal Document Analysis: CNN-Driven Classification of Case Law Texts
A 1D CNN with FastText embeddings classifies legal texts at 97.26% accuracy using 5.1 million parameters and runs over 13 times faster than BERT.
Reference graph
Works this paper leans on
-
[2]
Inlegalllama: Indian legal knowledge en- hanced large language model. InLKM@IJCAI. Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chen- hui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, et al. 2024. Chatglm: A family of large language models from glm-130b to glm-4 all tools.arXiv preprint arXiv:2406.12793. Aditi Godbole, Jabin Geevarghese Ge...
Pith/arXiv arXiv 2024
-
[4]
InAdvances in Neural Information Process- ing Systems, volume 36, pages 44123–44279
Legalbench: A collaboratively built bench- mark for measuring legal reasoning in large language models. InAdvances in Neural Information Process- ing Systems, volume 36, pages 44123–44279. Curran Associates, Inc. Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y . Wu, Y . K. Li, Fuli Luo, Yingfei Xiong, and We...
Pith/arXiv arXiv 2023
-
[5]
GenTranslate: Large language models are gen- erative multilingual speech and machine translators. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 74–90, Bangkok, Thailand. Association for Computational Linguistics. Jia-Hong Huang, Chao-Chun Yang, Yixian Shen, Alessio M. Pacces, and E...
Pith/arXiv arXiv 2024
-
[6]
Coliee 2022 summary: Methods for legal document retrieval and entailment. InNew Frontiers in Artificial Intelligence: JSAI-IsAI 2022 Workshop, JURISIN 2022, and JSAI 2022 International Session, Kyoto, Japan, June 12–17, 2022, Revised Selected Papers, page 51–67, Berlin, Heidelberg. Springer- Verlag. Seoyeon Kim, Kwangwook Seo, Hyungjoo Chae, Jinyoung Yeo,...
Pith/arXiv arXiv 2022
-
[7]
Interpretable long-form legal question an- swering with retrieval-augmented large language models. InProceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty- Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence, AAAI’24/IAAI’24/EAAI...
Pith/arXiv arXiv 2025
-
[8]
Yixiao Ma, Yunqiu Shao, Yueyue Wu, Yiqun Liu, Ruizhe Zhang, Min Zhang, and Shaoping Ma
Leveraging large language models for rele- vance judgments in legal case retrieval.Preprint, arXiv:2403.18405. Yixiao Ma, Yunqiu Shao, Yueyue Wu, Yiqun Liu, Ruizhe Zhang, Min Zhang, and Shaoping Ma. 2021. Lecard: A legal case retrieval dataset for chinese law system. InProceedings of the 44th Interna- tional ACM SIGIR Conference on Research and De- velopm...
Pith/arXiv arXiv 2021
-
[10]
Large language models for data annotation and synthesis: A survey. InProceedings of the 2024 Con- ference on Empirical Methods in Natural Language Processing, pages 930–957, Miami, Florida, USA. Association for Computational Linguistics. Chenhao Tang, Zhengliang Liu, Chong Ma, Zihao Wu, Yiwei Li, Wei Liu, Dajiang Zhu, Quanzheng Li, Xi- ang Li, Tianming Li...
Pith/arXiv arXiv 2024
-
[11]
Xue Zongyue, Liu Huanghai, Hu Yiran, Kong Kangle, Wang Chenlu, Liu Yun, and Shen Weixing
Lawgpt: A chinese legal knowledge-enhanced large language model.Preprint, arXiv:2406.04614. Xue Zongyue, Liu Huanghai, Hu Yiran, Kong Kangle, Wang Chenlu, Liu Yun, and Shen Weixing. 2023. Leec: A legal element extraction dataset with an extensive domain-specific label system.Preprint, arXiv:2310.01271. A Related Survey To the best of our knowledge, there ...
Pith/arXiv arXiv 2023
-
[12]
C Supervised Fine-Tuning Datasets We provide detailed information of the complete SFT datasets of legal LLMs in Table 4
Legal Case Retrieval: Retrieve the similar case in case database based on legal case documents. C Supervised Fine-Tuning Datasets We provide detailed information of the complete SFT datasets of legal LLMs in Table 4. It contains the language and data sources of SFT datasets used on different models. D Detailed Information of Legal LLMs In Table 5, we prov...
2024
-
[13]
97.79@A Legal Judge- ment Predic- tion CAIL2018 Chinese (Wu et al., 2023b) 87.07@A of Law Article, 94.99@A of Charge, 48.72@A of Prison Term ECHR English (Wang et al., 2024) 0.85@F Legal Ques- tion Answer- ing JEC-QA Chinese (Wan et al., 2024) 66.2@A LegalCQA- en English (Jiang et al., 2024b) 21.93@BLEU1 Legal Rea- soning COLIEE 2021 English (Yu et al., 2...
2024
-
[14]
72.77@P, 87.12@R, 80.85@F2 Self- Construct Dataset Burmese (Nigam et al., 2025) 87.32@ROUGE L Legal Sum- marization Claritin English (Chhikara et al.,
2025
-
[2021]
Swiss-judgment-prediction: A multilingual le- gal judgment prediction benchmark. InProceedings of the Natural Legal Language Processing Workshop 2021, pages 19–35, Punta Cana, Dominican Republic. Association for Computational Linguistics. Xiao Peng and Liang Chen. 2024. Athena: Retrieval- augmented legal judgment prediction with large lan- guage models.Pr...
Pith/arXiv arXiv 2021
-
[2023]
InPro- ceedings of the Nineteenth International Conference on Artificial Intelligence and Law, ICAIL ’23, page 472–480, New York, NY , USA
Summary of the competition on legal infor- mation, extraction/entailment (coliee) 2023. InPro- ceedings of the Nineteenth International Conference on Artificial Intelligence and Law, ICAIL ’23, page 472–480, New York, NY , USA. Association for Com- puting Machinery. Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Aditya K, Alex Chohlas-...
2023
-
[2024]
Chenlong Deng, Kelong Mao, and Zhicheng Dou
Automatic information extraction from em- ployment tribunal judgements using large language models.Preprint, arXiv:2403.12936. Chenlong Deng, Kelong Mao, and Zhicheng Dou. 2024a. Learning interpretable legal case retrieval via knowledge-guided case reformulation. InProceed- ings of the 2024 Conference on Empirical Methods in Natural Language Processing, p...
Pith/arXiv arXiv 2024
-
[2025]
62.66@ROUGE-Lsum Table 7: SOTA results for legal NLP tasks across different datasets and languages
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.