Pith. sign in

REVIEW 3 major objections 6 minor 3 cited by

This survey provides a systematic catalog of 16 legal LLM series, 47 LLM-based frameworks, 15 benchmarks, and 29 datasets, plus a capability taxonomy that shows where legal-LLM evaluation is concentrated and where it is thin.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A structured review of legal LLMs, LLM-based frameworks, benchmarks, and datasets, with a taxonomy and future directions.

T0 review reviewed 2026-08-04 challenge →

load-bearing objection Useful survey, but the 47-framework count is unverifiable and one SOTA entry is misattributed; revise before trusting as a catalog. the 3 major comments →

arxiv 2509.09969 v1 pith:IV46SFRP submitted 2025-09-12 cs.CL cs.AI

Large Language Models Meet Legal Artificial Intelligence: A Survey

classification cs.CL cs.AI
keywords legal AIlarge language modelslegal LLMsLLM-based frameworkslegal benchmarkslegal judgment predictionlegal question answeringsurvey
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish a comprehensive, beginner-usable map of large language model work in law: 16 legal LLM families, 47 frameworks that use LLMs without fine-tuning, 15 benchmarks, and 29 traditional datasets. It argues that the field can be organized by a taxonomy that separates fine-tuned legal LLMs from LLM-based frameworks, and that benchmark capabilities fall into seven groups: Arithmetic, Classification, Information Extraction, Knowledge Assessment, Question Answering, Reasoning, and Retrieval. The survey reports that Classification appears across every major benchmark it analyzes, while Arithmetic and Knowledge Assessment are underrepresented, and it identifies challenges including hallucination, multimodality, multilinguality, and interpretability. A sympathetic reader cares because the paper offers newcomers a single entry point and highlights evaluation blind spots that could shape future legal-AI research.

Core claim

On the paper's own terms, the discovery is the organization: to the authors' knowledge, this is the first systematic review that jointly covers legal datasets (pre-training, supervised fine-tuning, benchmark, and traditional), legal LLMs, and LLM-based frameworks. The central claim is that legal-LLM research can be structured by a taxonomy distinguishing fine-tuned legal LLMs from frameworks that leverage existing LLMs through prompting, retrieval, multi-agent simulation, or domain-model collaboration, and that benchmark tasks can be grouped into seven capabilities. The survey shows that Classification is the only capability represented in all ten general legal benchmarks it examines, while

What carries the argument

The central object is the survey's three-part organizational scheme: datasets (pre-training, SFT, benchmark, traditional), methods (legal LLMs versus LLM-based frameworks), and tasks. The load-bearing sub-mechanism is the seven-capability benchmark taxonomy (Arithmetic, Classification, Information Extraction, Knowledge Assessment, Question Answering, Reasoning, Retrieval), which lets the authors quantify which legal abilities are over-tested and which are neglected. The taxonomy does the argument's work by giving a coordinate system for placing any legal-LLM contribution and by exposing the field's evaluation gaps.

Load-bearing premise

The survey's comprehensiveness rests on an undocumented manual literature selection by two of the authors, so if a substantial number of relevant legal-LLM works published before the cutoff are missing, the central catalog claim fails.

What would settle it

A concrete check: count the actual entries in the paper's Appendix F tables and verify that at least 47 distinct LLM-based frameworks and 29 datasets are individually enumerated; then run a documented, reproducible search over arXiv, the ACL Anthology, and DBLP for legal-LLM papers up to the September 2025 cutoff and count how many relevant works are absent from the survey's 96 cited works. If the appendix lists fewer than 47 frameworks, or the search reveals a substantial number of missing works, the catalog claim collapses.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Newcomers can use the taxonomy to locate any legal-LLM approach as either a fine-tuned model or a framework, and any evaluation as one of seven capabilities.
  • The capability analysis implies that legal-LLM evaluation is lopsided: classification-style tasks dominate every major benchmark, so models may be over-fit to labeling and under-tested on arithmetic or legal-knowledge recall.
  • The state-of-the-art table (Appendix F) provides concrete reference points for legal information extraction, judgment prediction, question answering, reasoning, retrieval, and summarization across Chinese, English, French, and Burmese.
  • The five challenges identified—hallucination, multimodality, multilinguality, interpretability, and data quality—constitute a concrete research agenda for the next phase of legal AI.
  • The companion GitHub repository offers a continuously usable index alongside the paper, so the catalog can be extended after publication.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the measured capability imbalance is real, benchmark builders could rebalance evaluation by adding arithmetic and knowledge-recall tasks; this is a plausible consequence the paper does not itself recommend.
  • The paper's capability taxonomy could serve as the backbone of a living, continuously updated legal-LLM benchmark tracking models across all seven axes, since the paper's snapshot will age quickly.
  • The reported pattern that 6B–7B legal LLMs beat larger models on several Chinese benchmarks suggests that domain-specific fine-tuning can partly compensate for scale; the survey's data point to this hypothesis but do not prove it.
  • Because the survey's coverage rests on two authors' manual selection, an independent systematic review with a documented search protocol would either confirm the catalog's completeness or quantify what is missing; that verification is a natural next step the paper does not perform.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This survey reviews large language model (LLM) methods for legal artificial intelligence. Its stated contributions are a catalog of 16 legal LLM series, 47 LLM-based frameworks, 15 benchmarks, and 29 traditional datasets; a taxonomy distinguishing fine-tuned legal LLMs from frameworks that employ LLMs through prompting or hybrid systems; and an analysis of challenges and future directions. Section 2 covers pre-training, SFT, benchmark, and traditional datasets and metrics; Section 3 discusses legal LLMs (3.1) and framework-based approaches across six legal tasks (3.2); Sections 4–5 present challenges and future directions; six appendices and a GitHub repository supply supporting detail. The numeric catalog is the paper's central value proposition, and its verifiability is the main issue under review.

Significance. If accurate, the catalog is a serviceable entry point for the legal-NLP community: the 16 LLM series are enumerated in Table 5 and Figure 4, the 15 benchmarks are the 11+4 listed in §2.3, and the 29 datasets sum from Table 3. The taxonomy (fine-tuned legal LLMs vs. LLM-based frameworks), the SOTA summary, the honest limitation statement, and the public GitHub resource list are genuine assets. However, the headline '47 LLM-based frameworks' is not enumerated anywhere in the manuscript, and Appendix F contains at least one demonstrable citation error. The survey's usefulness will hinge on making the counts reproducible and correcting the factual errors; as it stands, the central comprehensiveness claim cannot be checked from the manuscript.

major comments (3)
  1. [Abstract; §6; §3.2] The Abstract and §6 claim '47 LLM-based frameworks,' but no enumeration exists in the manuscript. §3.2 describes frameworks per task with no table, numbered list, or per-task subtotals; Appendix F's Table 7 lists SOTA systems for only six tasks and cannot serve as the enumeration. The counting rule is also undefined: works cited only as 'demonstrating potential' (e.g., Deroy et al., Pont et al., Godbole et al. in Legal Summarization) and works discussed in multiple task sections (e.g., Deng et al. 2024a) are ambiguous. A hand count of named methods in §3.2 yields roughly 40 works, not 47. Since this number is a headline claim, it must be made reproducible—e.g., a table of all 47 frameworks with task, method, citation, and explicit inclusion criteria.
  2. [Appendix F, Table 7] The Legal Retrieval row reports a 'Self-Construct Dataset,' Burmese, with 87.32 ROUGE-L, citing Nigam et al. (2025). That work is NyayaAnumana/InLegalLLaMA, an Indian English legal judgment prediction dataset and model; nothing in it concerns Burmese retrieval. The low-resource multilingual RAG retrieval work cited in §3.2 is Phyu et al. (2024), which is the reference that could plausibly support this row. This is a concrete factual misattribution in a table whose purpose is to establish SOTA credibly; all rows of Table 7 should be fact-checked against primary sources and corrected.
  3. [Limitations] The Limitations section states that only two authors performed retrieval, investigation, and categorization, with no documented protocol (databases searched, keywords, date range, inclusion/exclusion criteria, screening counts). For a survey whose contribution is 'comprehensive' and whose abstract foregrounds integer counts, the absence of a reproducible selection process means a reader cannot distinguish 'meticulous' coverage from coverage that missed a substantial body of work. Please add a search-protocol summary or point to a versioned artifact (e.g., the cited GitHub) that records the full screening list and the per-task subtotals underlying the 47-framework count.
minor comments (6)
  1. [§2.3] The text says 'We categorize the tasks from 10 benchmark datasets' but the same paragraph lists 11 general benchmarks and Figure 3 shows 11 names; reconcile the count.
  2. [Figure 3] 'SMILDE' should be 'SMILED' (SMILED is the benchmark referenced in §2.3). Check that the figure legend matches the capability assignments in the text.
  3. [Appendix F, Table 7] The Legal Reasoning SOTA row cites COLIEE 2021 (via Yu et al. 2022a) although Table 3 lists COLIEE 2022/2023 for reasoning; clarify dataset-version naming. Use consistent ROUGE notation (ROUGE-L vs ROUGE-Lsum) and state evaluation settings (zero-shot vs. fine-tuned) for each row.
  4. [§3.1] The claim that 'Fuzi-Mingcha (6B) demonstrates the best performance' is unsupported by the sparse Table 6, where most cells are missing and Fuzi-Mingcha has scores on only two of four benchmarks. Qualify the claim or compute an aggregate over shared benchmarks.
  5. [Figure 1 caption; Table 2; Appendix B] Presentation typos: 'a system jointed by LLM and Domain Model' should read 'composed of'; 'community formus' should be 'forums'; Appendix B says 'give an introduce of detailed description' (should be 'give an introduction').
  6. [Figure 4] The timeline axis (Apr–Oct 2023; Apr–Oct 2024) does not accommodate LLaMandement (released Jan 2024 per Table 5) and omits Lawyer LLaMA 2 (2024.04) despite its presence in Table 5. Add explicit date markers or annotate approximate placements to avoid implying incorrect dates.

Circularity Check

0 steps flagged

No significant circularity: survey content is an aggregation of external cited works with no derived predictions or fitted parameters.

full rationale

This paper is a survey/catalog of legal LLMs, frameworks, benchmarks, and datasets. It contains no equations, no fitted parameters, no predictions, and no derivation chain that could reduce to its inputs. The headline counts (16 legal LLMs, 47 frameworks, 15 benchmarks, 29 datasets) are aggregations of externally published works; the 47-framework count is asserted without an enumeration, which is a verifiability/completeness limitation rather than a circularity. The only self-referential element is the authors' own GitHub resource list, which is not load-bearing evidence for any claim. The comparative benchmark table in Appendix E relies on independently published benchmark evaluations. There is no instance of a result being defined in terms of another result, no fitted input renamed as a prediction, and no load-bearing self-citation chain. Therefore the paper is not circular.

Axiom & Free-Parameter Ledger

0 free parameters · 2 axioms · 0 invented entities

A survey introduces no free parameters, mathematical axioms, or invented entities. Its load-bearing premises are coverage completeness and categorization accuracy, both acknowledged as fragile in the Limitations section.

axioms (2)
  • domain assumption The selection of 96 works and the 16/47/15/29 counts are complete and representative of the legal LLM literature up to 2025.
    The survey's comprehensiveness claim depends on this. No systematic search protocol or inclusion criteria are provided, and the Limitations section admits relevant works may be undiscovered.
  • domain assumption The authors' categorization of models, frameworks, datasets, and SOTA results accurately reflects the original papers.
    The taxonomy and SOTA table are the main content. The Limitations section acknowledges possible human errors in categorization, and Table 7 contains at least one suspicious attribution.

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Large Language Models Meet Legal Artificial Intelligence: A Survey." pith.science (2026). https://pith.science/paper/IV46SFRP

@misc{pith2026250909969,
  author       = {Pith},
  title        = {Pith review of: Large Language Models Meet Legal Artificial Intelligence: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IV46SFRP}},
  note         = {Machine review of arXiv:2509.09969}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large Language Models (LLMs) have significantly advanced the development of Legal Artificial Intelligence (Legal AI) in recent years, enhancing the efficiency and accuracy of legal tasks. To advance research and applications of LLM-based approaches in legal domain, this paper provides a comprehensive review of 16 legal LLMs series and 47 LLM-based frameworks for legal tasks, and also gather 15 benchmarks and 29 datasets to evaluate different legal capabilities. Additionally, we analyse the challenges and discuss future directions for LLM-based approaches in the legal domain. We hope this paper provides a systematic introduction for beginners and encourages future research in this field. Resources are available at https://github.com/ZhitianHou/LLMs4LegalAI.

Figures

Figures reproduced from arXiv: 2509.09969 by Kun Zeng, Nanli Zeng, Tianyong Hao, Zhitian Hou, Zihan Ye.

Figure 1
Figure 1. Figure 1: An example of LLMs in legal judgement prediction task. (a) is a pipeline of fine-tuning a new legal LLM. (b) and (c) are LLM-based frameworks for the task. (b) utilizes legal syllogism within the prompt. (c) uses a system jointed by LLM and Domain Model. fine-tuning) to address traditional legal tasks. Ad￾ditionally, datasets for the training and evaluation of LLMs have been developed. Despite the im￾press… view at source ↗
Figure 2
Figure 2. Figure 2: The organization of this survey. knowledge, this is the first systematic review of Legal AI datasets both traditional and LLM￾specific, and LLM-based approaches include legal LLMs and LLM-based frameworks. • Datasets Analysis: We provide an analysis of existing Legal AI datasets, offering a thorough overview of their characteristics. • Meticulous Taxonomy: We present a de￾tailed taxonomy, distinguishing be… view at source ↗
Figure 3
Figure 3. Figure 3: The capabilities assessed by benchmarks. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: The timeline of Legal LLMs, with only the largest model sizes plotted for visual clarity. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: The examples of 6 main legal tasks. tion extraction aims to to extract words or phrases that have important legal significance from legal texts. The deployment of LLMs for information ex￾traction in legal texts offers numerous potential ad￾vantages (de Faria et al., 2024). Some approaches use weak supervision or sequence labeling tech￾niques to generate annotated samples (Adhikary et al., 2024), while othe… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LegalCiteBench: Evaluating Citation Reliability in Legal Language Models

    cs.CL 2026-05 unverdicted novelty 6.0

    LegalCiteBench reveals that current LLMs achieve under 7% accuracy on closed-book legal citation retrieval and completion tasks, with misleading answer rates above 94% for nearly all tested models.

  2. Exploring Lightweight Large Language Models for Court View Generation

    cs.CL 2026-05 unverdicted novelty 4.0

    Lightweight LLMs are benchmarked for court view generation and charge prediction across architectures, sizes, DNN comparisons, and task ordering on three datasets using the new CVGEvalKit framework.

  3. Towards Intelligent Legal Document Analysis: CNN-Driven Classification of Case Law Texts

    cs.CL 2026-04 unverdicted novelty 4.0

    A 1D CNN with FastText embeddings classifies legal texts at 97.26% accuracy using 5.1 million parameters and runs over 13 times faster than BERT.

Reference graph

Works this paper leans on

15 extracted references · 10 linked inside Pith · cited by 3 Pith papers

  1. [2]

    InLKM@IJCAI

    Inlegalllama: Indian legal knowledge en- hanced large language model. InLKM@IJCAI. Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chen- hui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, et al. 2024. Chatglm: A family of large language models from glm-130b to glm-4 all tools.arXiv preprint arXiv:2406.12793. Aditi Godbole, Jabin Geevarghese Ge...

  2. [4]

    InAdvances in Neural Information Process- ing Systems, volume 36, pages 44123–44279

    Legalbench: A collaboratively built bench- mark for measuring legal reasoning in large language models. InAdvances in Neural Information Process- ing Systems, volume 36, pages 44123–44279. Curran Associates, Inc. Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y . Wu, Y . K. Li, Fuli Luo, Yingfei Xiong, and We...

  3. [5]

    InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 74–90, Bangkok, Thailand

    GenTranslate: Large language models are gen- erative multilingual speech and machine translators. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 74–90, Bangkok, Thailand. Association for Computational Linguistics. Jia-Hong Huang, Chao-Chun Yang, Yixian Shen, Alessio M. Pacces, and E...

  4. [6]

    Coliee 2022 summary: Methods for legal document retrieval and entailment. InNew Frontiers in Artificial Intelligence: JSAI-IsAI 2022 Workshop, JURISIN 2022, and JSAI 2022 International Session, Kyoto, Japan, June 12–17, 2022, Revised Selected Papers, page 51–67, Berlin, Heidelberg. Springer- Verlag. Seoyeon Kim, Kwangwook Seo, Hyungjoo Chae, Jinyoung Yeo,...

  5. [7]

    Interpretable long-form legal question an- swering with retrieval-augmented large language models. InProceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty- Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence, AAAI’24/IAAI’24/EAAI...

  6. [8]

    Yixiao Ma, Yunqiu Shao, Yueyue Wu, Yiqun Liu, Ruizhe Zhang, Min Zhang, and Shaoping Ma

    Leveraging large language models for rele- vance judgments in legal case retrieval.Preprint, arXiv:2403.18405. Yixiao Ma, Yunqiu Shao, Yueyue Wu, Yiqun Liu, Ruizhe Zhang, Min Zhang, and Shaoping Ma. 2021. Lecard: A legal case retrieval dataset for chinese law system. InProceedings of the 44th Interna- tional ACM SIGIR Conference on Research and De- velopm...

  7. [10]

    reasoning before responding

    Large language models for data annotation and synthesis: A survey. InProceedings of the 2024 Con- ference on Empirical Methods in Natural Language Processing, pages 930–957, Miami, Florida, USA. Association for Computational Linguistics. Chenhao Tang, Zhengliang Liu, Chong Ma, Zihao Wu, Yiwei Li, Wei Liu, Dajiang Zhu, Quanzheng Li, Xi- ang Li, Tianming Li...

  8. [11]

    Xue Zongyue, Liu Huanghai, Hu Yiran, Kong Kangle, Wang Chenlu, Liu Yun, and Shen Weixing

    Lawgpt: A chinese legal knowledge-enhanced large language model.Preprint, arXiv:2406.04614. Xue Zongyue, Liu Huanghai, Hu Yiran, Kong Kangle, Wang Chenlu, Liu Yun, and Shen Weixing. 2023. Leec: A legal element extraction dataset with an extensive domain-specific label system.Preprint, arXiv:2310.01271. A Related Survey To the best of our knowledge, there ...

  9. [12]

    C Supervised Fine-Tuning Datasets We provide detailed information of the complete SFT datasets of legal LLMs in Table 4

    Legal Case Retrieval: Retrieve the similar case in case database based on legal case documents. C Supervised Fine-Tuning Datasets We provide detailed information of the complete SFT datasets of legal LLMs in Table 4. It contains the language and data sources of SFT datasets used on different models. D Detailed Information of Legal LLMs In Table 5, we prov...

  10. [13]

    97.79@A Legal Judge- ment Predic- tion CAIL2018 Chinese (Wu et al., 2023b) 87.07@A of Law Article, 94.99@A of Charge, 48.72@A of Prison Term ECHR English (Wang et al., 2024) 0.85@F Legal Ques- tion Answer- ing JEC-QA Chinese (Wan et al., 2024) 66.2@A LegalCQA- en English (Jiang et al., 2024b) 21.93@BLEU1 Legal Rea- soning COLIEE 2021 English (Yu et al., 2...

  11. [14]

    72.77@P, 87.12@R, 80.85@F2 Self- Construct Dataset Burmese (Nigam et al., 2025) 87.32@ROUGE L Legal Sum- marization Claritin English (Chhikara et al.,

  12. [2021]

    InProceedings of the Natural Legal Language Processing Workshop 2021, pages 19–35, Punta Cana, Dominican Republic

    Swiss-judgment-prediction: A multilingual le- gal judgment prediction benchmark. InProceedings of the Natural Legal Language Processing Workshop 2021, pages 19–35, Punta Cana, Dominican Republic. Association for Computational Linguistics. Xiao Peng and Liang Chen. 2024. Athena: Retrieval- augmented legal judgment prediction with large lan- guage models.Pr...

  13. [2023]

    InPro- ceedings of the Nineteenth International Conference on Artificial Intelligence and Law, ICAIL ’23, page 472–480, New York, NY , USA

    Summary of the competition on legal infor- mation, extraction/entailment (coliee) 2023. InPro- ceedings of the Nineteenth International Conference on Artificial Intelligence and Law, ICAIL ’23, page 472–480, New York, NY , USA. Association for Com- puting Machinery. Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Aditya K, Alex Chohlas-...

  14. [2024]

    Chenlong Deng, Kelong Mao, and Zhicheng Dou

    Automatic information extraction from em- ployment tribunal judgements using large language models.Preprint, arXiv:2403.12936. Chenlong Deng, Kelong Mao, and Zhicheng Dou. 2024a. Learning interpretable legal case retrieval via knowledge-guided case reformulation. InProceed- ings of the 2024 Conference on Empirical Methods in Natural Language Processing, p...

  15. [2025]

    62.66@ROUGE-Lsum Table 7: SOTA results for legal NLP tasks across different datasets and languages

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.