Pith. sign in

REVIEW 6 cited by

LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.20288 v4 pith:XCUX2646 submitted 2024-09-30 cs.CL

classification cs.CL
keywords legallexevalllmschinesebenchmarkdatasetsevaluationlanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large language models (LLMs) have made significant progress in natural language processing tasks and demonstrate considerable potential in the legal domain. However, legal applications demand high standards of accuracy, reliability, and fairness. Applying existing LLMs to legal systems without careful evaluation of their potential and limitations could pose significant risks in legal practice. To this end, we introduce a standardized comprehensive Chinese legal benchmark LexEval. This benchmark is notable in the following three aspects: (1) Ability Modeling: We propose a new taxonomy of legal cognitive abilities to organize different tasks. (2) Scale: To our knowledge, LexEval is currently the largest Chinese legal evaluation dataset, comprising 23 tasks and 14,150 questions. (3) Data: we utilize formatted existing datasets, exam datasets and newly annotated datasets by legal experts to comprehensively evaluate the various capabilities of LLMs. LexEval not only focuses on the ability of LLMs to apply fundamental legal knowledge but also dedicates efforts to examining the ethical issues involved in their application. We evaluated 38 open-source and commercial LLMs and obtained some interesting findings. The experiments and findings offer valuable insights into the challenges and potential solutions for developing Chinese legal systems and LLM evaluation pipelines. The LexEval dataset and leaderboard are publicly available at \url{https://github.com/CSHaitao/LexEval} and will be continuously updated.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025)

    cs.CL 2026-07 conditional novelty 7.0 of 10

    BLAD releases 1,484 Bangladeshi legal acts (1799–2025) with structural annotations and historical government context.

  2. CitaLaw: Enhancing LLM with Citations in Legal Domain

    cs.CL 2024-12 conditional novelty 7.0 of 10

    CitaLaw is a Chinese legal benchmark that tests citation-grounded answers for laypeople and legal practitioners, with a syllogism-based evaluation that shows substantial agreement with human judges.

  3. Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An automated LLM-based evaluator finds that eight LLMs rarely hallucinate factors in legal argument generation but often omit relevant factors and usually fail to abstain when no common ground exists.

  4. LegalAgentBench: Evaluating LLM Agents in Legal Domain

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new Chinese legal-domain benchmark with 17 real-world corpora, 37 tools, 300 human-verified tasks, and a fine-grained evaluation metric shows GPT-4o leads with 79% success under ReAct.

  5. Evaluating LLM-based Approaches to Legal Citation Prediction: Domain-specific Pre-training, Fine-tuning, or RAG? A Benchmark and an Australian Law Case Study

    cs.CL 2024-12 conditional novelty 6.0 of 10

    The paper releases the AusLaw Citation Benchmark and shows that instruction-tuned 7B-8B LLMs plus retrieval re-ranking outperform general and law-specific pretrained LLMs for legal citation prediction, reaching about ...

  6. A Survey of LLM $\times$ DATA

    cs.DB 2025-05 conditional novelty 5.0 of 10

    A comprehensive survey of the bidirectional links between LLMs and data management, organized as DATA4LLM and LLM4DATA with a new 'IaaS' data-quality framework.

Pith tools