Pith. sign in

REVIEW 6 cited by

LegalAgentBench: Evaluating LLM Agents in Legal Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.17259 v1 pith:J76LGJJH submitted 2024-12-23 cs.CL cs.IR

classification cs.CLcs.IR
keywords legallegalagentbenchdomainagentsreal-worldbenchmarkcomplexitydesigned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the increasing intelligence and autonomy of LLM agents, their potential applications in the legal domain are becoming increasingly apparent. However, existing general-domain benchmarks cannot fully capture the complexity and subtle nuances of real-world judicial cognition and decision-making. Therefore, we propose LegalAgentBench, a comprehensive benchmark specifically designed to evaluate LLM Agents in the Chinese legal domain. LegalAgentBench includes 17 corpora from real-world legal scenarios and provides 37 tools for interacting with external knowledge. We designed a scalable task construction framework and carefully annotated 300 tasks. These tasks span various types, including multi-hop reasoning and writing, and range across different difficulty levels, effectively reflecting the complexity of real-world legal scenarios. Moreover, beyond evaluating final success, LegalAgentBench incorporates keyword analysis during intermediate processes to calculate progress rates, enabling more fine-grained evaluation. We evaluated eight popular LLMs, highlighting the strengths, limitations, and potential areas for improvement of existing models and methods. LegalAgentBench sets a new benchmark for the practical application of LLMs in the legal domain, with its code and data available at \url{https://github.com/CSHaitao/LegalAgentBench}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AppealCase: A Dataset and Benchmark for Civil Case Appeal Scenarios

    cs.CL 2025-05 conditional novelty 7.0 of 10

    AppealCase is a new paired first- and second-instance Chinese civil judgment benchmark with five appellate LegalAI tasks on which current models score below 50% F1 for reversal prediction from the first-instance perspective.

  2. KoBLEX: Open Legal Question Answering with Multi-hop Reasoning

    cs.CL 2025-09 conditional novelty 6.0 of 10

    KoBLEX is a bilingual 226-instance provision-grounded legal QA benchmark, and its ParSeR pipeline (generate pseudo-statutes, then retrieve-rerank-select real ones) beats baselines across five LLMs, graded by a new hum...

  3. LaQual: An Automated Framework for LLM App Quality Evaluation

    cs.SE 2025-08 reject novelty 5.0 of 10

    LaQual automates LLM app-store quality evaluation through scenario classification, static indicator filtering, and LLM-generated dynamic metrics, with Spearman correlations of about 0.6 against human ratings.

  4. BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

    cs.LG 2025-05 conditional novelty 5.0 of 10

    BenchHub is an automatically categorized, customizable LLM benchmark suite covering 303K questions across 38 benchmarks in English and Korean.

  5. Collaborative Editable Model

    cs.AI 2025-06 reject novelty 4.0 of 10

    CoEM scores user-contributed knowledge fragments using user ratings and LLM attribution, keeps the high scorers in a prompt-level knowledge pool, and reports 76% agreement with FinGPT on fragment value.

  6. ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation

    cs.AI 2025-06 reject novelty 4.0 of 10

    ChemAU adds a position penalty to token-level uncertainty estimates so that flagged reasoning steps are corrected by a fine-tuned chemistry model, reporting improved accuracy on GPQA, MMLU-Pro, and SuperGPQA chemistry...

Pith tools