Pith. sign in

REVIEW 2 cited by

LegalBench: Prototyping a Collaborative Benchmark for Legal Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.06120 v1 pith:MB7R2LNJ submitted 2022-09-13 cs.AI

classification cs.AI
keywords legaltasksbenchmarksciencecollaborativecommunitiescomputerefforts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Can foundation models be guided to execute tasks involving legal reasoning? We believe that building a benchmark to answer this question will require sustained collaborative efforts between the computer science and legal communities. To that end, this short paper serves three purposes. First, we describe how IRAC-a framework legal scholars use to distinguish different types of legal reasoning-can guide the construction of a Foundation Model oriented benchmark. Second, we present a seed set of 44 tasks built according to this framework. We discuss initial findings, and highlight directions for new tasks. Finally-inspired by the Open Science movement-we make a call for the legal and computer science communities to join our efforts by contributing new tasks. This work is ongoing, and our progress can be tracked here: https://github.com/HazyResearch/legalbench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BenGER Platform: A Collaborative Web Platform for End-to-End Benchmarking of German Legal Tasks

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    BenGER is a collaborative web platform that integrates end-to-end workflows for creating, annotating, running, and evaluating benchmarks on German legal tasks with large language models.

  2. LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents

    cs.AI 2025-09 reject novelty 5.0 of 10

    On CUAD legal contracts, a prompt-engineered QWEN-2 pipeline with chunking and two answer-selection heuristics reportedly outperforms the fine-tuned DeBERTa-large baseline by about 9%, reaching claimed state-of-the-ar...

Pith tools