Pith. sign in

REVIEW 15 cited by

CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.06268 v2 pith:QCMRM4BT submitted 2021-03-10 cs.CL cs.LG

CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review

classification cs.CL cs.LG
keywords contractcuaddatasetlegalreviewatticusexpertslarge
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Many specialized domains remain untouched by deep learning, as large labeled datasets require expensive expert annotators. We address this bottleneck within the legal domain by introducing the Contract Understanding Atticus Dataset (CUAD), a new dataset for legal contract review. CUAD was created with dozens of legal experts from The Atticus Project and consists of over 13,000 annotations. The task is to highlight salient portions of a contract that are important for a human to review. We find that Transformer models have nascent performance, but that this performance is strongly influenced by model design and training dataset size. Despite these promising results, there is still substantial room for improvement. As one of the only large, specialized NLP benchmarks annotated by experts, CUAD can serve as a challenging research benchmark for the broader NLP community.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

    cs.CL 2026-05 unverdicted novelty 8.0

    Multi-Legal-Bench creates a sparse 5x6 task-jurisdiction matrix across six countries and reports that few-shot effects replicate, no model dominates, cross-lingual transfer tracks label alignment more than language fa...

  2. UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning

    cs.CL 2026-05 unverdicted novelty 7.0

    UA-Legal-Bench is a new five-task benchmark for Ukrainian legal reasoning that demonstrates task-dependent few-shot prompting effects and the need for macro-F1 over accuracy on imbalanced classes.

  3. LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

    cs.CL 2026-07 conditional novelty 6.0

    A new benchmark, LegalCiteTrust, measures citation existence, fidelity, and applicability in Chinese legal research reports and shows that more legal retrieval does not automatically make citations more trustworthy.

  4. Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards

    cs.LG 2026-07 reject novelty 6.0

    Non-vacuous PAC-Bayes generalization bounds for billion-parameter RLVR models, obtained by a Gumbel-max reparameterization and aggressive TinyLoRA distillation/quantization, are claimed for four tasks.

  5. Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification

    cs.AI 2026-07 accept novelty 6.0

    A constraint-aware hierarchical search over a regulatory tree, using local candidates plus structured rule fields, beats strong RAG baselines on four new expert-validated benchmarks.

  6. LegalCiteBench: Evaluating Citation Reliability in Legal Language Models

    cs.CL 2026-05 unverdicted novelty 6.0

    LegalCiteBench reveals that current LLMs achieve under 7% accuracy on closed-book legal citation retrieval and completion tasks, with misleading answer rates above 94% for nearly all tested models.

  7. Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis

    cs.CV 2026-07 conditional novelty 5.0

    Scene-aware multi-agent document synthesis plus error-driven hard-example expansion improves compact Qwen3-VL models on constrained and open-category KIE, topping reported on-device baselines.

  8. Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free

    cs.CL 2026-05 unverdicted novelty 5.0

    Retrieval with frozen embeddings and k-NN delivers competitive accuracy, high data efficiency, and zero hallucinations on legal multi-label annotation across ECtHR and Eurlex datasets.

  9. A Few Good Clauses: Comparing LLMs vs Domain-Trained Small Language Models on Structured Contract Extraction

    cs.CL 2026-05 unverdicted novelty 5.0

    Domain-trained small language model Olava Extract outperforms frontier LLMs on structured contract extraction with macro F1 0.812, micro F1 0.842, highest precision, and 78-97% lower inference cost.

  10. A Benchmark for Gap and Overlap Analysis as a Test of KG Task Readiness

    cs.AI 2026-04 unverdicted novelty 5.0

    The paper releases a benchmark of ten life-insurance contracts, a domain ontology, and 58 evidence-linked scenarios that shows ontology-driven knowledge graph queries produce more consistent and diagnosable gap/overla...

  11. LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents

    cs.AI 2025-09 reject novelty 5.0

    On CUAD legal contracts, a prompt-engineered QWEN-2 pipeline with chunking and two answer-selection heuristics reportedly outperforms the fine-tuned DeBERTa-large baseline by about 9%, reaching claimed state-of-the-ar...

  12. Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)

    cs.CL 2026-06 unverdicted novelty 4.0

    A graph-augmented RAG system with vector and graph query tools halves hallucinations and raises factual correctness scores on the MoNaCo complex QA benchmark.

  13. Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)

    cs.CL 2026-06 conditional novelty 4.0

    Adding handwritten Cypher graph tools to an agentic RAG system roughly doubled factual-correctness precision and recall on MoNaCo complex questions and improved fine-grained truthfulness compared with vector-only RAG.

  14. Hybrid Topic-Semantic Labeling and Graph Embeddings for Unsupervised Legal Document Clustering

    stat.ML 2025-08 reject novelty 4.0

    Concatenating Top2Vec and Node2Vec embeddings, where the Node2Vec graph encodes Top2Vec's own topic labels, yields compact clusters, but the gain is largely circular.

  15. L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit

    cs.AI 2025-08 reject novelty 4.0

    The abstract reports that a judge-driven multi-agent loop improves legal-citation faithfulness from 0.13 to 0.25 strict F1 and cuts the no-citation rate from 34% to 13%, but the provided manuscript text does not conta...