Pith. sign in

REVIEW 7 cited by

LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.00976 v4 pith:JESY6CKG submitted 2021-10-03 cs.CL

classification cs.CL
keywords legalacrosslanguagetasksunderstandinganalysisbenchmarkevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Laws and their interpretations, legal arguments and agreements\ are typically expressed in writing, leading to the production of vast corpora of legal text. Their analysis, which is at the center of legal practice, becomes increasingly elaborate as these collections grow in size. Natural language understanding (NLU) technologies can be a valuable tool to support legal practitioners in these endeavors. Their usefulness, however, largely depends on whether current state-of-the-art models can generalize across various tasks in the legal domain. To answer this currently open question, we introduce the Legal General Language Understanding Evaluation (LexGLUE) benchmark, a collection of datasets for evaluating model performance across a diverse set of legal NLU tasks in a standardized way. We also provide an evaluation and analysis of several generic and legal-oriented models demonstrating that the latter consistently offer performance improvements across multiple tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents

    cs.AI 2025-09 reject novelty 5.0 of 10

    On CUAD legal contracts, a prompt-engineered QWEN-2 pipeline with chunking and two answer-selection heuristics reportedly outperforms the fine-tuned DeBERTa-large baseline by about 9%, reaching claimed state-of-the-ar...

  2. Standard Applicability Judgment and Cross-jurisdictional Reasoning: A RAG-based Framework for Medical Device Compliance

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A retrieval-augmented system classifies applicability of Chinese and US medical device standards from free-text device descriptions, reporting 73% accuracy and 87% top-5 recall on a 105-item benchmark.

  3. Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE

    cs.AI 2025-09 reject novelty 4.0 of 10

    KG-SMILE applies perturbation and linear regression to a knowledge graph to attribute which entities and relations drive a GraphRAG system's answers.

  4. Scaling Legal AI: Benchmarking Mamba and Transformers for Statutory Classification and Case Law Retrieval

    cs.CY 2025-08 reject novelty 4.0 of 10

    A benchmark reports that SSD-Mamba matches or surpasses transformer baselines on legal classification and retrieval with higher throughput, but omits reproducible experimental details.

  5. The Judge Variable: Challenging Judge-Agnostic Legal Judgment Prediction

    cs.CL 2025-07 reject novelty 4.0 of 10

    Models trained on individual judges' past child-custody rulings predict those judges' future rulings better than a model trained on all judges together, a result the paper reads as support for legal realism.

  6. When Large Language Models Meet Law: Dual-Lens Taxonomy, Technical Advances, and Ethical Governance

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A literature review that classifies LLM-for-law research using a dual-lens taxonomy of Toulmin argumentation components and legal practitioner roles.

  7. Enhancing Large Language Models with Reliable Knowledge Graphs

    cs.CL 2025-06 conditional novelty 2.0 of 10

    A thesis composed of four published papers proposes contrastive KG error detection, attribute-aware error-aware embedding, inductive graph completion, and KG prompting, but adds no new result beyond those papers.

Pith tools