Pith. sign in

REVIEW 3 cited by

MultiLegalPile: A 689GB Multilingual Legal Corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.02069 v3 pith:UB7ABR6N submitted 2023-06-03 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords modelscorpusenglishlegallicensesmultilegalpilemultilingualavailable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large, high-quality datasets are crucial for training Large Language Models (LLMs). However, so far, there are few datasets available for specialized critical domains such as law and the available ones are often only for the English language. We curate and release MultiLegalPile, a 689GB corpus in 24 languages from 17 jurisdictions. The MultiLegalPile corpus, which includes diverse legal data sources with varying licenses, allows for pretraining NLP models under fair use, with more permissive licenses for the Eurlex Resources and Legal mC4 subsets. We pretrain two RoBERTa models and one Longformer multilingually, and 24 monolingual models on each of the language-specific subsets and evaluate them on LEXTREME. Additionally, we evaluate the English and multilingual models on LexGLUE. Our multilingual models set a new SotA on LEXTREME and our English models on LexGLUE. We release the dataset, the trained models, and all of the code under the most open possible licenses.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new wheat-specific dataset with pretraining, quantitative, and instruction-tuning layers improves VLM performance on wheat stress diagnosis and growth-stage management tasks.

  2. LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents

    cs.AI 2025-09 reject novelty 5.0 of 10

    On CUAD legal contracts, a prompt-engineered QWEN-2 pipeline with chunking and two answer-selection heuristics reportedly outperforms the fine-tuned DeBERTa-large baseline by about 9%, reaching claimed state-of-the-ar...

  3. A Survey of Classification Tasks and Approaches for Legal Contracts

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A survey organizing legal contract classification into seven tasks, fourteen datasets, and a three-part methodology taxonomy.

Pith tools