Pith. sign in

REVIEW 5 cited by

Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.00220 v2 pith:WAI4EXNW submitted 2022-07-01 cs.CL cs.CY

classification cs.CLcs.CY
keywords filteringlegalpiledatadatasetpretrainingadministrativeapproaches
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

One concern with the rise of large language models lies with their potential for significant harm, particularly from pretraining on biased, obscene, copyrighted, and private information. Emerging ethical approaches have attempted to filter pretraining material, but such approaches have been ad hoc and failed to take context into account. We offer an approach to filtering grounded in law, which has directly addressed the tradeoffs in filtering material. First, we gather and make available the Pile of Law, a 256GB (and growing) dataset of open-source English-language legal and administrative data, covering court opinions, contracts, administrative rules, and legislative records. Pretraining on the Pile of Law may help with legal tasks that have the promise to improve access to justice. Second, we distill the legal norms that governments have developed to constrain the inclusion of toxic or private content into actionable lessons for researchers and discuss how our dataset reflects these norms. Third, we show how the Pile of Law offers researchers the opportunity to learn such filtering rules directly from the data, providing an exciting new research direction in model-based processing.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark

    cs.IR 2025-08 conditional novelty 7.0 of 10

    State-of-the-art LLMs with retrieval answer simplified boolean questions about state unemployment insurance law with at best 0.69 F1, well short of reliable end-to-end code simplification.

  2. Teach Old SAEs New Domain Tricks with Boosting

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Training a small secondary sparse autoencoder on the reconstruction error of a pretrained SAE improves domain-specific reconstruction and language-model perplexity without hurting general performance.

  3. DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domains

    cs.CL 2025-06 conditional novelty 6.0 of 10

    DivScore detects AI-written medical and legal text by dividing a domain-tuned model's entropy by its disagreement with a general model, beating baselines on a new benchmark.

  4. Inteligencia Artificial jur\'idica y el desaf\'io de la veracidad: an\'alisis de alucinaciones, optimizaci\'on de RAG y principios para una integraci\'on responsable

    cs.AI 2025-09 conditional novelty 4.0 of 10

    Legal AI hallucination persists in commercial RAG tools (17-34%+ of queries), so the report argues the fix is consultative, source-citing system design plus mandatory human oversight, not better generative models.

  5. L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit

    cs.AI 2025-08 reject novelty 4.0 of 10

    The abstract reports that a judge-driven multi-agent loop improves legal-citation faithfulness from 0.13 to 0.25 strict F1 and cuts the no-citation rate from 34% to 13%, but the provided manuscript text does not conta...

Pith tools