Pith. sign in

REVIEW 16 cited by

LEGAL-BERT: The Muppets straight out of Law School

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.02559 v1 pith:L23H3QAY submitted 2020-10-06 cs.CL

classification cs.CL
keywords bertlegaltasksapplyingcorporadomaindomain-specificdomains
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

BERT has achieved impressive performance in several NLP tasks. However, there has been limited investigation on its adaptation guidelines in specialised domains. Here we focus on the legal domain, where we explore several approaches for applying BERT models to downstream legal tasks, evaluating on multiple datasets. Our findings indicate that the previous guidelines for pre-training and fine-tuning, often blindly followed, do not always generalize well in the legal domain. Thus we propose a systematic investigation of the available strategies when applying BERT in specialised domains. These are: (a) use the original BERT out of the box, (b) adapt BERT by additional pre-training on domain-specific corpora, and (c) pre-train BERT from scratch on domain-specific corpora. We also propose a broader hyper-parameter search space when fine-tuning for downstream tasks and we release LEGAL-BERT, a family of BERT models intended to assist legal NLP research, computational law, and legal technology applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SteuerLLM: Local specialized large language model for German tax law analysis

    cs.CL 2026-02 reject novelty 6.0 of 10

    A tax-specialized 28B model beats larger general-purpose LLMs on a new authentic German tax-law exam benchmark, but its edge may be inflated by overlap between training and evaluation exams.

  2. Health Insurance Coverage Rule Interpretation Corpus: Law, Policy, and Medical Guidance for Health Insurance Coverage Understanding

    cs.CY 2025-07 conditional novelty 6.0 of 10

    A new corpus and a pseudo-annotated benchmark for predicting health insurance external appeal outcomes, with baseline transformer models.

  3. The Missing Link: Joint Legal Citation Prediction using Heterogeneous Graph Enrichment

    cs.SI 2025-06 accept novelty 6.0 of 10

    A graph neural network that enriches legal citation graphs with categorical metadata nodes predicts case and law citations more accurately than prior GNN baselines, and joint training boosts case citation prediction.

  4. BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    BriefMe introduces a legal brief benchmark with argument summarization, argument completion, and case retrieval, and shows LLMs beat human headings on the first two but struggle on the latter two.

  5. CHANCERY: Evaluating Corporate Governance Reasoning Capabilities in Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A 502-question benchmark for corporate governance reasoning shows current language models reach at most 78.1 percent accuracy.

  6. LAARA: Layer-Aware Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    LAARA allocates LoRA ranks per layer from diagonal Fisher (gradient-based) estimates, reporting improved accuracy with fewer trainable parameters on GLUE and MathInstruct.

  7. LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents

    cs.AI 2025-09 reject novelty 5.0 of 10

    On CUAD legal contracts, a prompt-engineered QWEN-2 pipeline with chunking and two answer-selection heuristics reportedly outperforms the fine-tuned DeBERTa-large baseline by about 9%, reaching claimed state-of-the-ar...

  8. EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes

    cs.LG 2025-07 unverdicted novelty 5.0 of 10

    A bias-corrected exponential moving average (BEMA) is claimed to remove the lag of standard EMA weight averaging during LLM fine-tuning, improving convergence and final performance over EMA and vanilla training.

  9. TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Trident-Bench provides 2,652 professionally validated harmful prompts across finance, law, and medicine, and shows that domain-specialized LLMs often comply with unethical requests more than generalist models.

  10. PromptAL: Sample-Aware Dynamic Soft Prompts for Few-Shot Active Learning

    cs.CL 2025-07 conditional novelty 5.0 of 10

    PromptAL combines sample-aware dynamic soft prompts with uncertainty and diversity scores to select better annotation candidates in few-shot active learning.

  11. ASP2LJ : An Adversarial Self-Play Laywer Augmented Legal Judgment Framework

    cs.CL 2025-06 conditional novelty 5.0 of 10

    ASP2LJ combines synthetic case generation with adversarial self-play for lawyer agents, improving legal judgment prediction on a Chinese benchmark and on a new rare-case dataset.

  12. Hybrid Topic-Semantic Labeling and Graph Embeddings for Unsupervised Legal Document Clustering

    stat.ML 2025-08 reject novelty 4.0 of 10

    Concatenating Top2Vec and Node2Vec embeddings, where the Node2Vec graph encodes Top2Vec's own topic labels, yields compact clusters, but the gain is largely circular.

  13. L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit

    cs.AI 2025-08 reject novelty 4.0 of 10

    The abstract reports that a judge-driven multi-agent loop improves legal-citation faithfulness from 0.13 to 0.25 strict F1 and cuts the no-citation rate from 34% to 13%, but the provided manuscript text does not conta...

  14. A Data Science Approach to Calcutta High Court Judgments: An Efficient LLM and RAG-powered Framework for Summarization and Similar Cases Retrieval

    cs.IR 2025-06 reject novelty 4.0 of 10

    Fine-tuning Pegasus on LLM-annotated headnotes improves part of the legal summarization pipeline, and a RAG framework retrieves similar Calcutta High Court cases, though retrieval quality is never measured.

  15. CoLA: Collaborative Low-Rank Adaptation

    cs.CL 2025-05 conditional novelty 4.0 of 10

    CoLA generalizes LoRA to multiple A and B matrices with a principal-component initialization and reports gains of roughly 2-4 accuracy points over PiSSA on low-sample fine-tuning benchmarks.

  16. LLMPR: A Novel LLM-Driven Transfer Learning based Petition Ranking Model

    cs.CL 2025-05 reject novelty 2.0 of 10

    A petition-ranking model that reports near-perfect accuracy, but its target ranking is derived from the same gap-days features it feeds the model, making the result circular.

Pith tools