Pith. sign in

REVIEW 4 cited by

Pre-trained Language Models for the Legal Domain: A Case Study on Indian Law

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.06049 v5 pith:RMXYPL7Y submitted 2022-09-13 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords legalindiandomainplmstextcountriespre-trainedcourt
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

NLP in the legal domain has seen increasing success with the emergence of Transformer-based Pre-trained Language Models (PLMs) pre-trained on legal text. PLMs trained over European and US legal text are available publicly; however, legal text from other domains (countries), such as India, have a lot of distinguishing characteristics. With the rapidly increasing volume of Legal NLP applications in various countries, it has become necessary to pre-train such LMs over legal text of other countries as well. In this work, we attempt to investigate pre-training in the Indian legal domain. We re-train (continue pre-training) two popular legal PLMs, LegalBERT and CaseLawBERT, on Indian legal data, as well as train a model from scratch with a vocabulary based on Indian legal text. We apply these PLMs over three benchmark legal NLP tasks -- Legal Statute Identification from facts, Semantic Segmentation of Court Judgment Documents, and Court Appeal Judgment Prediction -- over both Indian and non-Indian (EU, UK) datasets. We observe that our approach not only enhances performance on the new domain (Indian texts) but also over the original domain (European and UK texts). We also conduct explainability experiments for a qualitative comparison of all these different PLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LegalSeg: Unlocking the Structure of Indian Legal Judgments Through Rhetorical Role Classification

    cs.CL 2025-02 reject novelty 6.0 of 10

    LegalSeg provides 7,120 Indian judgments annotated with seven rhetorical roles and benchmarks several models, reporting that context-aware sequence models work best.

  2. NyayaAnumana & INLegalLlama: The Largest Indian Legal Judgment Prediction Dataset and Specialized Language Model for Enhanced Decision Analysis

    cs.CL 2024-12 reject novelty 6.0 of 10

    A new large corpus of Indian court cases and a legal LLaMA model report very high judgment-prediction accuracy, but the evaluation leaks the outcome from the input text.

  3. AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System

    cs.CL 2026-07 conditional novelty 4.0 of 10

    RAG with top-3 chunk retrieval lifts smaller LLMs on Indian legal QA (Llama2-70B: 45.7% to 51.7% on AIBE) but often hurts large models, and under the study's own rating protocol some AI answers outscored the reference...

  4. A Comprehensive Survey on Legal Summarization: Challenges and Future Directions

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A systematic survey of legal summarization finds a field dominated by English common-law datasets, ROUGE-based evaluation, and few human or expert validation studies.

Pith tools