Pith. sign in

REVIEW 1 cited by

Large Scale Legal Text Classification Using Transformer Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.12871 v1 pith:7OS6E4A5 submitted 2020-10-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords classificationlegaltextcreateddatasetseurlex57keurovocgradual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large multi-label text classification is a challenging Natural Language Processing (NLP) problem that is concerned with text classification for datasets with thousands of labels. We tackle this problem in the legal domain, where datasets, such as JRC-Acquis and EURLEX57K labeled with the EuroVoc vocabulary were created within the legal information systems of the European Union. The EuroVoc taxonomy includes around 7000 concepts. In this work, we study the performance of various recent transformer-based models in combination with strategies such as generative pretraining, gradual unfreezing and discriminative learning rates in order to reach competitive classification performance, and present new state-of-the-art results of 0.661 (F1) for JRC-Acquis and 0.754 for EURLEX57K. Furthermore, we quantify the impact of individual steps, such as language model fine-tuning or gradual unfreezing in an ablation study, and provide reference dataset splits created with an iterative stratification algorithm.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CHANCERY: Evaluating Corporate Governance Reasoning Capabilities in Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A 502-question benchmark for corporate governance reasoning shows current language models reach at most 78.1 percent accuracy.

Pith tools