Pith. sign in

REVIEW 1 cited by

The Right Model for the Job: An Evaluation of Legal Multi-Label Classification Baselines

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.11852 v1 pith:6L34WZPE submitted 2024-01-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords legalapproachesclassificationcomputationaldifferentevaluationlabelmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-Label Classification (MLC) is a common task in the legal domain, where more than one label may be assigned to a legal document. A wide range of methods can be applied, ranging from traditional ML approaches to the latest Transformer-based architectures. In this work, we perform an evaluation of different MLC methods using two public legal datasets, POSTURE50K and EURLEX57K. By varying the amount of training data and the number of labels, we explore the comparative advantage offered by different approaches in relation to the dataset properties. Our findings highlight DistilRoBERTa and LegalBERT as performing consistently well in legal MLC with reasonable computational demands. T5 also demonstrates comparable performance while offering advantages as a generative model in the presence of changing label sets. Finally, we show that the CrossEncoder exhibits potential for notable macro-F1 score improvements, albeit with increased computational costs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Supervised Machine Learning Approach for Assessing Grant Peer Review Reports

    econ.EM 2024-11 conditional novelty 6.0 of 10

    Fine-tuned transformer models classify Swiss National Science Foundation grant peer review sentences into twelve content categories with an average macro F1 of 0.85, enabling large-scale analysis of review reports.

Pith tools