Pith. sign in

REVIEW 1 cited by

Named Entity Recognition in the Legal Domain using a Pointer Generator Network

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.09936 v1 pith:OSZEDD7E submitted 2020-12-17 cs.CL cs.IRcs.LG

classification cs.CLcs.IRcs.LG
keywords dataentitiestextlegalnamedentitygeneratorgold
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Named Entity Recognition (NER) is the task of identifying and classifying named entities in unstructured text. In the legal domain, named entities of interest may include the case parties, judges, names of courts, case numbers, references to laws etc. We study the problem of legal NER with noisy text extracted from PDF files of filed court cases from US courts. The "gold standard" training data for NER systems provide annotation for each token of the text with the corresponding entity or non-entity label. We work with only partially complete training data, which differ from the gold standard NER data in that the exact location of the entities in the text is unknown and the entities may contain typos and/or OCR mistakes. To overcome the challenges of our noisy training data, e.g. text extraction errors and/or typos and unknown label indices, we formulate the NER task as a text-to-text sequence generation task and train a pointer generator network to generate the entities in the document rather than label them. We show that the pointer generator can be effective for NER in the absence of gold standard data and outperforms the common NER neural network architectures in long legal documents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection

    cs.CL 2024-12 conditional novelty 6.0 of 10

    PICLe shows that pseudo-annotated, partially correct demonstrations can replace gold labels for in-context named entity detection, outperforming few-shot ICL with scarce gold examples on biomedical datasets.

Pith tools