Pith. sign in

REVIEW 1 cited by

Text Chunking using Transformation-Based Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv cmp-lg/9505040 v1 pith:CH5E7N7T submitted 1995-05-23 cmp-lg cs.CL

classification cmp-lgcs.CL
keywords chunkslearningtransformation-basedbasenpchunkingtaggingtextaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Eric Brill introduced transformation-based learning and showed that it can do part-of-speech tagging with fairly high accuracy. The same method can be applied at a higher level of textual interpretation for locating chunks in the tagged text, including non-recursive ``baseNP'' chunks. For this purpose, it is convenient to view chunking as a tagging problem by encoding the chunk structure in new tags attached to each word. In automatic tests using Treebank-derived data, this technique achieved recall and precision rates of roughly 92% for baseNP chunks and 88% for somewhat more complex chunks that partition the sentence. Some interesting adaptations to the transformation-based learning approach are also suggested by this application.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MariNER: A Dataset for Historical Brazilian Portuguese Named Entity Recognition

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MariNER is a new manually annotated NER dataset for early 20th-century Brazilian Portuguese, with benchmark results showing fine-tuned transformers outperform large language models.

Pith tools