Pith. sign in

REVIEW 5 cited by

Portuguese Named Entity Recognition using BERT-CRF

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.10649 v2 pith:7FQGAPH2 submitted 2019-09-23 cs.CL cs.IRcs.LG

classification cs.CLcs.IRcs.LG
keywords languagebertportuguesebert-crfclassesentityfine-tuningmodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advances in language representation using neural networks have made it viable to transfer the learned internal states of a trained model to downstream natural language processing tasks, such as named entity recognition (NER) and question answering. It has been shown that the leverage of pre-trained language models improves the overall performance on many tasks and is highly beneficial when labeled data is scarce. In this work, we train Portuguese BERT models and employ a BERT-CRF architecture to the NER task on the Portuguese language, combining the transfer capabilities of BERT with the structured predictions of CRF. We explore feature-based and fine-tuning training strategies for the BERT model. Our fine-tuning approach obtains new state-of-the-art results on the HAREM I dataset, improving the F1-score by 1 point on the selective scenario (5 NE classes) and by 4 points on the total scenario (10 NE classes).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala

    cs.CL 2024-12 conditional novelty 7.0 of 10

    A new manually annotated English-Tamil-Sinhala parallel NER corpus of 3,835 sentences per language, with benchmarks showing XLM-R outperforms monolingual and Indic models, and a case study where NER output improves En...

  2. Learning the Language of NVMe Streams for Ransomware Detection

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Transformer models that tokenize NVMe command streams detect ransomware commands and estimate ransomware IO volume better than tabular baselines in the authors' evaluation.

  3. Challenges in Expanding Portuguese Resources: A View from Open Information Extraction

    cs.CL 2025-01 conditional novelty 6.0 of 10

    The authors introduce OIEC-PT, a 300-sentence Portuguese Open IE corpus with 473 extractions and a formal annotation rule set.

  4. Symbol-based entity marker highlighting for enhanced text mining in materials science with generative AI

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Symbol-tag entity highlighting before LLM-based structuring improves named entity recognition and downstream knowledge graph extraction in materials science text mining.

  5. Bangla Grammatical Error Detection Leveraging Transformer-based Token Classification

    cs.CL 2024-11 conditional novelty 4.0 of 10

    A token-classification ensemble of BanglaBERT models with rule-based post-processing detects grammatical errors in Bangla text with a reported Levenshtein distance score of 1.04 (or 1.054, the paper is inconsistent).

Pith tools