REVIEW 5 cited by
Portuguese Named Entity Recognition using BERT-CRF
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advances in language representation using neural networks have made it viable to transfer the learned internal states of a trained model to downstream natural language processing tasks, such as named entity recognition (NER) and question answering. It has been shown that the leverage of pre-trained language models improves the overall performance on many tasks and is highly beneficial when labeled data is scarce. In this work, we train Portuguese BERT models and employ a BERT-CRF architecture to the NER task on the Portuguese language, combining the transfer capabilities of BERT with the structured predictions of CRF. We explore feature-based and fine-tuning training strategies for the BERT model. Our fine-tuning approach obtains new state-of-the-art results on the HAREM I dataset, improving the F1-score by 1 point on the selective scenario (5 NE classes) and by 4 points on the total scenario (10 NE classes).
Forward citations
Cited by 5 Pith papers
-
A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala
A new manually annotated English-Tamil-Sinhala parallel NER corpus of 3,835 sentences per language, with benchmarks showing XLM-R outperforms monolingual and Indic models, and a case study where NER output improves En...
-
Learning the Language of NVMe Streams for Ransomware Detection
Transformer models that tokenize NVMe command streams detect ransomware commands and estimate ransomware IO volume better than tabular baselines in the authors' evaluation.
-
Challenges in Expanding Portuguese Resources: A View from Open Information Extraction
The authors introduce OIEC-PT, a 300-sentence Portuguese Open IE corpus with 473 extractions and a formal annotation rule set.
-
Symbol-based entity marker highlighting for enhanced text mining in materials science with generative AI
Symbol-tag entity highlighting before LLM-based structuring improves named entity recognition and downstream knowledge graph extraction in materials science text mining.
-
Bangla Grammatical Error Detection Leveraging Transformer-based Token Classification
A token-classification ensemble of BanglaBERT models with rule-based post-processing detects grammatical errors in Bangla text with a reported Levenshtein distance score of 1.04 (or 1.054, the paper is inconsistent).
Discussion (0). Continue with ORCID to comment.