Pith. sign in

REVIEW 1 cited by

Biomedical Named Entity Recognition via Dictionary-based Synonym Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13066 v2 pith:VG7ROZYZ submitted 2023-05-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords synonymgeneralizationapproachesbiomedicaldictionary-basedapproachnamedsyngen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Biomedical named entity recognition is one of the core tasks in biomedical natural language processing (BioNLP). To tackle this task, numerous supervised/distantly supervised approaches have been proposed. Despite their remarkable success, these approaches inescapably demand laborious human effort. To alleviate the need of human effort, dictionary-based approaches have been proposed to extract named entities simply based on a given dictionary. However, one downside of existing dictionary-based approaches is that they are challenged to identify concept synonyms that are not listed in the given dictionary, which we refer as the synonym generalization problem. In this study, we propose a novel Synonym Generalization (SynGen) framework that recognizes the biomedical concepts contained in the input text using span-based predictions. In particular, SynGen introduces two regularization terms, namely, (1) a synonym distance regularizer; and (2) a noise perturbation regularizer, to minimize the synonym generalization error. To demonstrate the effectiveness of our approach, we provide a theoretical analysis of the bound of synonym generalization error. We extensively evaluate our approach on a wide range of benchmarks and the results verify that SynGen outperforms previous dictionary-based models by notable margins. Lastly, we provide a detailed analysis to further reveal the merits and inner-workings of our approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Extracting Structured Requirements from Unstructured Building Technical Specifications for Building Information Modeling

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    A study showing that CamemBERT and Fr_core_news_lg achieve over 90% F1 for named entity recognition and Random Forest achieves over 80% F1 for relation extraction on French building technical specifications.

Pith tools