Pith. sign in

REVIEW 7 cited by

ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1902.07669 v3 pith:DUKSCABA submitted 2019-02-20 cs.CL

classification cs.CL
keywords processingmodelsscispacybiomedicallanguagenaturaltextavailable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite recent advances in natural language processing, many statistical models for processing text perform extremely poorly under domain shift. Processing biomedical and clinical text is a critically important application area of natural language processing, for which there are few robust, practical, publicly available models. This paper describes scispaCy, a new tool for practical biomedical/scientific text processing, which heavily leverages the spaCy library. We detail the performance of two packages of models released in scispaCy and demonstrate their robustness on several tasks and datasets. Models and code are available at https://allenai.github.io/scispacy/

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    MedLayBench-V is the first large-scale multimodal benchmark for expert-lay semantic alignment in medical vision-language models, constructed via a Structured Concept-Grounded Refinement pipeline that uses UMLS CUIs to...

  2. MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking

    cs.CL 2025-11 conditional novelty 6.0 of 10

    MedPath combines 513k+ expert-annotated biomedical mentions into a UMLS-normalized dataset with cross-vocabulary mappings and hierarchical paths for 11 vocabularies.

  3. HIVMedQA: Benchmarking large language models for HIV medical decision support

    cs.CL 2025-07 conditional novelty 6.0 of 10

    The HIVMedQA benchmark finds that Gemini 2.5 Pro leads on most clinical reasoning dimensions, medical fine-tuning does not guarantee gains, and LLM judges are more informative than lexical overlap.

  4. A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models

    cs.LG 2025-07 reject novelty 4.0 of 10

    A survey that taxonomizes EHR modeling research into data-centric, architectural, learning-focused, multimodal, and LLM-based categories, with datasets and metrics.

  5. DS@GT at CheckThat! 2025: Ensemble Methods for Detection of Scientific Discourse on Social Media

    cs.CL 2025-07 conditional novelty 4.0 of 10

    The DS@GT system achieved 0.8611 macro-F1 on the CheckThat! 2025 Task 4a development set by combining a fine-tuned DeBERTa model with GPT-4o few-shot prompting, outperforming the DeBERTaV3 baseline of 0.8375.

  6. Charting the Future of Scholarly Knowledge with AI: A Community Perspective

    cs.DL 2025-08 unverdicted novelty 2.0 of 10

    A community perspective on how AI can support scholarly knowledge extraction, organization, and communication, with a proposed classification and ethical considerations.

  7. Building Entity Association Mining Framework for Knowledge Discovery

    cs.CL 2025-06 reject novelty 2.0 of 10

    The paper presents a pluggable co-occurrence mining framework for financial text and illustrates it with two anecdotal use cases, without releasing artifacts or quantitative evaluation.

Pith tools