Pith. sign in

REVIEW 27 cited by

SciBERT: A Pretrained Language Model for Scientific Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.10676 v3 pith:RZL2AYFJ submitted 2019-03-26 cs.CL

SciBERT: A Pretrained Language Model for Scientific Text

classification cs.CL
keywords scientificsciberttaskspretrainedbertdatalanguagelarge-scale
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Obtaining large-scale annotated data for NLP tasks in the scientific domain is challenging and expensive. We release SciBERT, a pretrained language model based on BERT (Devlin et al., 2018) to address the lack of high-quality, large-scale labeled scientific data. SciBERT leverages unsupervised pretraining on a large multi-domain corpus of scientific publications to improve performance on downstream scientific NLP tasks. We evaluate on a suite of tasks including sequence tagging, sentence classification and dependency parsing, with datasets from a variety of scientific domains. We demonstrate statistically significant improvements over BERT and achieve new state-of-the-art results on several of these tasks. The code and pretrained models are available at https://github.com/allenai/scibert/.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding

    cs.CV 2026-01 unverdicted novelty 8.0

    S1-MMAlign is a new large-scale dataset of 15.5 million semantically enhanced scientific image-text pairs created via an AI recaptioning pipeline to improve multimodal understanding.

  2. From Efficiency to Leakage -- Privacy Backdoor in Federated Language Model Fine-Tuning

    cs.CR 2026-06 conditional novelty 7.0

    NeuroImprint attack assigns isolated memorization neurons to training samples in PEFT adapters, enabling closed-form reconstruction of 59-79% of samples across BERT, GPT-2, Qwen2, and Llama3.2 on multiple datasets.

  3. OOD-GraphLLM: Graph Large Language Model for Out-of-Distribution Generalized Drug Synergy Prediction

    cs.LG 2026-05 unverdicted novelty 7.0

    OOD-GraphLLM is a graphLLM framework that jointly optimizes molecular graph representations and biomedical semantic language representations for out-of-distribution drug synergy prediction.

  4. The software space of science

    cs.DL 2026-04 unverdicted novelty 7.0

    A network analysis of software mentions in 1.3 million papers identifies 520 tools in eight communities and shows disciplines maintain distinct, stable tool portfolios that are crystallizing toward common sets.

  5. MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature

    cs.IR 2026-04 unverdicted novelty 7.0

    MasterSet is a new large-scale benchmark for must-cite citation recommendation in AI/ML, using LLM-annotated tiers on 150k papers and Recall@K evaluation.

  6. PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark

    cs.CV 2026-07 conditional novelty 6.5

    A placeholder-first harness separates visual poster design from scientific figure grounding, turning poster generation into measurable instruction-following with a 12-paper pilot and failure taxonomy.

  7. LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction

    cs.CL 2026-07 conditional novelty 6.0

    Label-aware diagnostic reflection plus two-stage outcome GRPO improves same-backbone IE F1 over SFT, with larger gains under relation-extraction domain shift.

  8. Examining the Cognitive Gap Between Authors and Peer Reviewers on Academic Paper Novelty

    cs.DL 2026-06 unverdicted novelty 6.0

    Empirical analysis of peer review data reveals that author promotional language correlates with reviewer novelty disagreement only for moderately innovative papers.

  9. Aspect-Aware Content-Based Recommendations for Mathematical Research Papers

    cs.IR 2026-05 unverdicted novelty 6.0

    The authors introduce aspect-aware datasets GoldRiM and SilverRiM for math papers and AchGNN, a heterogeneous GNN that outperforms prior methods by jointly modeling textual semantics, citations, and author lineage acr...

  10. Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains

    cs.CL 2026-04 unverdicted novelty 6.0

    Automatic translation metrics show lower agreement with humans on unseen technical domains than humans show with each other, and their robustness claims weaken when benchmarked against inter-annotator agreement instea...

  11. Teaching Robots to Interpret Social Interactions through Lexically-guided Dynamic Graph Learning

    cs.HC 2026-04 conditional novelty 6.0

    A dynamic graph multi-task network with lexical task tokens improves joint prediction of intent, attitude, and actions in human-robot interaction and learns a time-varying task-affinity matrix.

  12. Teaching Robots to Interpret Social Interactions through Lexically-guided Dynamic Graph Learning

    cs.HC 2026-04 unverdicted novelty 6.0

    SocialLDG models six socio-cognitive tasks with lexical priors from language models and time-evolving task affinities via dynamic graphs, claiming state-of-the-art results on two public human-robot interaction dataset...

  13. Teaching Robots to Interpret Social Interactions through Lexically-guided Dynamic Graph Learning

    cs.HC 2026-04 unverdicted novelty 6.0

    SocialLDG is a multi-task framework using language models for lexical priors and dynamic graphs to model evolving task affinities among six social interaction tasks, claiming SOTA results on two public HRI datasets pl...

  14. CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics

    cs.CL 2025-09 unverdicted novelty 6.0

    CFDLLMBench is a new benchmark suite with CFDQuery, CFDCodeBench, and FoamBench to evaluate LLMs on graduate-level CFD knowledge, numerical reasoning, and context-dependent code implementation.

  15. Large Language Models for Market Research: A Data-augmentation Approach

    cs.AI 2024-12 unverdicted novelty 6.0

    A data-augmentation framework for conjoint analysis integrates LLM-generated data with human responses to yield consistent, asymptotically normal estimators and reported cost savings of 24.9-79.8% in two empirical studies.

  16. Gender Differences in Research Topic and Method Convergence among Collaborating Scholars in Library and Information Science

    cs.DL 2026-06 unverdicted novelty 5.0

    Analysis of 25,204 LIS papers finds female collaborating scholars exhibit lower convergence in topics and methods than male scholars.

  17. Explaining Black-Box Language Models: Learning to Optimize Linguistically-Structured Word Subsets

    cs.AI 2026-06 unverdicted novelty 5.0

    Amortized optimization with policy gradients and graph knowledge selects informative word subsets to explain black-box DLM outputs.

  18. Traditional statistical representations outperform generative AI in identifying expert peer reviewers

    cs.IR 2026-05 unverdicted novelty 5.0

    TF-IDF identifies labeled experts in the top 25 recommendations 79.5% of the time versus 51.5% for GPT-4o mini on an astronomy observatory dataset.

  19. Surface-Form Neural Sparse Retrieval: Robust Fuzzy Matching for Industrial Music Search

    cs.AI 2026-05 unverdicted novelty 5.0

    A neural sparse retrieval system with granular subword tokenization (max 3 chars) achieves 91.4% recall@10 on a 6M music document corpus versus 57.7% for trigrams, with improved HCI exploration efficiency and zero add...

  20. STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking

    cs.CL 2025-07 unverdicted novelty 5.0

    StructSense is a task-agnostic agentic framework for structured information extraction that reports 91-100% accuracy on schema-based tasks, 86-93% on metadata extraction, and 58-75% NER accuracy while adding entities ...

  21. Galactica: A Large Language Model for Science

    cs.CL 2022-11 unverdicted novelty 5.0

    Galactica, a science-specialized LLM, reports higher scores than GPT-3, Chinchilla, and PaLM on LaTeX knowledge, mathematical reasoning, and medical QA benchmarks while outperforming general models on BIG-bench.

  22. Though Language Models Err While They Strive: Conformal Prediction for Self-Correcting Scientific Generation

    cs.CL 2026-07 reject novelty 4.0

    A graph-based conformal wrapper that filters and regenerates LLM reasoning steps claims formal coverage guarantees on scientific validity, but its evaluation is circular and its gains are confounded with sampling effo...

  23. Exploring the relationship between team institutional composition and novelty in academic papers based on fine-grained knowledge entities

    cs.CL 2026-06 unverdicted novelty 4.0

    Mixed academic-industrial teams in NLP produce more novel papers than purely industrial teams, with mixed teams emphasizing method-metric novelty and industrial teams emphasizing method-tool novelty.

  24. Extracting Breast Cancer Phenotypes from Clinical Notes: Comparing LLMs with Classical Ontology Methods

    cs.CL 2026-03 unverdicted novelty 4.0

    LLM framework extracts breast cancer phenotypes from clinical notes with accuracy comparable to ontology-based methods and greater adaptability to new diseases.

  25. Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts

    cs.CL 2026-04 conditional novelty 3.0

    An ensemble of three Google LLMs achieves a 0.74 weighted F1-score for detecting EQ-5D reporting in 200 PubMed abstracts, marginally outperforming individual models.

  26. Data-Centric Foundation Models in Computational Healthcare: A Survey

    cs.LG 2024-01 unverdicted novelty 3.0

    The paper surveys data-centric strategies for foundation models in computational healthcare and supplies a curated list of related models and datasets.

  27. Natural Language Processing in the Legal Domain

    cs.CL 2023-02 unverdicted novelty 3.0

    A survey of nearly 1000 NLP & Law papers from 2013-2024 documenting increases in publication volume, scope, methodological sophistication, and data/code availability.