REVIEW 27 cited by
SciBERT: A Pretrained Language Model for Scientific Text
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
SciBERT: A Pretrained Language Model for Scientific Text
read the original abstract
Obtaining large-scale annotated data for NLP tasks in the scientific domain is challenging and expensive. We release SciBERT, a pretrained language model based on BERT (Devlin et al., 2018) to address the lack of high-quality, large-scale labeled scientific data. SciBERT leverages unsupervised pretraining on a large multi-domain corpus of scientific publications to improve performance on downstream scientific NLP tasks. We evaluate on a suite of tasks including sequence tagging, sentence classification and dependency parsing, with datasets from a variety of scientific domains. We demonstrate statistically significant improvements over BERT and achieve new state-of-the-art results on several of these tasks. The code and pretrained models are available at https://github.com/allenai/scibert/.
Forward citations
Cited by 27 Pith papers
-
S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding
S1-MMAlign is a new large-scale dataset of 15.5 million semantically enhanced scientific image-text pairs created via an AI recaptioning pipeline to improve multimodal understanding.
-
From Efficiency to Leakage -- Privacy Backdoor in Federated Language Model Fine-Tuning
NeuroImprint attack assigns isolated memorization neurons to training samples in PEFT adapters, enabling closed-form reconstruction of 59-79% of samples across BERT, GPT-2, Qwen2, and Llama3.2 on multiple datasets.
-
OOD-GraphLLM: Graph Large Language Model for Out-of-Distribution Generalized Drug Synergy Prediction
OOD-GraphLLM is a graphLLM framework that jointly optimizes molecular graph representations and biomedical semantic language representations for out-of-distribution drug synergy prediction.
-
The software space of science
A network analysis of software mentions in 1.3 million papers identifies 520 tools in eight communities and shows disciplines maintain distinct, stable tool portfolios that are crystallizing toward common sets.
-
MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature
MasterSet is a new large-scale benchmark for must-cite citation recommendation in AI/ML, using LLM-annotated tiers on 150k papers and Recall@K evaluation.
-
PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark
A placeholder-first harness separates visual poster design from scientific figure grounding, turning poster generation into measurable instruction-following with a 12-paper pilot and failure taxonomy.
-
LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction
Label-aware diagnostic reflection plus two-stage outcome GRPO improves same-backbone IE F1 over SFT, with larger gains under relation-extraction domain shift.
-
Examining the Cognitive Gap Between Authors and Peer Reviewers on Academic Paper Novelty
Empirical analysis of peer review data reveals that author promotional language correlates with reviewer novelty disagreement only for moderately innovative papers.
-
Aspect-Aware Content-Based Recommendations for Mathematical Research Papers
The authors introduce aspect-aware datasets GoldRiM and SilverRiM for math papers and AchGNN, a heterogeneous GNN that outperforms prior methods by jointly modeling textual semantics, citations, and author lineage acr...
-
Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains
Automatic translation metrics show lower agreement with humans on unseen technical domains than humans show with each other, and their robustness claims weaken when benchmarked against inter-annotator agreement instea...
-
Teaching Robots to Interpret Social Interactions through Lexically-guided Dynamic Graph Learning
A dynamic graph multi-task network with lexical task tokens improves joint prediction of intent, attitude, and actions in human-robot interaction and learns a time-varying task-affinity matrix.
-
Teaching Robots to Interpret Social Interactions through Lexically-guided Dynamic Graph Learning
SocialLDG models six socio-cognitive tasks with lexical priors from language models and time-evolving task affinities via dynamic graphs, claiming state-of-the-art results on two public human-robot interaction dataset...
-
Teaching Robots to Interpret Social Interactions through Lexically-guided Dynamic Graph Learning
SocialLDG is a multi-task framework using language models for lexical priors and dynamic graphs to model evolving task affinities among six social interaction tasks, claiming SOTA results on two public HRI datasets pl...
-
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics
CFDLLMBench is a new benchmark suite with CFDQuery, CFDCodeBench, and FoamBench to evaluate LLMs on graduate-level CFD knowledge, numerical reasoning, and context-dependent code implementation.
-
Large Language Models for Market Research: A Data-augmentation Approach
A data-augmentation framework for conjoint analysis integrates LLM-generated data with human responses to yield consistent, asymptotically normal estimators and reported cost savings of 24.9-79.8% in two empirical studies.
-
Gender Differences in Research Topic and Method Convergence among Collaborating Scholars in Library and Information Science
Analysis of 25,204 LIS papers finds female collaborating scholars exhibit lower convergence in topics and methods than male scholars.
-
Explaining Black-Box Language Models: Learning to Optimize Linguistically-Structured Word Subsets
Amortized optimization with policy gradients and graph knowledge selects informative word subsets to explain black-box DLM outputs.
-
Traditional statistical representations outperform generative AI in identifying expert peer reviewers
TF-IDF identifies labeled experts in the top 25 recommendations 79.5% of the time versus 51.5% for GPT-4o mini on an astronomy observatory dataset.
-
Surface-Form Neural Sparse Retrieval: Robust Fuzzy Matching for Industrial Music Search
A neural sparse retrieval system with granular subword tokenization (max 3 chars) achieves 91.4% recall@10 on a 6M music document corpus versus 57.7% for trigrams, with improved HCI exploration efficiency and zero add...
-
STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking
StructSense is a task-agnostic agentic framework for structured information extraction that reports 91-100% accuracy on schema-based tasks, 86-93% on metadata extraction, and 58-75% NER accuracy while adding entities ...
-
Galactica: A Large Language Model for Science
Galactica, a science-specialized LLM, reports higher scores than GPT-3, Chinchilla, and PaLM on LaTeX knowledge, mathematical reasoning, and medical QA benchmarks while outperforming general models on BIG-bench.
-
Though Language Models Err While They Strive: Conformal Prediction for Self-Correcting Scientific Generation
A graph-based conformal wrapper that filters and regenerates LLM reasoning steps claims formal coverage guarantees on scientific validity, but its evaluation is circular and its gains are confounded with sampling effo...
-
Exploring the relationship between team institutional composition and novelty in academic papers based on fine-grained knowledge entities
Mixed academic-industrial teams in NLP produce more novel papers than purely industrial teams, with mixed teams emphasizing method-metric novelty and industrial teams emphasizing method-tool novelty.
-
Extracting Breast Cancer Phenotypes from Clinical Notes: Comparing LLMs with Classical Ontology Methods
LLM framework extracts breast cancer phenotypes from clinical notes with accuracy comparable to ontology-based methods and greater adaptability to new diseases.
-
Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts
An ensemble of three Google LLMs achieves a 0.74 weighted F1-score for detecting EQ-5D reporting in 200 PubMed abstracts, marginally outperforming individual models.
-
Data-Centric Foundation Models in Computational Healthcare: A Survey
The paper surveys data-centric strategies for foundation models in computational healthcare and supplies a curated list of related models and datasets.
-
Natural Language Processing in the Legal Domain
A survey of nearly 1000 NLP & Law papers from 2013-2024 documenting increases in publication volume, scope, methodological sophistication, and data/code availability.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.