REVIEW 10 cited by
Structured information extraction from complex scientific text with fine-tuned large language models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Intelligently extracting and linking complex scientific information from unstructured text is a challenging endeavor particularly for those inexperienced with natural language processing. Here, we present a simple sequence-to-sequence approach to joint named entity recognition and relation extraction for complex hierarchical information in scientific text. The approach leverages a pre-trained large language model (LLM), GPT-3, that is fine-tuned on approximately 500 pairs of prompts (inputs) and completions (outputs). Information is extracted either from single sentences or across sentences in abstracts/passages, and the output can be returned as simple English sentences or a more structured format, such as a list of JSON objects. We demonstrate that LLMs trained in this way are capable of accurately extracting useful records of complex scientific knowledge for three representative tasks in materials chemistry: linking dopants with their host materials, cataloging metal-organic frameworks, and general chemistry/phase/morphology/application information extraction. This approach represents a simple, accessible, and highly-flexible route to obtaining large databases of structured knowledge extracted from unstructured text. An online demo is available at http://www.matscholar.com/info-extraction.
Forward citations
Cited by 10 Pith papers
-
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
ScheMatiQ uses LLMs to automatically generate schemas and extract structured data from text corpora based on natural language questions, supported by interactive user steering.
-
Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
TCoT uses a single VLM to select question-relevant video frames from segments, then answers from that curated context, improving video QA accuracy across four benchmarks and three VLMs.
-
MIMDE: Exploring the Use of Synthetic vs Human Data for Evaluating Multi-Insight Multi-Document Extraction Tasks
LLMs rank similarly on human and synthetic data for extracting insights, but synthetic data does not predict how well models map insights back to source documents.
-
A Hybrid Framework for Subject Analysis: Integrating Embedding-Based Regression Models with Large Language Models
Using an ML-predicted label count to constrain LLM generation and post-editing outputs to the LCSH vocabulary lifts subject-heading prediction F1 from 0.135 to 0.300 on a 2,100-book test set.
-
Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case
Reformulating generative product IE as LLM correction of PLM drafts improves F1 on weakly expressed attributes and lets mid-size local models approach larger ones.
-
Advancing Scientific Text Classification: Fine-Tuned Models with Dataset Expansion and Hard-Voting
Fine-tuning BERT-family models on a Web of Science dataset expanded with model-generated pseudo-labels and hard voting yields higher self-reported F1 scores, but the evaluation lacks independent ground truth.
-
A Layered Architecture for Developing and Enhancing Capabilities in Large Language Model-based Software Systems
A layered architecture with model, inference, and application layers, plus a capability-mapping process, guides where to implement features like structured output and domain knowledge in LLM systems.
-
ByteScience: Bridging Unstructured Scientific Literature and Structured Data with Auto Fine-tuned Large Language Model in Token Granularity
A cloud platform fine-tunes the DARWIN language model on a few annotated papers to extract structured data from scientific literature, but the reported accuracy is not accompanied by a reproducible evaluation.
-
Harnessing multiple LLMs for Information Retrieval: A case study on Deep Learning methodologies in Biodiversity publications
An ensemble of five RAG-assisted LLMs identifies the presence of deep-learning methodology details in biodiversity papers, agreeing with human annotations on 417 of 600 comparisons.
-
Fine-Tuning Large Language Models for Scientific Text Classification: A Comparative Study
A benchmark comparison of fine-tuned BERT, SciBERT, BioBERT, and BlueBERT on Web of Science scientific text classification, claiming SciBERT performs best.
Discussion (0). Continue with ORCID to comment.