Pith. sign in

REVIEW 10 cited by

Structured information extraction from complex scientific text with fine-tuned large language models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.05238 v1 pith:K6O24LVF submitted 2022-12-10 cs.CL cond-mat.mtrl-sci

classification cs.CLcond-mat.mtrl-sci
keywords informationcomplexscientifictextapproachextractionlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Intelligently extracting and linking complex scientific information from unstructured text is a challenging endeavor particularly for those inexperienced with natural language processing. Here, we present a simple sequence-to-sequence approach to joint named entity recognition and relation extraction for complex hierarchical information in scientific text. The approach leverages a pre-trained large language model (LLM), GPT-3, that is fine-tuned on approximately 500 pairs of prompts (inputs) and completions (outputs). Information is extracted either from single sentences or across sentences in abstracts/passages, and the output can be returned as simple English sentences or a more structured format, such as a list of JSON objects. We demonstrate that LLMs trained in this way are capable of accurately extracting useful records of complex scientific knowledge for three representative tasks in materials chemistry: linking dopants with their host materials, cataloging metal-organic frameworks, and general chemistry/phase/morphology/application information extraction. This approach represents a simple, accessible, and highly-flexible route to obtaining large databases of structured knowledge extracted from unstructured text. An online demo is available at http://www.matscholar.com/info-extraction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    ScheMatiQ uses LLMs to automatically generate schemas and extract structured data from text corpora based on natural language questions, supported by interactive user steering.

  2. Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

    cs.LG 2025-07 conditional novelty 6.0 of 10

    TCoT uses a single VLM to select question-relevant video frames from segments, then answers from that curated context, improving video QA accuracy across four benchmarks and three VLMs.

  3. MIMDE: Exploring the Use of Synthetic vs Human Data for Evaluating Multi-Insight Multi-Document Extraction Tasks

    cs.CL 2024-11 conditional novelty 6.0 of 10

    LLMs rank similarly on human and synthetic data for extracting insights, but synthetic data does not predict how well models map insights back to source documents.

  4. A Hybrid Framework for Subject Analysis: Integrating Embedding-Based Regression Models with Large Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Using an ML-predicted label count to constrain LLM generation and post-editing outputs to the LCSH vocabulary lifts subject-heading prediction F1 from 0.135 to 0.300 on a 2,100-book test set.

  5. Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Reformulating generative product IE as LLM correction of PLM drafts improves F1 on weakly expressed attributes and lets mid-size local models approach larger ones.

  6. Advancing Scientific Text Classification: Fine-Tuned Models with Dataset Expansion and Hard-Voting

    cs.CL 2025-04 reject novelty 4.0 of 10

    Fine-tuning BERT-family models on a Web of Science dataset expanded with model-generated pseudo-labels and hard voting yields higher self-reported F1 scores, but the evaluation lacks independent ground truth.

  7. A Layered Architecture for Developing and Enhancing Capabilities in Large Language Model-based Software Systems

    cs.SE 2024-11 conditional novelty 4.0 of 10

    A layered architecture with model, inference, and application layers, plus a capability-mapping process, guides where to implement features like structured output and domain knowledge in LLM systems.

  8. ByteScience: Bridging Unstructured Scientific Literature and Structured Data with Auto Fine-tuned Large Language Model in Token Granularity

    cs.CL 2024-11 reject novelty 4.0 of 10

    A cloud platform fine-tunes the DARWIN language model on a few annotated papers to extract structured data from scientific literature, but the reported accuracy is not accompanied by a reproducible evaluation.

  9. Harnessing multiple LLMs for Information Retrieval: A case study on Deep Learning methodologies in Biodiversity publications

    cs.IR 2024-11 conditional novelty 4.0 of 10

    An ensemble of five RAG-assisted LLMs identifies the presence of deep-learning methodology details in biodiversity papers, agreeing with human annotations on 417 of 600 comparisons.

  10. Fine-Tuning Large Language Models for Scientific Text Classification: A Comparative Study

    cs.CL 2024-11 conditional novelty 2.0 of 10

    A benchmark comparison of fine-tuned BERT, SciBERT, BioBERT, and BlueBERT on Web of Science scientific text classification, claiming SciBERT performs best.

Pith tools