Pith. sign in

REVIEW 4 cited by

Structured information extraction from complex scientific text with fine-tuned large language models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.05238 v1 pith:K6O24LVF submitted 2022-12-10 cs.CL cond-mat.mtrl-sci

classification cs.CLcond-mat.mtrl-sci
keywords informationcomplexscientifictextapproachextractionlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Intelligently extracting and linking complex scientific information from unstructured text is a challenging endeavor particularly for those inexperienced with natural language processing. Here, we present a simple sequence-to-sequence approach to joint named entity recognition and relation extraction for complex hierarchical information in scientific text. The approach leverages a pre-trained large language model (LLM), GPT-3, that is fine-tuned on approximately 500 pairs of prompts (inputs) and completions (outputs). Information is extracted either from single sentences or across sentences in abstracts/passages, and the output can be returned as simple English sentences or a more structured format, such as a list of JSON objects. We demonstrate that LLMs trained in this way are capable of accurately extracting useful records of complex scientific knowledge for three representative tasks in materials chemistry: linking dopants with their host materials, cataloging metal-organic frameworks, and general chemistry/phase/morphology/application information extraction. This approach represents a simple, accessible, and highly-flexible route to obtaining large databases of structured knowledge extracted from unstructured text. An online demo is available at http://www.matscholar.com/info-extraction.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    ScheMatiQ uses LLMs to automatically generate schemas and extract structured data from text corpora based on natural language questions, supported by interactive user steering.

  2. Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

    cs.LG 2025-07 conditional novelty 6.0 of 10

    TCoT uses a single VLM to select question-relevant video frames from segments, then answers from that curated context, improving video QA accuracy across four benchmarks and three VLMs.

  3. A Hybrid Framework for Subject Analysis: Integrating Embedding-Based Regression Models with Large Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Using an ML-predicted label count to constrain LLM generation and post-editing outputs to the LCSH vocabulary lifts subject-heading prediction F1 from 0.135 to 0.300 on a 2,100-book test set.

  4. Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Reformulating generative product IE as LLM correction of PLM drafts improves F1 on weakly expressed attributes and lets mid-size local models approach larger ones.

Pith tools