Pith. sign in

REVIEW 6 cited by

An Empirical Study on Information Extraction using Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14450 v2 pith:N6CLQI7R submitted 2023-05-23 cs.CL

classification cs.CL
keywords informationextractionllmsabilitygpt-4languagemethodshuman-like
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human-like large language models (LLMs), especially the most powerful and popular ones in OpenAI's GPT family, have proven to be very helpful for many natural language processing (NLP) related tasks. Therefore, various attempts have been made to apply LLMs to information extraction (IE), which is a fundamental NLP task that involves extracting information from unstructured plain text. To demonstrate the latest representative progress in LLMs' information extraction ability, we assess the information extraction ability of GPT-4 (the latest version of GPT at the time of writing this paper) from four perspectives: Performance, Evaluation Criteria, Robustness, and Error Types. Our results suggest a visible performance gap between GPT-4 and state-of-the-art (SOTA) IE methods. To alleviate this problem, considering the LLMs' human-like characteristics, we propose and analyze the effects of a series of simple prompt-based methods, which can be generalized to other LLMs and NLP tasks. Rich experiments show our methods' effectiveness and some of their remaining issues in improving GPT-4's information extraction ability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fine-Grained Chinese Hate Speech Understanding: Span-Level Resources, Coded Term Lexicon, and Enhanced Detection Frameworks

    cs.CL 2025-07 reject novelty 7.0 of 10

    The paper creates a span-level Chinese hate speech dataset and a 830-term coded hate lexicon, but its two-stage training method's reported superiority is contradicted by the paper's own COLD results.

  2. Discourse-Aware Policy Analysis with Argumentation: A Hybrid LLM-Symbolic Framework for Disaster Governance

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A hybrid LLM-symbolic pipeline maps disaster-policy text to typed argumentation graphs using new frame-mediated relation subtypes, with a new 100-document dataset from USA, UK, Canada, and Australia.

  3. From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A weak-supervision pipeline using LLM-extracted concepts and rank fusion creates silver-standard patent-to-SDG labels that recover known citation-derived associations and show high network modularity.

  4. MPL: Multiple Programming Languages with Large Language Models for Information Extraction

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Using multiple programming languages as code-style prompts during fine-tuning improves LLM information extraction accuracy over single-language prompting.

  5. GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A single 205M-parameter encoder model unifies named entity recognition, text classification, and hierarchical structured extraction through declarative schemas.

  6. MExplore: an entity-based visual analytics approach for medical expertise acquisition

    cs.HC 2025-07 conditional novelty 5.0 of 10

    An entity-based multi-level visual analytics system for acquiring medical expertise from unstructured medical text, evaluated in a small user study without significance testing.

Pith tools