Pith. sign in

REVIEW 2 cited by

Semi-automatic Data Enhancement for Document-Level Relation Extraction with Distant Supervision from Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.07314 v1 pith:P23IVM2W submitted 2023-11-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords relationlanguagedocument-levelextractionlargemethodcomprehensiondocre
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Document-level Relation Extraction (DocRE), which aims to extract relations from a long context, is a critical challenge in achieving fine-grained structural comprehension and generating interpretable document representations. Inspired by recent advances in in-context learning capabilities emergent from large language models (LLMs), such as ChatGPT, we aim to design an automated annotation method for DocRE with minimum human effort. Unfortunately, vanilla in-context learning is infeasible for document-level relation extraction due to the plenty of predefined fine-grained relation types and the uncontrolled generations of LLMs. To tackle this issue, we propose a method integrating a large language model (LLM) and a natural language inference (NLI) module to generate relation triples, thereby augmenting document-level relation datasets. We demonstrate the effectiveness of our approach by introducing an enhanced dataset known as DocGNRE, which excels in re-annotating numerous long-tail relation types. We are confident that our method holds the potential for broader applications in domain-specific relation type definitions and offers tangible benefits in advancing generalized language semantic comprehension.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Attack Effect Model based Malicious Behavior Detection

    cs.CR 2025-06 conditional novelty 5.0 of 10

    FEAD derives security monitoring items from attack reports with an LLM, decomposes them across existing collectors, and applies locality-aware graph analysis, reporting an 8.23% higher F1-score than ThreaTrace with 5....

  2. Large Language Model for Extracting Complex Contract Information in Industrial Scenes

    cs.CL 2025-07 conditional novelty 3.0 of 10

    Clustering contracts, LLM-based labeling, augmentation, and LoRA fine-tuning improve Chinese industrial contract field extraction over traditional TF-IDF/TextRank/SNOWNLP/KeyBERT baselines.

Pith tools