Pith. sign in

REVIEW 2 cited by

DocRED: A Large-Scale Document-Level Relation Extraction Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.06127 v3 pith:77K36YYR submitted 2019-06-14 cs.CL

classification cs.CL
keywords docreddocument-levelmethodsrelationsdatasetdocumententitiesmultiple
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multiple entities in a document generally exhibit complex inter-sentence relations, and cannot be well handled by existing relation extraction (RE) methods that typically focus on extracting intra-sentence relations for single entity pairs. In order to accelerate the research on document-level RE, we introduce DocRED, a new dataset constructed from Wikipedia and Wikidata with three features: (1) DocRED annotates both named entities and relations, and is the largest human-annotated dataset for document-level RE from plain text; (2) DocRED requires reading multiple sentences in a document to extract entities and infer their relations by synthesizing all information of the document; (3) along with the human-annotated data, we also offer large-scale distantly supervised data, which enables DocRED to be adopted for both supervised and weakly supervised scenarios. In order to verify the challenges of document-level RE, we implement recent state-of-the-art methods for RE and conduct a thorough evaluation of these methods on DocRED. Empirical results show that DocRED is challenging for existing RE methods, which indicates that document-level RE remains an open problem and requires further efforts. Based on the detailed analysis on the experiments, we discuss multiple promising directions for future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SSCard: Substring Cardinality Estimation using Suffix Tree-Guided Learned FM-Index

    cs.DB 2025-05 conditional novelty 7.0 of 10

    SSCard estimates substring cardinality using a pruned suffix tree over an FM-index, with learned spline rank functions and an error bound, achieving lower q-error and smaller space than prior methods on five datasets.

  2. Multi-Relation Extraction in Entity Pairs using Global Context

    cs.CL 2025-07 reject novelty 3.0 of 10

    A BERT input format that appends the head and tail entity names after the document is claimed to beat all prior document-level relation extraction systems, but the reported gains rest on misaligned evaluation protocols.

Pith tools