Pith. sign in

REVIEW 1 cited by

OPIEC: An Open Information Extraction Corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.12324 v1 pith:E3GVEDJX submitted 2019-04-28 cs.CL

classification cs.CL
keywords opieccorpusinformationknowledgeopenbasefactsvaluable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open information extraction (OIE) systems extract relations and their arguments from natural language text in an unsupervised manner. The resulting extractions are a valuable resource for downstream tasks such as knowledge base construction, open question answering, or event schema induction. In this paper, we release, describe, and analyze an OIE corpus called OPIEC, which was extracted from the text of English Wikipedia. OPIEC complements the available OIE resources: It is the largest OIE corpus publicly available to date (over 340M triples) and contains valuable metadata such as provenance information, confidence scores, linguistic annotations, and semantic annotations including spatial and temporal information. We analyze the OPIEC corpus by comparing its content with knowledge bases such as DBpedia or YAGO, which are also based on Wikipedia. We found that most of the facts between entities present in OPIEC cannot be found in DBpedia and/or YAGO, that OIE facts often differ in the level of specificity compared to knowledge base facts, and that OIE open relations are generally highly polysemous. We believe that the OPIEC corpus is a valuable resource for future research on automated knowledge base construction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new open-domain dialogue dataset shows that users' own judgments of stylistic similarity correlate with their preference (Spearman r=0.67-0.75), but third-party stylistic similarity judgments do not, indicating a ga...

Pith tools