Pith. sign in

REVIEW 1 cited by

From Dataset Recycling to Multi-Property Extraction and Beyond

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.03228 v1 pith:7R2XSXI6 submitted 2020-11-06 cs.CL cs.IR

classification cs.CLcs.IR
keywords datasetextractionwikireadingmodeladditionanalysisarchitecturesbeyond
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper investigates various Transformer architectures on the WikiReading Information Extraction and Machine Reading Comprehension dataset. The proposed dual-source model outperforms the current state-of-the-art by a large margin. Next, we introduce WikiReading Recycled-a newly developed public dataset and the task of multiple property extraction. It uses the same data as WikiReading but does not inherit its predecessor's identified disadvantages. In addition, we provide a human-annotated test set with diagnostic subsets for a detailed analysis of model performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Extract Information from Hybrid Long Documents Leveraging LLMs: A Framework and Dataset

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A retrieve-summarize-extract pipeline with simple table-to-text serialization improves LLM extraction from hybrid long documents, and a new financial KPI dataset (FINE) is introduced to support evaluation.

Pith tools