A comparison of NLP, computer vision, and multimodal methods for PDF metadata extraction, including a new TextMap approach, evaluated on two newly built datasets.
Towards Better UD Parsing: Deep Contextualized Word Embeddings, Ensemble, and Treebank Concatenation
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This paper describes our system (HIT-SCIR) submitted to the CoNLL 2018 shared task on Multilingual Parsing from Raw Text to Universal Dependencies. We base our submission on Stanford's winning system for the CoNLL 2017 shared task and make two effective extensions: 1) incorporating deep contextualized word embeddings into both the part of speech tagger and parser; 2) ensembling parsers trained with different initialization. We also explore different ways of concatenating treebanks for further improvements. Experimental results on the development data show the effectiveness of our methods. In the final evaluation, our system was ranked first according to LAS (75.84%) and outperformed the other systems by a large margin.
fields
cs.IR 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Comparison of Feature Learning Methods for Metadata Extraction from PDF Scholarly Documents
A comparison of NLP, computer vision, and multimodal methods for PDF metadata extraction, including a new TextMap approach, evaluated on two newly built datasets.