REVIEW 3 cited by
ERNIE 2.0: A Continual Pre-training Framework for Language Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recently, pre-trained models have achieved state-of-the-art results in various language understanding tasks, which indicates that pre-training on large-scale corpora may play a crucial role in natural language processing. Current pre-training procedures usually focus on training the model with several simple tasks to grasp the co-occurrence of words or sentences. However, besides co-occurring, there exists other valuable lexical, syntactic and semantic information in training corpora, such as named entity, semantic closeness and discourse relations. In order to extract to the fullest extent, the lexical, syntactic and semantic information from training corpora, we propose a continual pre-training framework named ERNIE 2.0 which builds and learns incrementally pre-training tasks through constant multi-task learning. Experimental results demonstrate that ERNIE 2.0 outperforms BERT and XLNet on 16 tasks including English tasks on GLUE benchmarks and several common tasks in Chinese. The source codes and pre-trained models have been released at https://github.com/PaddlePaddle/ERNIE.
Forward citations
Cited by 3 Pith papers
-
Specializing Unsupervised Pretraining Models for Word-Level Semantic Similarity
LIBERT, a BERT variant pretrained with an auxiliary word-pair similarity task, outperforms BERT on 9/10 GLUE tasks and on three lexical simplification datasets.
-
NEZHA: Neural Contextualized Representation for Chinese Language Understanding
NEZHA is a BERT-style Chinese pretrained model whose parameter-free functional relative positional encoding and training improvements push scores higher on several Chinese NLU benchmarks.
-
A Morpho-Syntactically Informed LSTM-CRF Model for Named Entity Recognition
An LSTM-CRF named entity recognizer for Bulgarian reaches F1 92.20 by adding part-of-speech tags and, to a lesser extent, morphological features to word and character embeddings.
Discussion (0). Continue with ORCID to comment.