REVIEW 5 cited by
MC-BERT: Efficient Language Pre-Training via a Meta Controller
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
MC-BERT: Efficient Language Pre-Training via a Meta Controller
read the original abstract
Pre-trained contextual representations (e.g., BERT) have become the foundation to achieve state-of-the-art results on many NLP tasks. However, large-scale pre-training is computationally expensive. ELECTRA, an early attempt to accelerate pre-training, trains a discriminative model that predicts whether each input token was replaced by a generator. Our studies reveal that ELECTRA's success is mainly due to its reduced complexity of the pre-training task: the binary classification (replaced token detection) is more efficient to learn than the generation task (masked language modeling). However, such a simplified task is less semantically informative. To achieve better efficiency and effectiveness, we propose a novel meta-learning framework, MC-BERT. The pre-training task is a multi-choice cloze test with a reject option, where a meta controller network provides training input and candidates. Results over GLUE natural language understanding benchmark demonstrate that our proposed method is both efficient and effective: it outperforms baselines on GLUE semantic tasks given the same computational budget.
Forward citations
Cited by 5 Pith papers
-
MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
MedStruct-S benchmark shows encoder-only models outperform larger decoder-only ones on key-conditioned QA from noisy OCR clinical reports, with fine-tuned large models winning only when scale is ignored.
-
Atom-level Protein Representation Learning Improves Protein Structure Prediction
TriProRep pretrains structure-aware protein representations from three aligned views via token recovery and improves results on the new RepSP benchmark for structure-predictive tasks.
-
Atom-level Protein Representation Learning Improves Protein Structure Prediction
TriProRep learns structure-aware protein representations via three-view VQ-VAE pretraining and shows gains on the new RepSP benchmark for dimer co-folding, interaction properties, and monomer structure prediction.
-
Atom-level Protein Representation Learning Improves Protein Structure Prediction
TriProRep pretrains on three aligned protein views and improves results on homodimer co-folding and related structure tasks in the new RepSP benchmark.
-
Enriched text-guided variational multimodal knowledge distillation network (VMD) for automated diagnosis of plaque vulnerability in 3D carotid artery MRI
VMD, a student-teacher-expert distillation network, improves plaque vulnerability classification in unannotated 3D carotid MRI by transferring knowledge from limited annotations and radiology reports.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.