Pith. sign in

REVIEW 5 cited by

MC-BERT: Efficient Language Pre-Training via a Meta Controller

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.05744 v2 pith:RWWQ6VDL submitted 2020-06-10 cs.CL cs.LG

MC-BERT: Efficient Language Pre-Training via a Meta Controller

classification cs.CL cs.LG
keywords pre-trainingtaskefficientlanguageachievecontrollerelectraglue
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Pre-trained contextual representations (e.g., BERT) have become the foundation to achieve state-of-the-art results on many NLP tasks. However, large-scale pre-training is computationally expensive. ELECTRA, an early attempt to accelerate pre-training, trains a discriminative model that predicts whether each input token was replaced by a generator. Our studies reveal that ELECTRA's success is mainly due to its reduced complexity of the pre-training task: the binary classification (replaced token detection) is more efficient to learn than the generation task (masked language modeling). However, such a simplified task is less semantically informative. To achieve better efficiency and effectiveness, we propose a novel meta-learning framework, MC-BERT. The pre-training task is a multi-choice cloze test with a reject option, where a meta controller network provides training input and candidates. Results over GLUE natural language understanding benchmark demonstrate that our proposed method is both efficient and effective: it outperforms baselines on GLUE semantic tasks given the same computational budget.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports

    cs.CL 2026-05 unverdicted novelty 7.0

    MedStruct-S benchmark shows encoder-only models outperform larger decoder-only ones on key-conditioned QA from noisy OCR clinical reports, with fine-tuned large models winning only when scale is ignored.

  2. Atom-level Protein Representation Learning Improves Protein Structure Prediction

    q-bio.BM 2026-05 unverdicted novelty 6.0

    TriProRep pretrains structure-aware protein representations from three aligned views via token recovery and improves results on the new RepSP benchmark for structure-predictive tasks.

  3. Atom-level Protein Representation Learning Improves Protein Structure Prediction

    q-bio.BM 2026-05 unverdicted novelty 6.0

    TriProRep learns structure-aware protein representations via three-view VQ-VAE pretraining and shows gains on the new RepSP benchmark for dimer co-folding, interaction properties, and monomer structure prediction.

  4. Atom-level Protein Representation Learning Improves Protein Structure Prediction

    q-bio.BM 2026-05 unverdicted novelty 5.0

    TriProRep pretrains on three aligned protein views and improves results on homodimer co-folding and related structure tasks in the new RepSP benchmark.

  5. Enriched text-guided variational multimodal knowledge distillation network (VMD) for automated diagnosis of plaque vulnerability in 3D carotid artery MRI

    cs.CV 2025-09 conditional novelty 4.0

    VMD, a student-teacher-expert distillation network, improves plaque vulnerability classification in unannotated 3D carotid MRI by transferring knowledge from limited annotations and radiology reports.