Pith. sign in

REVIEW 3 cited by

Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.01853 v1 pith:UFD47FVF submitted 2022-12-04 cs.CL

classification cs.CL
keywords languagetaskssupergluedownstreamknowledgemodelpretraininggoal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This technical report briefly describes our JDExplore d-team's Vega v2 submission on the SuperGLUE leaderboard. SuperGLUE is more challenging than the widely used general language understanding evaluation (GLUE) benchmark, containing eight difficult language understanding tasks, including question answering, natural language inference, word sense disambiguation, coreference resolution, and reasoning. [Method] Instead of arbitrarily increasing the size of a pretrained language model (PLM), our aim is to 1) fully extract knowledge from the input pretraining data given a certain parameter budget, e.g., 6B, and 2) effectively transfer this knowledge to downstream tasks. To achieve goal 1), we propose self-evolution learning for PLMs to wisely predict the informative tokens that should be masked, and supervise the masked language modeling (MLM) process with rectified smooth labels. For goal 2), we leverage the prompt transfer technique to improve the low-resource tasks by transferring the knowledge from the foundation model and related downstream tasks to the target task. [Results] According to our submission record (Oct. 2022), with our optimized pretraining and fine-tuning strategies, our 6B Vega method achieved new state-of-the-art performance on 4/8 tasks, sitting atop the SuperGLUE leaderboard on Oct. 8, 2022, with an average score of 91.3.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Open Your Model's Eyes: Video and Context-Aware Multimodal Backchannel Prediction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Adding video via a two-stage cross-modal alignment (CAMA-BC) raises macro-F1 for backchannel prediction to 58.53 on KC-Dialog and 39.20 on BACKSpeech, outperforming audio/text baselines.

  2. MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism

    cs.DC 2025-06 conditional novelty 6.0 of 10

    MPipeMoE speeds up MoE training by adaptively pipelining token batches and reusing memory buffers across partitions, achieving up to 2.8x speedup and 47% memory reduction over FasterMoE.

  3. Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A survey of NLU diagnostics benchmarks finds no shared naming convention or standard set of linguistic phenomena, and asks whether the field should build an ISO-like evaluation standard.

Pith tools