Pith. sign in

REVIEW 1 cited by

oBERTa: Improving Sparse Transfer Learning via improved initialization, distillation, and pruning regimes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.17612 v3 pith:SWXJ6HLW submitted 2023-03-30 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords obertamodelcompressiondistillationlanguagemodelspruningbroad
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we introduce the range of oBERTa language models, an easy-to-use set of language models which allows Natural Language Processing (NLP) practitioners to obtain between 3.8 and 24.3 times faster models without expertise in model compression. Specifically, oBERTa extends existing work on pruning, knowledge distillation, and quantization and leverages frozen embeddings improves distillation and model initialization to deliver higher accuracy on a broad range of transfer tasks. In generating oBERTa, we explore how the highly optimized RoBERTa differs from the BERT for pruning during pre-training and finetuning. We find it less amenable to compression during fine-tuning. We explore the use of oBERTa on seven representative NLP tasks and find that the improved compression techniques allow a pruned oBERTa model to match the performance of BERTbase and exceed the performance of Prune OFA Large on the SQUAD V1.1 Question Answering dataset, despite being 8x and 2x, respectively faster in inference. We release our code, training regimes, and associated model for broad usage to encourage usage and experimentation

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MeVe: A Modular System for Memory Verification and Effective Context Control in Language Models

    cs.CL 2025-09 conditional novelty 4.0 of 10

    MeVe, a five-stage modular RAG pipeline, reduces average context tokens by 57-75% versus plain top-k retrieval in a simulated proof-of-concept, without improving answer-quality metrics.

Pith tools