Pith. sign in

REVIEW 1 cited by

DEPT: Decoupled Embeddings for Pre-training Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.05021 v5 pith:LBZSB3XB submitted 2024-10-07 cs.LG cs.CL

classification cs.LGcs.CL
keywords datadeptpre-trainingcommunicationtrainingbodycostsembedding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language Model pre-training uses broad data mixtures to enhance performance across domains and languages. However, training on such heterogeneous text corpora requires extensive and expensive efforts. Since these data sources vary significantly in lexical, syntactic, and semantic aspects, they cause negative interference or the ``curse of multilinguality''. To address these challenges we propose a communication-efficient pre-training framework, DEPT. Our method decouples embeddings from the transformer body while simultaneously training the latter on multiple data sources without requiring a shared vocabulary. DEPT can: (1) train robustly and effectively under significant data heterogeneity, (2) minimize token embedding parameters to only what the data source vocabulary requires, while cutting communication costs in direct proportion to both the communication frequency and the reduction in parameters, (3) enhance transformer body plasticity and generalization, improving both average perplexity (up to 20%) and downstream task performance, and (4) enable training with custom optimized vocabularies per data source. We demonstrate DEPT's potential via the first vocabulary-agnostic federated pre-training of billion-scale models, reducing communication costs by orders of magnitude and embedding memory by 4-5x.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Structure and Destructure: Dual Forces in the Making of Knowledge Engines

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Language modeling objectives induce recoverable structure in both graph and text models, and periodically resetting embeddings improves their plasticity; together these two forces unify structured and unstructured kno...

Pith tools