REVIEW 3 cited by
Curriculum learning for language modeling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language Models like ELMo and BERT have provided robust representations of natural language, which serve as the language understanding component for a diverse range of downstream tasks.Curriculum learning is a method that employs a structured training regime instead, which has been leveraged in computer vision and machine translation to improve model training speed and model performance. While language models have proven transformational for the natural language processing community, these models have proven expensive, energy-intensive, and challenging to train. In this work, we explore the effect of curriculum learning on language model pretraining using various linguistically motivated curricula and evaluate transfer performance on the GLUE Benchmark. Despite a broad variety of training methodologies and experiments we do not find compelling evidence that curriculum learning methods improve language model training.
Forward citations
Cited by 3 Pith papers
-
Estimating the Effects of Sample Training Orders for Large Language Models without Retraining
A framework using Taylor expansions and random projections estimates LLM performance under arbitrary training batch orders from one reference run.
-
Data Efficacy for Language Model Training
Ordering training data by a gradient-based score, using a folding scheme that interleaves multiple curriculum passes, improves small-scale LM accuracy by roughly 1.5 to 2 points on average benchmarks.
-
Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives
A survey proposes a macro-meso-micro value framework for agentic AI alignment and maps applications, methods, and benchmarks onto it.
Discussion (0). Sign in to comment.