Pith. sign in

REVIEW 3 cited by

Curriculum learning for language modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.02170 v1 pith:E32HR2XL submitted 2021-08-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords languagecurriculumlearningmodeltrainingmodelsimprovenatural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language Models like ELMo and BERT have provided robust representations of natural language, which serve as the language understanding component for a diverse range of downstream tasks.Curriculum learning is a method that employs a structured training regime instead, which has been leveraged in computer vision and machine translation to improve model training speed and model performance. While language models have proven transformational for the natural language processing community, these models have proven expensive, energy-intensive, and challenging to train. In this work, we explore the effect of curriculum learning on language model pretraining using various linguistically motivated curricula and evaluate transfer performance on the GLUE Benchmark. Despite a broad variety of training methodologies and experiments we do not find compelling evidence that curriculum learning methods improve language model training.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Estimating the Effects of Sample Training Orders for Large Language Models without Retraining

    cs.LG 2025-05 reject novelty 6.0 of 10

    A framework using Taylor expansions and random projections estimates LLM performance under arbitrary training batch orders from one reference run.

  2. Data Efficacy for Language Model Training

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Ordering training data by a gradient-based score, using a folding scheme that interleaves multiple curriculum passes, improves small-scale LM accuracy by roughly 1.5 to 2 points on average benchmarks.

  3. Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey proposes a macro-meso-micro value framework for agentic AI alignment and maps applications, methods, and benchmarks onto it.

Pith tools