Pith. sign in

REVIEW 1 cited by

Training a T5 Using Lab-sized Resources

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.12097 v1 pith:3TQWANOP submitted 2022-08-25 cs.CL

classification cs.CL
keywords languagelargeresourcesmodelmodelstraintrainingamount
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training large neural language models on large datasets is resource- and time-intensive. These requirements create a barrier to entry, where those with fewer resources cannot build competitive models. This paper presents various techniques for making it possible to (a) train a large language model using resources that a modest research lab might have, and (b) train it in a reasonable amount of time. We provide concrete recommendations for practitioners, which we illustrate with a case study: a T5 model for Danish, the first for this language.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SnakModel: Lessons Learned from Training an Open Danish Large Language Model

    cs.CL 2024-12 conditional novelty 5.0 of 10

    SnakModel, an open Danish 7B LLM, outperforms other Llama2-7B-based models on the ScandEval Danish benchmark, with analyses of training dynamics and data curation.

Pith tools