Pith. sign in

REVIEW 3 cited by

The Cost of Training NLP Models: A Concise Overview

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.08900 v1 pith:LGFMCGGK submitted 2020-04-19 cs.CL cs.LGcs.NE

classification cs.CLcs.LGcs.NE
keywords costlanguagemodelstrainingaudiencebudgetingconcisecosts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We review the cost of training large-scale language models, and the drivers of these costs. The intended audience includes engineers and scientists budgeting their model-training experiments, as well as non-practitioners trying to make sense of the economics of modern-day Natural Language Processing (NLP).

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 115 citations worldwide. Full citation record

  1. GNN-CNN: An Efficient Hybrid Model of Convolutional and Graph Neural Networks for Text Representation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A hybrid GNN-CNN text classifier with real-time graph generation and injected LLM embeddings achieves near-transformer accuracy at linear complexity.

  2. SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A new benchmark of 15 small language models across 23 datasets and 11 metrics shows clear accuracy-versus-energy trade-offs, with no single model dominating.

  3. Empirical Evaluation of Large Language Models in Automated Program Repair

    cs.SE 2025-06 conditional novelty 5.0 of 10

    An empirical study of four open-source LLMs across six benchmarks shows code-specialized models often outperform larger general models, and most correct repairs appear early in generation.

Pith tools