Pith. sign in

REVIEW 2 cited by

Data Scaling Laws in NMT: The Effect of Noise and Architecture

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.01994 v1 pith:NPZSSHYU submitted 2022-02-04 cs.LG cs.CL

classification cs.LGcs.CL
keywords datascalingtrainingarchitecturenoiseaddingeffectfind
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we study the effect of varying the architecture and training data quality on the data scaling properties of Neural Machine Translation (NMT). First, we establish that the test loss of encoder-decoder transformer models scales as a power law in the number of training samples, with a dependence on the model size. Then, we systematically vary aspects of the training setup to understand how they impact the data scaling laws. In particular, we change the following (1) Architecture and task setup: We compare to a transformer-LSTM hybrid, and a decoder-only transformer with a language modeling loss (2) Noise level in the training distribution: We experiment with filtering, and adding iid synthetic noise. In all the above cases, we find that the data scaling exponents are minimally impacted, suggesting that marginally worse architectures or training data can be compensated for by adding more data. Lastly, we find that using back-translated data instead of parallel data, can significantly degrade the scaling exponent.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Pre-training to One Hundred Billion Data for Vision Language Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Scaling VLM pretraining from 10B to 100B image-text pairs yields saturation on standard benchmarks but large gains on cultural diversity, low-resource language retrieval, and subgroup disparity.

  2. Recursive Inference Scaling: A Winning Path to Scalable Inference in Language and Multimodal Systems

    cs.AI 2025-02 conditional novelty 6.0 of 10

    Recursively applying the first half of a transformer before the second half (RINS) improves language modeling and vision-language accuracy under compute-matched comparisons.

Pith tools