Pith. sign in

REVIEW 4 cited by

Massive Exploration of Neural Machine Translation Architectures

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1703.03906 v2 pith:AMBPOFYW submitted 2017-03-11 cs.CL

classification cs.CL
keywords architecturesneuraltranslationexpensivemachinenovelresultsadvice
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural Machine Translation (NMT) has shown remarkable progress over the past few years with production systems now being deployed to end-users. One major drawback of current architectures is that they are expensive to train, typically requiring days to weeks of GPU time to converge. This makes exhaustive hyperparameter search, as is commonly done with other neural network architectures, prohibitively expensive. In this work, we present the first large-scale analysis of NMT architecture hyperparameters. We report empirical results and variance numbers for several hundred experimental runs, corresponding to over 250,000 GPU hours on the standard WMT English to German translation task. Our experiments lead to novel insights and practical advice for building and extending NMT architectures. As part of this contribution, we release an open-source NMT framework that enables researchers to easily experiment with novel techniques and reproduce state of the art results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Anomaly Detection for IoT Global Connectivity

    cs.NI 2025-08 conditional novelty 6.0 of 10

    An unsupervised roaming-signaling pipeline flags IoT fleet connectivity issues, but the headline results are weakened by training/test overlap.

  2. AsserT5: Test Assertion Generation Using a Fine-Tuned Code Language Model

    cs.SE 2025-02 conditional novelty 6.0 of 10

    A fine-tuned CodeT5 model generates exact-match test assertions in up to 59.5% of cases, but detects only 33 of 138 real Defects4J bugs.

  3. A Survey on Large Language Models with some Insights on their Capabilities and Limitations

    cs.CL 2025-01 unverdicted novelty 3.0 of 10

    A broad survey of LLM methods and applications, plus an empirical section on how code-rich pretraining may influence chain-of-thought reasoning, the details of which are not visible in the supplied text.

  4. Sequence-to-Sequence Natural Language to Humanoid Robot Sign Language

    cs.RO 2019-07 unverdicted novelty 3.0 of 10

    Applies established seq2seq neural networks to convert text to Spanish sign language for humanoid robot TEO, proposing OpenPose for skeleton data collection to handle sequence length differences and non-manual markers.

Pith tools