REVIEW 4 cited by
Massive Exploration of Neural Machine Translation Architectures
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Neural Machine Translation (NMT) has shown remarkable progress over the past few years with production systems now being deployed to end-users. One major drawback of current architectures is that they are expensive to train, typically requiring days to weeks of GPU time to converge. This makes exhaustive hyperparameter search, as is commonly done with other neural network architectures, prohibitively expensive. In this work, we present the first large-scale analysis of NMT architecture hyperparameters. We report empirical results and variance numbers for several hundred experimental runs, corresponding to over 250,000 GPU hours on the standard WMT English to German translation task. Our experiments lead to novel insights and practical advice for building and extending NMT architectures. As part of this contribution, we release an open-source NMT framework that enables researchers to easily experiment with novel techniques and reproduce state of the art results.
Forward citations
Cited by 4 Pith papers
-
Anomaly Detection for IoT Global Connectivity
An unsupervised roaming-signaling pipeline flags IoT fleet connectivity issues, but the headline results are weakened by training/test overlap.
-
AsserT5: Test Assertion Generation Using a Fine-Tuned Code Language Model
A fine-tuned CodeT5 model generates exact-match test assertions in up to 59.5% of cases, but detects only 33 of 138 real Defects4J bugs.
-
A Survey on Large Language Models with some Insights on their Capabilities and Limitations
A broad survey of LLM methods and applications, plus an empirical section on how code-rich pretraining may influence chain-of-thought reasoning, the details of which are not visible in the supplied text.
-
Sequence-to-Sequence Natural Language to Humanoid Robot Sign Language
Applies established seq2seq neural networks to convert text to Spanish sign language for humanoid robot TEO, proposing OpenPose for skeleton data collection to handle sequence length differences and non-manual markers.
Discussion (0). Continue with ORCID to comment.