With per-method hyperparameter tuning, 50x top-k and DGC sparsification improved PTB LSTM perplexity by up to 0.06 over the uncompressed baseline, while QSGD and stronger compression performed at or below baseline.
Neural machine translation by jointly learning to align and translate,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Efficient Distributed Training through Gradient Compression with Sparsification and Quantization Techniques
With per-method hyperparameter tuning, 50x top-k and DGC sparsification improved PTB LSTM perplexity by up to 0.06 over the uncompressed baseline, while QSGD and stronger compression performed at or below baseline.