Torque-Aware Momentum damps momentum updates by the alignment between new gradients and previous momentum, giving small gains on some benchmarks but mixed results on large model fine-tuning.
Bert: Pre-training of deep bidirectional transformers for language understanding
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Torque-Aware Momentum
Torque-Aware Momentum damps momentum updates by the alignment between new gradients and previous momentum, giving small gains on some benchmarks but mixed results on large model fine-tuning.