A bidirectional, interleaved interaction module with a separate Cross Arch for selective summarization improves CTR prediction over unidirectional fusion baselines by small margins on public and industrial data.
Highway Transformer: Self-Gating Enhanced Self-Attentive Networks
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Self-attention mechanisms have made striking state-of-the-art (SOTA) progress in various sequence learning tasks, standing on the multi-headed dot product attention by attending to all the global contexts at different locations. Through a pseudo information highway, we introduce a gated component self-dependency units (SDU) that incorporates LSTM-styled gating units to replenish internal semantic importance within the multi-dimensional latent space of individual representations. The subsidiary content-based SDU gates allow for the information flow of modulated latent embeddings through skipped connections, leading to a clear margin of convergence speed with gradient descent algorithms. We may unveil the role of gating mechanism to aid in the context-based Transformer modules, with hypothesizing that SDU gates, especially on shallow layers, could push it faster to step towards suboptimal points during the optimization process.
fields
cs.IR 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
InterFormer: Effective Heterogeneous Interaction Learning for Click-Through Rate Prediction
A bidirectional, interleaved interaction module with a separate Cross Arch for selective summarization improves CTR prediction over unidirectional fusion baselines by small margins on public and industrial data.