Appending the candidate with cross-attention matches concatenation-based user history fusion and enables amortized multi-candidate inference that cuts latency by about 30% in production.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Efficient user history modeling with amortized inference for deep learning recommendation models
Appending the candidate with cross-attention matches concatenation-based user history fusion and enables amortized multi-candidate inference that cuts latency by about 30% in production.