A full-attention teacher distills knowledge to token-merging students, letting a recommender use 20K-token behavior sequences at near-baseline serving cost.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Teacher Retains Full Tokens, Student Merges Efficiently: TM20K for E-Commerce Sequence Modeling in Ad Recommendation
A full-attention teacher distills knowledge to token-merging students, letting a recommender use 20K-token behavior sequences at near-baseline serving cost.