REVIEW 5 major objections 8 minor 39 references
Teacher Retains Full Tokens, Student Merges Efficiently: TM20K for E-Commerce Sequence Modeling in Ad Recommendation
T0 review · 5 major / 8 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read TM20K shows that extending ad sequences to 20K tokens is practical: a once-trained teacher supervises a token-merging student, lifting an online advertiser metric by 1.036% with only 5.6% higher serving latency.
desk verdict Well-engineered industrial sequence modeling worth refereeing, but the offline recovery ratio is statistically fragile. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanisms are (1) full attention (FA) as the sequence encoder, which the paper shows beats target-only attention by up to 0.25% AUC at 20K and uses sequence-to-sequence interactions that are nearly as informative as target-to-sequence ones; (2) three hierarchical token-merge rules derived from attention statistics: LITM merges consecutive same-product-ID tokens, PATM applies stronger compression to older positions, and LPTM halves the sequence length every two transformer layers; and (3) a two-stage distillation setup where a teacher trained once on full 20K tokens caches its logits, and the student optimizes a main prediction head plus a distillation head with loss $\lambda\ell_{\mathrm{kd}}$ matched in scale to the cross-entropy loss. Together they move the heavy computation into offline training while keeping the online model's input short.
What would settle it
A concrete check: compute the same attention statistics on the fully trained teacher using all query positions and all heads, then rebuild the student's merge masks from that full distribution. If the resulting student closes most of the remaining 15% gap to the teacher, the paper's rule-based merges are not preserving all the signal they claim; if the performance is comparable, the small-sample statistics were representative.
Extended reading notes
Core claim
The central discovery is that the conflict between effectiveness and efficiency for ultra-long sequences can be separated into a one-time training problem and an online serving problem. A teacher transformer trained with full attention on all 20K tokens (with causal masking) achieves the best prediction quality, improving CVR AUC by 0.26% over the 5K baseline. Three rule-based merge strategies—merging repeated product-ID interactions in a short window (LITM), compressing older tokens more aggressively than recent ones (PATM), and halving sequence length every few transformer layers (LPTM)—shrink the student's average sequence from 8.8K to 1.8K tokens. The student alone gains +0.15% AUC, and after distillation from the teacher's cached logits it rises to +0.22%, roughly 85% of the teacher's gain, at a training throughput of 83K compared with the baseline's 88K and a serving latency increase of only 5.6%.
Load-bearing premise
The three token-merge rules are based on attention statistics gathered from only five query tokens and the first attention head of a single model; if those statistics do not represent the full multi-head attention distribution, the merges could be discarding signal that the ablations would not reveal.
Editorial extensions
If this is right
- Other advertising and recommendation systems can extend behavior sequences to tens of thousands of tokens without expensive online serving changes, if they re-estimate the attention priors on their own data.
- A one-time teacher can be reused across several student deployments with different latency budgets, since the teacher's logits are cached and the merge hyperparameters are tuned per student.
- Full attention should be preferred over target-only attention for long sequences in this setting, contradicting the implicit assumption in some earlier sequence-compression work that target interactions capture most of the signal.
- The 85% recovery means the majority of the long-sequence benefit can be obtained without serving a full 20K transformer, leaving a small residual gap that future learned-merge methods might close.
Reading between the lines
- If the same attention statistics are re-measured with all heads and full query coverage, the merge rules might change; this would be a direct test of whether the paper's small-sample analysis is representative.
- The framework could transfer to non-advertisement domains such as feed or short-video recommendation, but the recency and same-ID priors may differ, so the Sec. 3.3 statistics would need to be recomputed there.
- Because the teacher is trained once and its logits are cached, one teacher could serve as the distillation source for many students with varying sequence budgets, turning the one-time cost into an even smaller amortized fraction of deployment.
- A learned token-merge policy that imitates the teacher's attention would likely close the remaining 15% performance gap, at the cost of the extra engineering the paper explicitly avoids.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes TM20K, a two-stage knowledge-distillation framework for ultra-long e-commerce behavior sequences (up to 20K raw tokens) in a deployed advertising recommender system. The teacher model is a full-attention transformer trained once on uncompressed 20K sequences; the student model applies three attention-motivated token merge strategies (LITM, PATM, LPTM) to compress sequences for efficient training and serving, and is distilled from the teacher's cached logits. Offline experiments on a proprietary industrial dataset compare TM20K with several long-sequence baselines, report throughput/memory numbers, and include ablations of each merge strategy and of distillation weights. An online A/B test shows ADSS +1.036% and ADVV +0.780% with serving latency +5.6% relative to a deployed 5K baseline. The paper's headline quantitative claim is that the distilled student recovers about 85% of the teacher's AUC gain, computed as +0.22%/+0.26%.
Significance. If the reported results hold, the paper provides a practically valuable industrial solution: it demonstrates that a one-time full-attention teacher can transfer long-sequence knowledge to a token-merging student, making 20K-scale sequence modeling feasible online with modest overhead. The three merge strategies are simple, interpretable, and directly tied to observed attention statistics, and the paper includes transparent efficiency measurements, a broad ablation study, and real deployment evidence. The main weakness is statistical: the central recovery ratio and most offline deltas are single runs with no confidence intervals or significance tests, and the teacher and student are trained under different optimization budgets, so the quantitative strength of the claims is currently overstated. The online A/B gains are encouraging but lack test details. Overall, this is a useful applied contribution with a defensible architecture, but the evidence base needs tightening before the central claims can be accepted at face value.
major comments (5)
- [Table 2 / Sec. 5.2] The paper's headline 'recovering around 85% of the teacher's total performance improvement' is computed as 0.22/0.26 from single-run ΔAUC values that differ by only 0.04 percentage points, yet no confidence intervals, standard errors, or repeated-seed results are reported for any offline metric in Tables 2-6. Given that several ablation deltas in Tables 3-6 are of the same 0.01-0.06% magnitude, run-to-run variance could easily move the ratio from 85% to near parity or above 100%. Please report multiple seeds or a significance test for the key comparisons, especially the 5K baseline, teacher, and student-with-KD rows of Table 2.
- [Sec. 5.1.3 / Sec. 5.2] The teacher is trained with batch size 96 while the student uses batch size 320, and the abstract describes the teacher as 'heavily trained,' but the paper does not specify training steps, epochs, or data repetition for either model. The teacher's +0.26% AUC gain is therefore not a clean measure of what is preserved by full tokens versus token merge; part of the gap may reflect a larger training budget. Please report the exact optimization budget for both models, or include a control where the teacher is trained with the student's batch size and protocol.
- [Sec. 5.4] The online A/B section states that 'TM20K yields statistically significant positive gains across all business metrics' but gives no test details, confidence intervals, or evaluation window. Since the ADSS +1.036% and ADVV +0.780% numbers carry the paper's practical claim, please specify the statistical test, number of days, sample sizes, and confidence intervals for the reported deltas.
- [Sec. 3.3] The attention analysis that motivates the three token merge strategies uses an unstated number of training instances, m=5 query tokens, and only the first attention head. The paper should state the instance count and, ideally, verify that the observed patterns (same-ID attention closeness, recency concentration, layer-wise entropy differences) hold across heads and layers; alternatively, temper the 'well-motivated' characterization, since Table 3 is the direct empirical test of the merge strategies.
- [Sec. 5.2 / Table 2] All deltas are relative to the deployed 5K baseline, which itself uses LITM0 and UTM2 compression and has an average sequence length of 1.2K. Table 3 shows that a full 5K sequence outperforms this compressed baseline by 0.05% AUC, so the teacher's +0.26% conflates the benefit of extending to 20K with the benefit of removing the 5K compression. To support the interpretation that TM20K extends the sequence length usefully, report deltas relative to a full 5K model or explicitly decompose the length and compression effects.
minor comments (8)
- [Eq. (10)] Equation (10) uses q_T both for the teacher's logit and for its sigmoid-transformed probability; please distinguish the logit (e.g., z_T) from the probability to avoid notational collision.
- [Table 5 caption] The caption 'AUC gains relative to the TM20K-S model with different distillation weights' is ambiguous; clarify whether the reference is the student without KD or the best lambda configuration.
- [Sec. 5.3.4] The row 'w/o KD in Late Period' does not define what 'Late Period' means; specify the training stage at which the distillation loss is dropped.
- [Fig. 3(b) / Sec. 4.2.2] State the position-ordering convention once (smaller index = more recent) and keep it consistent in the figure, the text, and the algorithm descriptions.
- [Abstract] The abstract's phrase 'nearly the same' training cost should acknowledge the one-time teacher training cost; the main text does this, but the abstract could be more precise.
- [Sec. 4.2.3 / Eq. (9)] Equation (9) merges adjacent token pairs via reshape to (L_n/2, 2, d), which requires L_n to be even; specify how odd lengths are handled, especially after LITM/PATM produce variable-length sequences.
- [Sec. 5.1.2] The exclusion of DIN and TWIN results is justified qualitatively; consider reporting a single performance number to substantiate the claim of 'obvious performance drops.'
- [Reference [14]] Reference [14] lists 'Ads Recommendation' as the author, which appears to be a venue or organization name rather than a person; please check the bibliographic entry.
Circularity Check
No significant circularity: the token-merge strategies and the ~85% recovery claim are measured results, not reductions to inputs or self-citation chains.
full rationale
The paper's central deliverables—the three token-merge strategies (Sec. 4.2) and the claim that TM20K-S w/ KD recovers ~85% of the teacher's 20K AUC gain—are empirical outputs, not consequences of an equation re-defining a fit as a prediction. The merge rules are motivated by attention-score statistics (Sec. 3.3), but their performance is measured directly in Table 3 against the uncompressed 20K full-sequence baseline, so the design rationale and the evaluation are independent. The 0.22/0.26 recovery ratio is computed from two measured ΔAUC values in Table 2; it is not a fitted parameter renamed as a prediction, despite being statistically fragile because no confidence intervals are reported and the teacher and student use different batch sizes. Same-group citations (LONGER [2], Rec-Distill [7], IAT [17], RankMixer [39]) appear as baselines, components, or prior KD frameworks; none is invoked as a premise that forces the result. The token-merge operation borrowed from [2] is fully specified in Algorithms 1–3 and benchmarked against the full-sequence baseline. No uniqueness theorem or self-citation chain is used to forbid alternatives. The paper is self-contained in that its claims are supported by its own experiments; the noted statistical caveats belong to correctness risk, not circularity.
Assumptions & free parameters
free parameters (5)
- LITM gap threshold T =
3
- PATM segment ranges and compression factors =
Example R=[[1,1000],[1001,5000],[5001,10000],[10001,20000]], K=[1,2,3,4]; deployed config described as 'moderately…
- KD weight lambda =
50
- LPTM merge frequency =
every 2 transformer layers
- Teacher batch size =
96
assumptions (5)
- domain assumption Full attention (FA) extracts sequential information better than target attention (TA) for ultra-long user behavior sequences.
- domain assumption Sum-pooling merged tokens preserves most of the predictive information when the merged tokens have similar attention behavior.
- domain assumption Knowledge distillation from cached teacher logits lets the student recover most of the teacher's performance.
- ad hoc to paper Attention statistics from a few training instances and the first attention head are representative of the full model.
- domain assumption Smaller position indices correspond to more recent user interactions.
Cite this review
Pith. "Pith review of Teacher Retains Full Tokens, Student Merges Efficiently: TM20K for E-Commerce Sequence Modeling in Ad Recommendation." pith.science (2026). https://pith.science/paper/4EJUQYN5
@misc{pith2026260807055,
author = {Pith},
title = {Pith review of: Teacher Retains Full Tokens, Student Merges Efficiently: TM20K for E-Commerce Sequence Modeling in Ad Recommendation},
year = {2026},
howpublished = {\url{https://pith.science/paper/4EJUQYN5}},
note = {Machine review of arXiv:2608.07055}
}
read the original abstract
Benefiting from ultra-long behavior sequence modeling, existing recommender systems bring users a better experience via simultaneously considering their long-term and short-term interests. Nevertheless, extended sequence lengths introduce substantial burdens on training efficiency and serving throughput. Prior approaches typically utilize search-based or cluster-based compression on ultra-long sequences at the cost of fine-grained information, or rely on various lightweight target attention structures incapable of sufficient sequential feature extraction. In this paper, we balance the effectiveness and efficiency for ultra-long sequence modeling via full transformer modeling accompanied with a two-stage knowledge distillation framework. First, both teacher and student models take the full attention mechanism rather than pure target-sequence attention for effective sequence scaling. For student models, we propose several simple yet well-motivated token merge approaches, significantly compressing the sequence length while maintaining an acceptable performance. Then, a one-time teacher is heavily trained with full sequence tokens, further boosting the performance of student models via knowledge distillation. The proposed paradigm named TM20K has been successfully deployed in ByteDance's e-commerce advertising recommender system that extends the e-commerce sequence length to 20K, delivering substantial improvements in key business metrics (e.g., ADSS +1.036\%) while keeping the training and serving cost nearly the same as the online state-of-the-art model (e.g., serving latency only +5.6\%).
Figures
Reference graph
Works this paper leans on
-
[1]
Amit Ben Artzy and Roy Schwartz. 2024. Attend first, consolidate later: On the importance of attention in different llm layers. InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP. 177–184
work page 2024
-
[2]
Zheng Chai, Qin Ren, Xijun Xiao, Huizhi Yang, Bo Han, Sijun Zhang, Di Chen, Hui Lu, Wenlin Zhao, Lele Yu, et al . 2025. Longer: Scaling up long sequence modeling in industrial recommenders. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 247–256
2025
-
[3]
Jianxin Chang, Chenbin Zhang, Zhiyi Fu, Xiaoxue Zang, Lin Guan, Jing Lu, Yiqun Hui, Dewei Leng, Yanan Niu, Yang Song, et al. 2023. TWIN: TWo-stage interest network for lifelong user behavior modeling in CTR prediction at kuaishou. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3785–3794
2023
-
[4]
Chenglong Chu, Guorui Zhou, Guowang Zhang, Han Li, Hao Peng, Hongtao Cheng, Jian Liang, Jiangxia Cao, Kun Gai, Lingzhi Zhou, et al . 2026. Kwai Summary Attention Technical Report.arXiv preprint arXiv:2604.24432(2026)
arXiv 2026
-
[5]
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. Flashat- tention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems35 (2022), 16344–16359
2022
-
[6]
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Peter Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al. 2023. Scaling vision transformers to 22 billion parameters. InInternational conference on machine learning. PMLR, 7480–7512
work page 2023
-
[7]
Haoran Ding, Wenlin Zhao, Yuchen Jiang, Juren Li, Jie Zhu, Xinchun Li, Yishujie Zhao, Yi Zhang, Ao Qiao, Jianhui Dong, et al. 2026. Rec-Distill: An Industrial Distillation Pipeline for Large-Scale Recommendation Models.arXiv preprint arXiv:2605.29755(2026)
arXiv 2026
-
[8]
Qin Ding, Kevin Course, Linjian Ma, Jianhui Sun, Ruochen Liu, Zhao Zhu, Chunx- ing Yin, Wei Li, Dai Li, Yu Shi, et al . 2026. Bending the scaling law curve in large-scale recommendation systems.arXiv preprint arXiv:2602.16986(2026)
arXiv 2026
Show all 39 references
-
[9]
Lin Guan, Jia-Qi Yang, Zhishan Zhao, Beichuan Zhang, Bo Sun, Xuanyuan Luo, Jinan Ni, Xiaowen Li, Yuhang Qi, Zhifang Fan, et al. 2025. Make It Long, Keep It Fast: End-to-End 10k-Sequence Modeling at Billion Scale on Douyin.arXiv preprint arXiv:2511.06077(2025)
2025 arXiv
-
[10]
Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji-Rong Wen. 2022. Towards universal sequence representation learning for recommender systems. InProceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining. 585–593
2022
-
[11]
Yunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin, Xuanyuan Luo, Zhe Wang, Zheng Chai, Shikang Wu, Yuchao Zheng, and Jingjian Lin. 2026. HyFormer: Revis- iting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction. arXiv preprint arXiv:2601.12681(2026)
2026
-
[12]
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. 2021. Perceiver: General perception with iterative attention. InInternational conference on machine learning. PMLR, 4651–4664
2021
-
[13]
Shali Jiang, Hua Zheng, Boyang Liu, Laming Chen, Kenny Lov, Chuanqi Xu, Lisang Ding, Qinghai Zhou, Can Cui, Xiaolong Liu, et al. 2026. LoopFM: Learning frOm HistOrical RePresentations of Foundation Model for Recommendation. arXiv preprint arXiv:2605.29280(2026)
2026 arXiv
-
[14]
Nikhil Khani, Li Wei, Aniruddh Nath, Shawn Andrews, Shuo Yang, Yang Liu, Pendo Abbo, Maciej Kula, Jarrod Kahn, Zhe Zhao, et al. 2024. Bridging the gap: Unpacking the hidden challenges in knowledge distillation for online ranking systems. InProceedings of the 18th ACM Conferenc...
2024
-
[15]
Weijiang Lai, Beihong Jin, Jiongyan Zhang, Yiyuan Zheng, Jian Dong, Jia Cheng, Jun Lei, and Xingxing Wang. 2025. Exploring Scaling Laws of CTR Model for Online Performance Improvement. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 114–123
2025
-
[16]
Kaiyuan Li, Yongxiang Tang, Yanhua Cheng, Yong Bai, Yanxiang Zeng, Chao Wang, Xialong Liu, and Peng Jiang. 2025. VQL: An End-to-End Context-Aware Vector Quantization Attention for Ultra-Long User Behavior Modeling.arXiv preprint arXiv:2508.17125(2025)
2025 arXiv
-
[17]
Xinchun Li, Ning Zhang, Qianqian Yang, Fei Teng, Wenlin Zhao, Huizhi Yang, Heng Shi, Linlan Chen, Yixin Wu, Zhen Wang, et al. 2026. IAT: Instance-As-Token Compression for Historical User Sequence Modeling in Industrial Recommender Systems.arXiv preprint arXiv:2604.08933(2026)
2026 arXiv
-
[18]
Mingyang Liu, Yong Bai, Zhangming Chan, Sishuo Chen, Xiang-Rong Sheng, Han Zhu, Jian Xu, and Xinyang Chen. 2026. EST: Towards Efficient Scaling Laws in Click-Through Rate Prediction via Unified Modeling.arXiv preprint arXiv:2602.10811(2026)
2026
-
[19]
Zhiwei Liu, Ziwei Fan, Yu Wang, and Philip S Yu. 2021. Augmenting sequential recommendation with pseudo-prior items via reversely pre-training transformer. InProceedings of the 44th international ACM SIGIR conference on Research and development in information retrieval. 1608–1612
2021
-
[20]
Xiao Lv, Jiangxia Cao, Shijie Guan, Xiaoyou Zhou, Zhiguang Qi, Yaqiang Zang, Ben Wang, and Guorui Zhou. 2025. MARM: Unlocking the Recommendation Cache Scaling-Law through Memory Augmentation and Scalable Complexity. InProceedings of the 34th ACM International Conference on Inf...
2025
-
[21]
Wenhan Lyu, Devashish Tyagi, Yihang Yang, Ziwei Li, Ajay Somani, Karthikeyan Shanmugasundaram, Nikola Andrejevic, Ferdi Adeputra, Curtis Zeng, Arun K Singh, et al. 2025. DV365: Extremely Long User History Modeling at Instagram. InProceedings of the 31st ACM SIGKDD Conference o...
2025
-
[22]
Qi Pi, Guorui Zhou, Yujing Zhang, Zhe Wang, Lejian Ren, Ying Fan, Xiaoqiang Zhu, and Kun Gai. 2020. Search-based user interest modeling with lifelong sequential behavior data for click-through rate prediction. InProceedings of the 29th ACM International Conference on Informati...
2020
-
[23]
Ads Recommendation. 2025. External Large Foundation Model: How to Efficiently Serve Trillions of Parameters for Online Ads Recommendation.arXiv preprint arXiv:2502.17494(2025)
2025 arXiv
-
[24]
Zihua Si, Lin Guan, ZhongXiang Sun, Xiaoxue Zang, Jing Lu, Yiqun Hui, Xingchao Cao, Zeyu Yang, Yichen Zheng, Dewei Leng, et al. 2024. Twin v2: Scaling ultra- long user behavior sequence modeling for enhanced ctr prediction at kuaishou. InProceedings of the 33rd ACM Internation...
2024
-
[25]
Xin Song, Zhilin Guan, Ruidong Han, Binghao Tang, Tianwen Chen, Bing Li, Zihao Li, Han Zhang, Fei Jiang, Qing Wang, et al. 2026. MTFM: A Scalable and Alignment-free Foundation Model for Industrial Recommendation in Meituan. arXiv preprint arXiv:2602.11235(2026)
2026
-
[26]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.Advances in neural information processing systems30 (2017)
2017
-
[27]
Jesse Vig and Yonatan Belinkov. 2019. Analyzing the structure of attention in a transformer language model. InProceedings of the 2019 ACL workshop Black- boxNLP: analyzing and interpreting neural networks for NLP. 63–76
2019
-
[28]
Shuli Wang. 2026. Sample Is Feature: Beyond Item-Level, Toward Sample-Level Tokens for Unified Large Recommender Models.arXiv preprint arXiv:2604.15650 (2026)
2026 arXiv
-
[29]
Zhuoxing Wei, Qi Liu, and Qingchen Xie. 2025. Deep Multiple Quantization Network on Long Behavior Sequence for Click-Through Rate Prediction. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval. 3090–3094
2025
-
[30]
Xue Xia, Saurabh Joshi, Kousik Rajesh, Kangnan Li, Yangyi Lu, Nikil Pancha, Dhruvil Badani, Jiajing Xu, and Pong Eksombatchai. 2025. TransAct V2: Lifelong User Action Sequence Modeling on Pinterest Recommendation. InProceedings of the 34th ACM International Conference on Infor...
2025
-
[31]
Lei Xin, Yuhao Zheng, Ke Cheng, Changjiang Jiang, Zifan Zhang, and Fanhu Zeng. 2026. Hytrec: A hybrid temporal-aware attention architecture for long behavior sequential recommendation.arXiv preprint arXiv:2602.18283(2026)
2026
-
[32]
Lee Xiong, Zhirong Chen, Rahul Mayuranath, Shangran Qiu, Arda Ozdemir, Lu Li, Yang Hu, Dave Li, Jingtao Ren, Howard Cheng, et al. 2026. LLaTTE: Scaling Laws for Multi-Stage Sequence Modeling in Large-Scale Ads Recommendation. arXiv preprint arXiv:2601.20083(2026)
2026
-
[33]
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Yuxing Wei, Lean Wang, Zhiping Xiao, et al. 2025. Native sparse attention: Hardware-aligned and natively trainable sparse attention. In Proceedings of the 63rd Annual Meeting of the Associ...
2025
-
[34]
Kun Yuan, Junyu Bi, Daixuan Cheng, Changfa Wu, Shuwen Xiao, Binbin Cao, Jian Wu, and Yuning Jiang. 2026. HiSAC: Hierarchical Sparse Activation Com- pression for Ultra-long Sequence Modeling in Recommenders.arXiv preprint arXiv:2602.21009(2026)
2026 arXiv
-
[35]
Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhao- jie Gong, Fangda Gu, Michael He, et al. 2024. Actions speak louder than words: Trillion-parameter sequential transducers for generative recommendations.arXiv preprint arXiv:2402.17152(2024)
2024 arXiv
-
[36]
Shangyu Zhang, Shijie Quan, Zhongren Wang, Junwei Pan, Tianqu Zhuang, Bo Fu, Yilong Sun, Jieying Lin, Jushuo Chen, Xiaotian Li, et al . 2025. Large Foundation Model for Ads Recommendation.arXiv preprint arXiv:2508.14948 (2025)
2025 arXiv
-
[37]
Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Deep interest evolution network for click-through rate prediction. InProceedings of the AAAI conference on artificial intelligence, Vol. 33. 5941–5948
2019
-
[38]
Wen-Ji Zhou, Yuhang Zheng, Yinfu Feng, Yunan Ye, Rong Xiao, Long Chen, Xiaosong Yang, and Jun Xiao. 2024. ENCODE: Breaking the trade-off between performance and efficiency in long-term user behavior modeling.IEEE Transac- tions on Knowledge and Data Engineering37, 1 (2024), 265–277
2024
-
[39]
Jie Zhu, Zhifang Fan, Xiaoxie Zhu, Yuchen Jiang, Hangyu Wang, Xintian Han, Haoran Ding, Xinmin Wang, Wenlin Zhao, Zhen Gong, et al. 2025. Rankmixer: Scaling up ranking models in industrial recommenders. InProceedings of the Conference acronym ’XX, June 03–05, 2018, Woodstock, ...
2025
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.