Pith. sign in

REVIEW 3 major objections 5 minor 46 references

CCFormer ranks long user histories under industrial latency by separating cross-field attention from compressed subspace sequence mixing.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 18:44 UTC pith:ZPW4W7N7

load-bearing objection Solid production ranking backbone with real deployment evidence; novelty is recombination, and the headline CTR lift is vs a DLRM, not the strong sequential baselines beaten offline. the 3 major comments →

arxiv 2607.28070 v1 pith:ZPW4W7N7 submitted 2026-07-30 cs.IR

CCFormer: Efficient Cross-Field Interaction and Hierarchical Sequence Compression for Industrial Recommendation at Tencent

classification cs.IR
keywords recommender systemssequential recommendationlong-sequence modelingcross-field attentiontoken mixingsequence compressionscaling lawCTR prediction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Industrial recommenders want longer behavior histories and bigger models because scaling improves accuracy, but full self-attention over long sequences is too slow and costly to serve. This paper argues that you can keep the accuracy gains without paying the full quadratic bill by splitting the job: directed cross-attention lets user, history, and target fields interact early and explicitly, while a lightweight subspace token mixer plus progressive convolutional compression handles the long history itself. Compression shrinks the sequence layer by layer so deeper layers see larger receptive fields at much lower cost. On public and Tencent-scale data the design beats strong sequential baselines; live A/B tests report a 3.57% CTR lift in video recommendation and a 1.71% ad-revenue lift, with training about 2.21× faster than HSTU, and the model now carries main production traffic.

Core claim

CCFormer shows that feature-field separated cross-attention, combined with long-sequence subspace token mixing and hierarchical Conv1D compression, is enough to exploit long-term preference signals across heterogeneous fields while staying inside industrial latency and resource budgets—outperforming strong sequential baselines offline and delivering statistically significant online CTR and revenue gains at higher training throughput.

What carries the argument

The CCFormer block: three directed cross-attention flows (user→sequence, target→sequence, target→user), relative temporal-positional encoding inside local subspaces, per-channel subspace token mixing (not full self-attention), then hierarchical 1D convolutional downsampling that expands each token’s receptive field across layers.

Load-bearing premise

That repeatedly compressing the behavior sequence with strided convolution keeps the preference signal that ranking actually needs, rather than throwing away rare but decisive long-range behaviors.

What would settle it

On the same industrial traffic, keep all other parts fixed and compare hierarchical compression against an otherwise identical model that retains the full-length sequence (or uses lossless retrieval of the same history): if full length clearly wins on AUC/GAUC and online CTR/revenue while compression’s gains disappear, the efficiency story fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Longer user histories can be used in production ranking without full quadratic self-attention cost.
  • Training throughput can rise (reported ~2.21× vs HSTU) while accuracy still scales with sequence length and model width.
  • Candidate items can be scored in one shared user/sequence pass, raising serving QPS under fixed hardware.
  • The same backbone can serve both content recommendation and ad ranking traffic once deployed.
  • Future lifelong-scale sequences become more practical if compression-aware modeling extends to retrieval as the paper plans.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If compression is nearly lossless on average, the next bottleneck shifts from sequence length to how well rare, high-value behaviors are protected before downsampling.
  • Field-separated early fusion may transfer to other multi-entity ranking settings (search, social feed) where user, context, and candidate must cross without full token self-attention.
  • A natural stress test is traffic slices with sparse or bursty interests, where progressive pooling is most likely to erase decisive events.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. CCFormer is an industrial Transformer backbone for long-sequence CTR/ranking that decouples cross-field interaction from intra-sequence modeling. It uses (i) feature-field separated directed cross-attention among user, behavior-sequence, and target tokens, (ii) relative temporal-positional encoding inside local sequence groups, (iii) subspace token mixing via per-channel FFNs instead of full self-attention over the behavior sequence, and (iv) hierarchical Conv1D downsampling that shortens the sequence across layers while expanding receptive fields. The paper reports gains on Taobao and KuaiRec, modest but consistent AUC/GAUC lifts on a >4B-sample Tencent log versus HSTU/OneTrans/STCA, sequence- and width-scaling curves, ablations with training-throughput speedups (up to 2.21× vs HSTU), and two online A/B tests (video recommendation and ad ranking) culminating in full production deployment.

Significance. If the results hold under clearer baseline and systems attribution, this is a useful systems-level contribution for industrial sequential ranking: a deployable recipe that couples early cross-field fusion with cheaper long-sequence mixing and progressive compression, backed by public repeats, industrial scaling plots, component ablations that separate accuracy from throughput, and real A/B lifts with full traffic rollout. Strengths include explicit multi-field attention design, GFLOPs-aware scaling tables, and production details (parallel multi-candidate scoring, sparse compression). Novelty is largely compositional rather than a new theoretical principle, but the empirical package is in the standard valuable range for KDD-style industrial recommender work.

major comments (3)
  1. [§4.5 Online A/B Results; Abstract] §4.5 and the abstract juxtapose the 3.57% CTR lift with the 2.21× training speedup “over the strong HSTU baseline,” but Scenario 1’s online baseline is explicitly “a long-standing, iteratively optimized DLRM,” while only Scenario 2 uses deployed HSTU. Offline industrial gains vs HSTU/STCA are modest (Table 3: +0.28 AUC / +0.50 GAUC vs HSTU; +0.21 AUC vs STCA). The headline online CTR number therefore does not establish online superiority over the same strong sequential models used offline and in Table 5. Please restate online baselines per scenario in the abstract and §4.5, and avoid implying a single HSTU-relative online story.
  2. [§3.5 Training and Deployment Optimization; Table 5] §3.5 co-ships mixed precision, INT8 embedding quantization, double-hashing (~50% sparse-table cut), and parallel multi-candidate target packing (+30% peak QPS despite claimed 20× compute). Online revenue/QPS and even training speedup are therefore not isolated to field-separated attention, subspace mixing, or hierarchical Conv1D. Either (a) report an architecture-only ablation under fixed systems stack for the industrial offline metrics and serving QPS, or (b) clearly scope the 2.21× / QPS claims as full-stack results and separate “model” vs “systems” contributions in Table 5 and §3.5.
  3. [§3.4; Table 5; Abstract] §3.4 and the abstract claim hierarchical compression enables long-sequence modeling “with reduced information loss,” illustrated by expanding receptive fields (Fig. 3). Table 5 shows removing compression yields essentially unchanged or slightly higher AUC/GAUC (77.96/71.25 vs 77.94/71.36) while collapsing speedup (2.21×→1.29×). That supports compression as primarily an efficiency device on average metrics, not a demonstrated preference-preserving multi-granularity encoder. Soften the lossless/reduced-loss language, or add evidence that rare/long-tail behaviors and slice-level metrics are not systematically hurt (e.g., long-history users, cold items).
minor comments (5)
  1. [Table 3; §4.1] Table 2 reports mean±std over 5 runs; Table 3 industrial results have no uncertainty or significance tests. Add repeats or bootstrap CIs for industrial AUC/GAUC, especially given small absolute deltas.
  2. [§3.2] Eqs. (9)–(11): clarify whether W_time and W_pos are applied as a multiplicative/additive bias on tokens before mixing or as a true pairwise reweighting inside each group; the matrix–slice product notation is easy to misread relative to standard relative-bias attention.
  3. [Figure 1] Fig. 1’s three-column comparison is dense; label which panel is CCFormer and align terminology with §3 (cross-field vs “Feature Interaction Block”).
  4. [§4 Baselines; Table 6] RelaImpr formula in §4 uses (Metric−0.5)/(base−0.5); state explicitly that this is only for AUC-like metrics bounded at 0.5, and do not apply the same formula language to revenue/CTR lifts in Table 6.
  5. [Throughout; References] Typos/style: “CCFormer” sometimes missing space after propose; arXiv-dated refs and KDD ’27 placeholder venue should be cleaned for camera-ready; ensure HSTU/OneTrans/STCA citations match the exact production variants compared.

Circularity Check

0 steps flagged

No circularity: empirical architecture paper evaluated on external offline/online metrics, not a self-defining derivation.

full rationale

CCFormer is an engineering architecture paper. Its load-bearing claims are comparative performance (AUC/GAUC on public and industrial logs; online CTR/revenue lifts; training throughput vs HSTU), not first-principles predictions derived from fitted constants or uniqueness theorems. Module definitions (field-separated cross-attention Eqs. 2–8, subspace PFFN mixing Eqs. 12–15, Conv1D hierarchical compression Eq. 16) do not define the evaluation targets; those targets are held-out click labels and live A/B business metrics. Ablations (Table 5) and scaling sweeps (Figs. 4–5, Table 4) vary components and measure external outcomes rather than recovering fitted inputs by construction. Self-references to the internal stack (Numerous-Torch, mixed precision, INT8/double-hash) describe deployment infrastructure and do not force the accuracy claims. No self-citation uniqueness theorem, no ansatz smuggled as a theorem, and no renaming of a known identity presented as a derivation. Baseline asymmetry and systems confounding (noted by the skeptic) are evaluation-validity concerns, not circularity. Score 0 is appropriate.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 3 invented entities

Load-bearing content is architectural and empirical, not axiomatic physics. Claims rest on standard deep-learning practice, recsys task assumptions, chosen hyperparameters, and proprietary production data. No new physical entities; invented pieces are named modules whose only evidence is ablation and A/B performance.

free parameters (6)
  • feature dim d / depth L = d=256, L=8 (default industrial)
    Model capacity knobs set for industrial runs (default d=256, L=8) and swept in scaling tables; performance and GFLOPs depend on these choices.
  • subspace sizes m, n = m=8, n=16
    Sequence group size and channel group size define the token-mixing reshape; chosen (m=8, n=16) and swept in Fig. 6.
  • conv kernel k and stride s = k=3, s=2
    Hierarchical compression rate and local fusion width; s=2 fixed industrially, k swept 2–7.
  • temporal encoding α, β, γ = learned (β∈(0,1))
    Learnable/time-decay parameters in relative temporal weights W_time inside each group (§3.2); fitted during training.
  • sequence length L_s and truncation = 200 public; 1000 industrial default
    Public runs truncate at 200; industrial default 1000 with scaling to 0.5k–2k—directly affects reported scaling laws.
  • Adam lr and batch size = lr=1e-4, batch=4096
    Optimization hyperparameters fixed for fair comparison but still free experimental choices.
axioms (5)
  • domain assumption Directed cross-attention among user/sequence/target fields without full all-to-all self-attention is sufficient to capture the CTR-relevant interactions (user↔history, target↔history, target↔user).
    Core design premise in §3.1; if important interactions require denser mixing across all tokens, the efficiency design under-models them.
  • domain assumption Local subspace token mixing (PFFN on reshaped m×n blocks) plus hierarchical compression approximates the useful long-range sequential dependencies that full self-attention would provide for ranking.
    Stated motivation in §3.3–3.4; ablation replaces mixing with self-attention and finds small metric change at large speed cost.
  • domain assumption Offline AUC/GAUC improvements and two-week A/B lifts with p<0.05 predict stable production value under continual traffic and feedback shift.
    Standard industrial recsys evaluation assumption; paper reports post-launch stability but cannot be externally audited.
  • standard math Standard Transformer/MLP training facts (softmax attention, residual+RMSNorm paths, Adam, mixed precision) behave as usual and do not introduce unmodeled failure modes.
    Background DL machinery used throughout §3 without new proof obligations.
  • ad hoc to paper INT8 embedding quantization and double-hashing of ID tables preserve enough representation quality for the reported lifts.
    Deployment optimization in §3.5 claimed not to degrade training effectiveness; supports industrial practicality claim but is stack-specific.
invented entities (3)
  • CCFormer block (field-separated dual cross-attention + subspace mixing + conv compression) no independent evidence
    purpose: Name the stacked unit that jointly does cross-field interaction and cheap long-sequence modeling for industrial ranking.
    Composite architecture defined in §3; evidence is empirical tables/A/B only, not an external physical referent.
  • Long-sequence subspace token mixing (PFFN on behavior subspaces) no independent evidence
    purpose: Replace O(L_s²) self-attention over behaviors with localized multi-token/channel mixing.
    Adaptation of token-mixing ideas to full behavior sequences; validated only via ablations inside this paper’s tasks.
  • Hierarchical long-sequence token compression with expanding receptive fields no independent evidence
    purpose: Progressively shorten S across layers while claiming multi-granularity interest modeling.
    Conv1D k,s schedule in §3.4/Fig. 3; independent evidence limited to speed/quality tradeoff tables here.

pith-pipeline@v1.2.0-daily-grok45 · 20024 in / 4159 out tokens · 85721 ms · 2026-07-31T18:44:56.899974+00:00 · methodology

0 comments
read the original abstract

Recent studies in industrial recommendation systems have demonstrated that sequential recommendation models built upon self-attention can benefit from predictable scaling laws by increasing sequence length and model capacity. However, practical recommender systems impose strict latency and resource constraints, making it challenging to balance computational overhead with fine-grained feature interaction. In this paper, we propose CCFormer, an efficient Transformer backbone that unifies cross-field feature interaction and compressed long-sequence modeling for industrial recommendation. Specifically, CCFormer combines feature-field separated cross attention with long-sequence subspace token mixing to exploit long-term preference signals across heterogeneous feature domains. A hierarchical sequence compression strategy with progressively expanded receptive fields enables efficient long-sequence modeling with reduced information loss. Extensive experiments on two public benchmarks and a large-scale industrial dataset demonstrate that CCFormer consistently outperforms state-of-the-art baselines. Online A/B tests in a video recommendation scenario and an advertising ranking scenario at Tencent further validate its industrial practicality, yielding a 3.57% CTR gain and a 1.71% advertising revenue lift, respectively, while accelerating model training by 2.21x over the strong HSTU baseline. CCFormer has been fully deployed in Tencent's production recommendation system, serving the main traffic of both scenarios.

Figures

Figures reproduced from arXiv: 2607.28070 by Bing Wen, Chengxiang Zhuo, Haonan Hu, Huizhe Zhang, Jianchao Tu, Yudong Li, Yunlong Wang, Zang Li.

Figure 1
Figure 1. Figure 1: Architectural comparison. traffic under strict latency and resource constraints. Consequently, directly extending the input sequence length or increasing model capacity often incurs unacceptable training and inference costs, making Transformer-based sequential models difficult to scale in practical systems. To alleviate the efficiency bottleneck of long-sequence modeling, existing long-sequence recommendat… view at source ↗
Figure 2
Figure 2. Figure 2: The overview of CCFormer. 𝑝-th sequence group, the relative temporal-positional encoding is defined as follows: Wtime 𝑝,𝑖,𝑗 = 𝛼 · 𝛽 |𝑡𝑝,𝑖 −𝑡𝑝, 𝑗 | 𝛾 , (9) where 𝛼 and 𝛾 are positive learnable parameters, 𝛽 ∈ (0, 1) is a time-decay parameter, and 𝑡𝑝,𝑖 is the 𝑖-th behavior timestamp in the 𝑝-th group. We construct a relative temporal encoding ma￾trix Wtime 𝑝 ∈ R 𝑚×𝑚 for the 𝑝-th group. Since 𝛽 ∈ (0, 1), the … view at source ↗
Figure 3
Figure 3. Figure 3: Change in sequence token receptive field with in [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: AUC comparison of four sequential recommenda [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Scaling law with sequence length. We vary the se [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

46 extracted references · 7 linked inside Pith

  1. [1]

    James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl. 2011. Algorithms for Hyper-Parameter Optimization. InAdvances in Neural Information Processing Systems, J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, and K. Weinberger CCFormer: Efficient Cross-Field Interaction and Sequence Compression KDD ’27, August 2027, San Jose, CA, USA (Eds.), Vol...

  2. [2]

    Zheng Chai, Qin Ren, Xijun Xiao, Huizhi Yang, Bo Han, Sijun Zhang, Di Chen, Hui Lu, Wenlin Zhao, Lele Yu, et al . 2025. Longer: Scaling up long sequence modeling in industrial recommenders. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 247–256

  3. [3]

    Jianxin Chang, Chenbin Zhang, Zhiyi Fu, Xiaoxue Zang, Lin Guan, Jing Lu, Yiqun Hui, Dewei Leng, Yanan Niu, Yang Song, and Kun Gai. 2023. TWIN: TWo-stage Interest Network for Lifelong User Behavior Modeling in CTR Prediction at Kuaishou. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3785–3794

  4. [4]

    Jin Chen, Shangyu Zhang, Bin Hu, Chao Zhou, Junwei Pan, Gengsheng Xue, Wentao Ning, Gengyu Weng, Wang Zheng, Shaohua Liu, et al. 2026. RankUp: Towards High-rank Representations for Large Scale Advertising Recommender Systems.arXiv preprint arXiv:2604.17878(2026)

  5. [5]

    Qiwei Chen, Huan Zhao, Wei Li, Pipei Huang, and Wenwu Ou. 2019. Behavior sequence transformer for e-commerce recommendation in alibaba. InProceedings of the 1st international workshop on deep learning practice for high-dimensional sparse data. 1–4

  6. [6]

    Xiangyi Chen, Kousik Rajesh, Matthew Lawhon, Zelun Wang, Hanyu Li, Haomiao Li, Saurabh Vishwas Joshi, Pong Eksombatchai, Jaewon Yang, Yi-Ping Hsu, et al

  7. [7]

    Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al

  8. [8]

    Sunhao Dai, Jiakai Tang, Jiahua Wu, Kun Wang, Yuxuan Zhu, Bingjun Chen, Bangyang Hong, Yu Zhao, Cong Fu, Kangle Wu, et al. 2025. Onepiece: Bringing context engineering and reasoning to industrial cascade ranking system.arXiv preprint arXiv:2509.18091(2025)

  9. [9]

    Qin Ding, Kevin Course, Linjian Ma, Jianhui Sun, Ruochen Liu, Zhao Zhu, Chunx- ing Yin, Wei Li, Dai Li, Yu Shi, et al . 2026. Bending the scaling law curve in large-scale recommendation systems.arXiv preprint arXiv:2602.16986(2026)

  10. [10]

    Yufei Feng, Fuyu Lv, Weichen Shen, Menghan Wang, Fei Sun, Yu Zhu, and Keping Yang. 2019. Deep session interest network for click-through rate prediction. arXiv preprint arXiv:1905.06482(2019)

  11. [11]

    Yulong Gu, Lixin Zou, and Chenliang Li. 2026. Deep Learning to Rank in Industrial Search Engines, Recommender Systems, and Online Advertising: An Overview and New Perspectives.ACM Transactions on Information Systems44, 4 (2026), 1–52

  12. [12]

    Lin Guan, Jia-Qi Yang, Zhishan Zhao, Beichuan Zhang, Bo Sun, Xuanyuan Luo, Jinan Ni, Xiaowen Li, Yuhang Qi, Zhifang Fan, et al. 2026. Make it long, keep it fast: End-to-end 10k-sequence modeling at billion scale on Douyin Recommendation. InProceedings of the ACM Web Conference 2026. 7989–7998

  13. [13]

    Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: a factorization-machine based neural network for CTR prediction.arXiv preprint arXiv:1703.04247(2017)

  14. [14]

    Mingming Ha, Guanchen Wang, Linxun Chen, Xuan Rao, Yuexin Shi, Tianbao Ma, Zhaojie Liu, Yunqian Fan, Zilong Lu, Yanan Niu, et al . 2026. UniMixer: A Unified Architecture for Scaling Laws in Recommendation Systems.arXiv preprint arXiv:2604.00590(2026)

  15. [15]

    Ruidong Han, Bin Yin, Shangyu Chen, et al. 2025. MTGR: Industrial-Scale Gener- ative Recommendation Framework in Meituan.arXiv preprint arXiv:2505.18654 (2025)

  16. [16]

    Xu Huang, Hao Zhang, Zhifang Fan, Yunwen Huang, Zhuoxing Wei, Zheng Chai, Jinan Ni, Yuchao Zheng, and Qiwei Chen. 2026. MixFormer: Co-Scaling Up Dense and Sequence in Industrial Recommenders.arXiv preprint arXiv:2602.14110 (2026)

  17. [17]

    Yuchen Jiang, Jie Zhu, Xintian Han, Hui Lu, Kunmin Bai, Mingyu Yang, Shikang Wu, Ruihao Zhang, Wenlin Zhao, Shipeng Bai, et al. 2026. TokenMixer-Large: Scaling Up Large Ranking Models in Industrial Recommenders.arXiv preprint arXiv:2602.06563(2026)

  18. [18]

    Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recom- mendation. In2018 IEEE international conference on data mining (ICDM). IEEE, 197–206

  19. [19]

    Weijiang Lai, Beihong Jin, Di Zhang, Siru Chen, Jiongyan Zhang, Yuhang Gou, Jian Dong, and Xingxing Wang. 2026. Unleashing the Potential of Sparse Atten- tion on Long-term Behaviors for CTR Prediction. InProceedings of the ACM Web Conference 2026. 8041–8050

  20. [20]

    Yang Li, Tong Chen, Peng-Fei Zhang, and Hongzhi Yin. 2021. Lightweight self- attentive sequential recommendation. InProceedings of the 30th ACM international conference on information & knowledge management. 967–977

  21. [21]

    Guanyu Lin, Jinwei Luo, Yinfeng Li, Chen Gao, Qun Luo, and Depeng Jin. 2025. Iterative sparse attention for long-sequence recommendation. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 12147–12155

  22. [22]

    Mingyang Liu, Yong Bai, Zhangming Chan, Sishuo Chen, Xiang-Rong Sheng, Han Zhu, Jian Xu, and Xinyang Chen. 2026. EST: Towards Efficient Scaling Laws in Click-Through Rate Prediction via Unified Modeling.arXiv preprint arXiv:2602.10811(2026)

  23. [23]

    Mengyang Ma, Xiaopeng Li, Wanyu Wang, Zhaocheng Du, Jingtong Gao, Pengyue Jia, Yuyang Ye, Yiqi Wang, Yunpeng Weng, Weihong Luo, et al. 2026. Blossomrec: Block-level fused sparse attention mechanism for sequential recom- mendations. InProceedings of the ACM Web Conference 2026. 6389–6399

  24. [24]

    Qi Pi, Weijie Bian, Guorui Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Practice on long sequential user behavior modeling for click-through rate prediction. InProceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining. 2671–2679

  25. [25]

    Qi Pi, Guorui Zhou, Yujing Zhang, Zhe Wang, Lejian Ren, Ying Fan, Xiaoqiang Zhu, and Kun Gai. 2020. Search-based User Interest Modeling with Lifelong Sequential Behavior Data for Click-Through Rate Prediction. InProceedings of the 29th ACM International Conference on Information & Knowledge Management. 2685–2692

  26. [26]

    Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang

  27. [27]

    Dan Svenstrup, Jonas Meinertz Hansen, and Ole Winther. 2017. Hash Embeddings for Efficient Word Representations. InAdvances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. ...

  28. [28]

    Qiaoyu Tan, Jianwei Zhang, Jiangchao Yao, Ninghao Liu, Jingren Zhou, Hongxia Yang, and Xia Hu. 2021. Sparse-interest network for sequential recommendation. InProceedings of the 14th ACM international conference on web search and data mining. 598–606

  29. [29]

    Jiakai Tang, Sunhao Dai, Teng Shi, Jun Xu, Xu Chen, Wen Chen, Jian Wu, and Yuning Jiang. 2026. Think before recommend: Unleashing the latent reasoning power for sequential recommendation.IEEE Transactions on Knowledge and Data Engineering(2026)

  30. [30]

    Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al. 2021. Mlp-mixer: An all-mlp architecture for vision.Advances in neural information processing systems34 (2021), 24261–24272

  31. [31]

    Yifei Xia, Suhan Ling, Fangcheng Fu, Yujie Wang, Huixia Li, Xuefeng Xiao, and Bin Cui. 2025. Training-free and adaptive sparse attention for efficient long video generation. InProceedings of the IEEE/CVF International Conference on Computer Vision. 15982–15993

  32. [32]

    Weinan Xu, Hengxu He, Minshi Tan, Yunming Li, Jun Lang, and Dongbai Guo

  33. [33]

    Bencheng Yan, Yuejie Lei, Zhiyuan Zeng, Di Wang, Kaiyi Lin, Pengjie Wang, Jian Xu, and Bo Zheng. 2025. From Scaling to Structured Expressivity: Rethinking Transformers for CTR Prediction.arXiv preprint arXiv:2511.12081(2025)

  34. [34]

    Yuhao Yang, Zhi Ji, Zhaopeng Li, Yi Li, Zhonglin Mo, Yue Ding, Kai Chen, Zijian Zhang, Jie Li, LIU LIN, et al . 2026. Sparse meets dense: Unified generative recommendations with cascaded sparse-dense representations.Advances in Neural Information Processing Systems38 (2026), 93746–93770

  35. [35]

    Yanwu Yang and Panyu Zhai. 2022. Click-through rate prediction in online advertising: A literature review.Information Processing & Management59, 2 (2022), 102853

  36. [36]

    Dezhi Yi, Wei Guo, Wenyang Cui, Wenxuan He, Huifeng Guo, Yong Liu, Zhen- hua Dong, and Ye Lu. 2026. FuXi-𝛾: Efficient Sequential Recommendation with Exponential-Power Temporal Encoder and Diagonal-Sparse Positional Mecha- nism. InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1. 1797–1808

  37. [37]

    Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Yuxing Wei, Lean Wang, Zhiping Xiao, et al. 2025. Native sparse attention: Hardware-aligned and natively trainable sparse attention. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 23078–23097

  38. [38]

    Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhao- jie Gong, Fangda Gu, Michael He, et al. 2024. Actions speak louder than words: Trillion-parameter sequential transducers for generative recommendations.arXiv preprint arXiv:2402.17152(2024)

  39. [39]

    Biao Zhang and Rico Sennrich. 2019. Root Mean Square Layer Normaliza- tion. InAdvances in Neural Information Processing Systems 32: Annual Con- ference on Neural Information Processing Systems 2019, NeurIPS 2019, Decem- ber 8-14, 2019, Vancouver, BC, Canada, Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roma...

  40. [40]

    Zhaoqi Zhang, Haolei Pei, Jun Guo, Tianyu Wang, Yufei Feng, Hui Sun, Shaowei Liu, and Aixin Sun. 2026. Onetrans: Unified feature interaction and sequence modeling with one transformer in industrial recommender. InProceedings of the ACM Web Conference 2026. 8162–8170

  41. [41]

    Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai. 2018. Deep interest network for click-through rate prediction. InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining. 1059–1068

  42. [42]

    Jie Zhu, Zhifang Fan, Xiaoxie Zhu, Yuchen Jiang, Hangyu Wang, Xintian Han, Haoran Ding, Xinmin Wang, Wenlin Zhao, Zhen Gong, et al. 2025. Rankmixer: Scaling up ranking models in industrial recommenders. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 6309–6316

  43. [2016]

    InProceedings of the 1st workshop on deep learning for recommender systems

    Wide & deep learning for recommender systems. InProceedings of the 1st workshop on deep learning for recommender systems. 7–10

  44. [2019]

    InProceedings of the 28th ACM international conference on information and knowledge management

    BERT4Rec: Sequential recommendation with bidirectional encoder rep- resentations from transformer. InProceedings of the 28th ACM international conference on information and knowledge management. 1441–1450

  45. [2020]

    InProceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval

    Deep interest with hierarchical attention network for click-through rate prediction. InProceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval. 1905–1908

  46. [2025]

    InProceedings of the Nineteenth ACM Conference on Recommender Systems

    Pinfm: foundation model for user activity sequences at a billion-scale visual discovery platform. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 381–390