Pith. sign in

REVIEW 3 major objections 6 minor 48 references

ROCS restructures recommendation models so that a user request's shared computation is done once and reused across all candidate items, cutting serving cost by up to 3x while preserving or improving prediction quality.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 02:05 UTC pith:SRATHTIS

load-bearing objection The GLM dependency contract is the real contribution; the quality-recovery story via RRR is the soft spot, and the production numbers need artifacts before they can be taken at face value. the 3 major comments →

arxiv 2607.27744 v1 pith:SRATHTIS submitted 2026-07-30 cs.LG cs.AIcs.IR

ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation

classification cs.LG cs.AIcs.IR
keywords request-oriented compute sharingrecommendation inferencecompute sharingfeature interactionsequence modelingGPU kernel optimizationscaling lawsefficient serving
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

ROCS starts from a simple observation: in recommendation serving, each user request is evaluated against many candidate items, but standard architectures recompute request-side representations separately for every candidate. The paper's central claim is that if every operator obeys a block lower-triangular dependency contract — request-side outputs depend only on request-side inputs while candidate-side outputs may consume both — then the request-side pathway can be evaluated once per request and reused verbatim across all candidates, with Eq. (11) guaranteeing the transformed module's outputs are unchanged relative to broadcasted execution. To make this practical, the paper adds a sequence-model extension (Deep Cross Attention), a capacity-reallocation rule (Request-Oriented Resource Reallocation), and a GPU execution scheme (In-Kernel Broadcast Optimization). If the claim holds, serving throughput rises by up to 3x on retrieval and 47–62% on ranking without quality loss — sometimes with a small quality gain — which lets systems afford larger candidate sets or heavier models at the same cost.

Core claim

The central discovery is that request–candidate redundancy is a structural property of recommendation models, not an implementation artifact. ROCS formalizes the condition for an intermediate result to be reusable — it must be invariant to the candidate — and turns it into an operator-level contract, the GLM invariant, which is enforced by applying block lower-triangular masks inside each operator and is closed under composition. The paper proves that every ROCSified module satisfies Eq. (11): for any set of candidates sharing one request, the request-side output is identical across all candidates and equals a function of the request alone, so it can be computed once and reused without pertu

What carries the argument

The load-bearing object is the block lower-triangular mask M_ij = 1{j≤i} applied inside every transformed operator — linear and linear-compression layers, group-local norms, factorization machines, and elementwise merges — so that output group i depends only on input groups up to i. This enforces the GLM dependency contract and, because the pattern is preserved under composition, yields Eq. (11), the exact reuse guarantee. Deep Cross Attention applies the same contract to sequences by computing a request-only encoding and K/V projections once, then letting request- and candidate-side queries attend to that shared representation. In-Kernel Broadcast Optimization realizes the contract on GPUs

Load-bearing premise

The quality-recovery story rests on the empirical bet that after GLM cuts off candidate-to-request information flow, the lost expressiveness can be made up by scaling the request-side pathway (the ratio r) in line with recommendation scaling laws; if that scaling does not deliver, ROCS trades accuracy for efficiency.

What would settle it

Run a fixed-FLOP comparison: train a vanilla early-fusion model and a ROCS model (with GLM/DCA and the same r-sweep) on a public benchmark, holding amortized per-candidate FLOPs equal at N=100; if the best ROCS model does not meet or beat the vanilla model's AUC/LogLoss across several backbones, then the RRR recovery premise fails. A simpler check: measure the AUC gap between ROCS-Base and Vanilla at N=1 (where there is no amortization benefit) — if scaling r cannot close that gap, the central quality claim collapses.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Retrieval and ranking systems can serve far more candidates per request at the same latency: up to 3x QPS on retrieval, 47–62% on ranking workloads.
  • Because Eq. (11) guarantees the request-side path is reused without changing the module's deterministic output, the efficiency gain is not an approximation; any quality change comes from the structural weight change alone.
  • Sequence models no longer pay per-candidate sequence encoding: DCA computes the user-sequence encoder and K/V projections once per request and still keeps candidate-conditioned retrieval via separate queries.
  • The compute saved can be reinvested as a quality lever (RRR): the paper shows quality improving as the request-to-candidate FLOPs ratio increases, so systems can trade throughput for accuracy in a controlled way.
  • ROCS is complementary to model unification, distillation, and low-precision execution, so it stacks with existing cost-reduction techniques.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same one-to-many scoring pattern occurs outside recommendation — a prompt scored against many candidate responses, a query against many documents — so the dependency-contract formulation could become a general serving principle for any large-scale scoring engine.
  • Because the reuse guarantee is exact, the practical question shifts to setting the request-side scaling ratio r; a natural extension the paper leaves open is adapting r per workload or per request based on candidate multiplicity and request-side FLOPs share.
  • The paper's boundary condition implies that ROCS's value is largest where candidate multiplicity is high and the request-side subgraph is a large fraction of total compute; measuring that fraction in an existing model is itself a useful diagnostic.
  • If RRR's scaling behavior holds, it suggests an additional axis for recommendation scaling laws — accuracy as a function of request-side capacity at fixed candidate-side capacity — which could guide compute budgeting across serving stages.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces ROCS, a modeling and inference paradigm for large-scale recommendation serving that exploits the many-candidates-per-request structure to share request-side computation. The method has four components: Generalized Layer Masking (GLM) is a block-triangular dependency contract that makes request-side outputs invariant to candidate inputs and is closed under composition; Deep Cross Attention (DCA) extends this to sequence models by separating a request-only shared sequence encoder from candidate-conditioned cross-attention; Request-Oriented Resource Reallocation (RRR) reinvests saved compute into a scaled request-side pathway controlled by a ratio r; and In-Kernel Broadcast Optimization (IKBO) implements the resulting computation on GPUs without materializing request-to-candidate broadcasts. Experiments on three public datasets and three backbones report consistent efficiency gains at N=100, and production/online results report large QPS gains with maintained or improved log-loss.

Significance. The mathematical core of ROCS, especially the GLM invariant and its composition closure in §3.1, is clean, parameter-free, and genuinely useful: it gives a precise condition under which request-side computation is reusable across candidates, and Eq. (11) correctly states the resulting invariant. The public experiments are FLOP-controlled and span diverse backbones, which is a strength. If the RRR-based quality-recovery claim is robust, ROCS would be a significant contribution to efficient recommendation serving. However, the quality-maintenance claim is currently supported only by small, single-run AUC differences, and the RRR scaling-ratio selection is opaque. The production and online numbers are unauditable at the detail provided. Thus the significance is real but not yet fully established by the evidence as submitted.

major comments (3)
  1. [§5.1, Table 1] All public-benchmark results are single runs with no variance information. The AUC deltas that carry the quality-maintenance claim are small (e.g., DCNv2/KuaiRand +0.0003, FinalMLP/KKBox +0.0007, Wukong/KuaiRand +0.0042) and are likely within seed-level noise for these models. Since the central claim is 'maintaining or improving prediction quality' on the quality–efficiency frontier, the paper should report repeated-seed means and standard deviations (or a full seed-level table), including for Figure 4 and Table 3. Without this, the positive-AUC conclusion cannot be distinguished from noise.
  2. [§3.3, Figure 4] ROCS's quality-recovery argument rests entirely on RRR, and RRR is controlled by the request-side scaling ratio r, a free parameter whose selection is not specified. The paper states in §3.3 that 'directly ROCSifying a model may therefore reduce prediction quality,' and Figure 4 sweeps r rather than deriving it from the cited scaling laws. The public ROCS-Scaled setup in Appendix A.1.2 only says capacity was 'progressively enlarged.' Please provide the exact selection rule for r, report the r values used for each dataset/backbone, and include a sensitivity analysis across r and across training seeds. Otherwise the improved-AUC results could be driven by favorable r selection.
  3. [§5.2.2, Table 2 / §5.3] The headline production claims (up to 3x QPS, 0.5% relative LogLoss improvement) are based on Table 2 and deployment paragraphs that supply no absolute MFLOPs, no model sizes, no candidate-distribution details, no measurement variance, and no explicit definition of 'on par.' These numbers are not externally checkable. If the journal accepts industrial evidence at this level of abstraction, the public experiments must carry the full weight of the quality claim; the authors should state exactly how the production metrics are measured and, where possible, provide sanitized configuration files or a reproducibility script. Footnote 1 promises an open-source repository but provides no URL or commit hash.
minor comments (6)
  1. [Footnote 1] The open-source statement should include the actual repository URL and, ideally, a commit hash or version identifier.
  2. [Table 6] KKBox is missing #Users and #Items. Add the values or state explicitly that they are unavailable in the source.
  3. [§3.3] The scaling ratio r is described only informally. Define precisely how r maps to layer widths or hidden dimensions, and state whether it is global or per-module when varied.
  4. [Figure 4] The axis labels and tick values are hard to read. Provide readable axes and include the numeric range of r and RLL improvement.
  5. [§5.1.2] The ROCS-Base configuration is said to 'approximately match' the vanilla FLOPs at N=1, but Table 1 shows non-trivial deviations (e.g., DCNv2/KuaiVideo: 1.19 vs 1.49). State the tolerance and how the closest feasible configuration is chosen.
  6. [§5.1.1] Define 'macro-average AUC' explicitly, including how multiple feedback types (Click/Like/Follow) are aggregated and whether the aggregation is per-example or per-evaluation-set.

Circularity Check

0 steps flagged

No load-bearing circularity: Eq. (11) is a definitional invariant, RRR is a tuned empirical knob, and quality claims are measured against vanilla baselines.

full rationale

The central derivation in §3.1 is a functional-dependency construction: Eq. (1) defines GLM compliance as request-side output invariance to candidates, and Eq. (11) restates that invariant under the closure argument in §3.1.4. The paper claims Eq. (11) only as 'semantic validity of compute sharing,' not as a quality prediction, so no empirical target is being derived by construction. The paper explicitly concedes 'Directly ROCSifying a model may therefore reduce prediction quality' (§3.3), which shows it is not conflating compute reuse with accuracy. Quality recovery is delegated to RRR, whose ratio r is a tunable capacity-allocation hyperparameter ('We control this allocation using a request-side scaling ratio r' §3.3) and is swept in Figure 4 rather than derived from the invariant. The public and production quality numbers are empirical comparisons against Vanilla baselines, not outputs of the GLM/DCA equations. The scaling-law citations [40,45] include the authors' own Wukong [40], but this is only motivational for RRR and is not load-bearing: the reported AUC deltas and RLL improvements are measured, not implied by the cited scaling law. No equation is fitted to the reported metrics, and no fitted input is renamed as a prediction. The admitted limitation and the minor self-citation do not make the derivation circular; they mildly reduce and mildly raise the concern, respectively.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 0 invented entities

The central GLM/DCA derivations are self-contained functional-dependency constructions, not fitted predictions. The main free parameter is the RRR scaling ratio r, which is tuned per workload. The domain axioms—candidate multiplicity, scaling-law quality recovery, and FLOP-based cost accounting—are empirically grounded in cited work but not independently proven for ROCSified architectures.

free parameters (2)
  • Request-side scaling ratio r = not reported (tuned per workload)
    §3.3 RRR; controls how much saved computation is reinvested into the request-side pathway. Figure 4 sweeps r for two production workloads; quality improves monotonically but the value is selected by validation rather than derived.
  • Vanilla model FLOP budget = 10 MFLOPs per instance
    Appendix A.1.2; a hand-picked computational budget used to select hyperparameters for public benchmarks and to define the ROCS-Base and ROCS-Scaled matching procedure.
axioms (3)
  • domain assumption Recommendation inference evaluates one request against many candidates, with request-side features identical across candidates.
    Opening of §1 and Figure 1. ROCS amortization is only valuable when candidate multiplicity is high; the paper itself notes in §6 that benefits diminish in low-candidate settings.
  • domain assumption Prediction quality improves with model compute following scaling laws, so RRR can recover quality by scaling the request-side pathway.
    §3.3 and Figure 4. This is the load-bearing premise for quality recovery after ROCSification; it is not proved for ROCSified models and is supported by scaling-law references that are partly co-authored by the same team.
  • domain assumption FLOPs are a valid proxy for inference cost in the public benchmark comparisons.
    §5.1.1 defines amortized FLOPs and uses them to compare quality-efficiency frontiers. Production QPS measurements partially validate this, but the public results do not report latency or energy.

pith-pipeline@v1.3.0-daily-deepseek · 17596 in / 11551 out tokens · 117334 ms · 2026-08-01T02:05:06.458625+00:00 · methodology

0 comments
read the original abstract

Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale. In this work, we propose Request-Oriented Compute Sharing (ROCS), a modeling and inference paradigm that exploits a unique property of recommendation inference: each user request is evaluated against many candidates, while request-side features are shared across candidates. ROCS defers request-candidate interactions as late as possible, isolates candidate-dependent representations, and evaluates substantial portions of the model once per request rather than once per candidate, significantly improving inference efficiency while maintaining or improving prediction quality. To realize this paradigm, we develop Generalized Layer Masking (GLM) to enforce candidate isolation in feature-interaction architectures, and Deep Cross Attention (DCA) to extend request-oriented sharing to sequence architectures. To support efficient GPU deployment, we co-design In-Kernel Broadcast Optimization (IKBO) that significantly accelerates ROCS model execution. Experiments on public benchmarks show that ROCS consistently improves the quality-efficiency tradeoff across recommendation backbones. On production-scale workloads, ROCS achieves up to a 3x QPS improvement on retrieval models without quality degradation and a 0.5% relative LogLoss improvement with a 50% QPS gain on a short-form video ranking model. ROCS has been deployed across large-scale recommendation systems spanning ads and organic surfaces, retrieval and ranking stages, and more than two orders of magnitude in inference complexity, delivering significant online gains at reduced infrastructure cost.

Figures

Figures reproduced from arXiv: 2607.27744 by Ao Cai, Bin Yu, Birmingham Guan, Boda Li, Buyun Zhang, Chunqiang Tang, Ellie Dingqiao Wen, GP Musumeci, Han Liu, Haoyu Wang, Hongtao Yu, Jianbo Sun, Jianbo Xiao, Jian Jiao, Jiawei Li, Jiaxin Lu, Liang Luo, Longhao Jin, Neng Shi, Qianru Li, Ryan Dick, Santanu Kolay, Shuqi Yang, Shuyao Bi, Sihan Zeng, Sijia Chen, Tongyi Tang, Venkatesh Ranganathan, Wei Ling, Wenlin Chen, Wenyi Xie, Xiaohan Wei, Yang Chen, Yantao Yao, Yichen Ruan, Yinbin Ma, Yong Ler Lee, Yuanwei Fang, Yuchen Hao, Yuxin Chen, Zeliang Chen, Zhengkai Zhang, Zhengyu Zhang, Zhuoran Zhao, Zijian Li, Zijian Shen, Zikun Liu.

Figure 1
Figure 1. Figure 1: Overview of ROCS. (a) Conventional models broadcast request-side inputs and introduce request–candidate interactions early, causing substantial redundant computation across candidates. (b) ROCS maintains a reusable request-side pathway by ROCSifying operators to satisfy a dependency contract and eliminate redundant computation. (c) Deep Cross Attention (DCA) extends ROCS to sequence modeling through deep c… view at source ↗
Figure 2
Figure 2. Figure 2: Construction of GLM modules, illustrated with two-group scenario where group 0 is request-side and group 1 is candidate-side. (a) linear and linear compression. (b) group-local operators. (c) factorization machine. (d) GLM invariant is preserved under composition. 3.1.3 Factorization Machine. A factorization machine explicitly computes pairwise interactions between input embeddings. To satisfy the GLM inva… view at source ↗
Figure 3
Figure 3. Figure 3: (a) Dense LCB implementation. (b) IKBO LCB implementation. at batch size 𝐵𝑟 :  Y 𝑅 Z 𝑅→𝐶  =  𝑊𝑅𝑅 𝑊𝐶𝑅 X 𝑅 , where Z 𝑅→𝐶 = 𝑊𝐶𝑅X 𝑅 corresponds to the request-derived contri￾bution in Eq. 6. The candidate-side output is then Y 𝐶 𝑖 = Z 𝑅→𝐶 𝑚𝑖 +𝑊𝐶𝐶X 𝐶 𝑖 . (14) DCA admits an analogous decomposition. Let S 𝑅 denote the request-side sequence representation stored at batch size 𝐵𝑟 . The linear K/V projections ca… view at source ↗
Figure 4
Figure 4. Figure 4: RLL improvement increases as request-side computation scales. 5.2.4 Effectiveness of DCA [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

48 extracted references · 16 linked inside Pith

  1. [1]

    KKBox Music Recommendation Challenge

    2018. KKBox Music Recommendation Challenge. https://www.kaggle.com/c/ kkbox-music-recommendation-challenge

  2. [2]

    KuaiVideo

    2022. KuaiVideo. https://www.kuaishou.com/activity/uimc

  3. [3]

    Newsha Ardalani, Carole-Jean Wu, Zeliang Chen, Bhargav Bhushanam, and Adnan Aziz. 2022. Understanding scaling laws for recommendation models. arXiv preprint arXiv:2208.08489(2022)

  4. [4]

    Zheng Chai, Qin Ren, Xijun Xiao, Huizhi Yang, Bo Han, Sijun Zhang, Di Chen, Hui Lu, Wenlin Zhao, Lele Yu, et al . 2025. Longer: Scaling up long sequence modeling in industrial recommenders. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 247–256

  5. [5]

    Jianxin Chang, Chenbin Zhang, Zhiyi Fu, Xiaoxue Zang, Lin Guan, Jing Lu, Yiqun Hui, Dewei Leng, Yanan Niu, Yang Song, et al. 2023. TWIN: TWo-stage interest network for lifelong user behavior modeling in CTR prediction at kuaishou. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3785–3794

  6. [6]

    Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep neural networks for YouTube recommendations. InProceedings of the 10th ACM Conference on Recommender Systems. 191–198

  7. [7]

    Tri Dao. 2024. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. InInternational Conference on Learning Representa- tions, B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun (Eds.), Vol. 2024. 35549–35562. https://proceedings.iclr.cc/paper_files/paper/2024/file/ 98ed250b203d1ac6b24bbcf263e3d4a7-Paper-Conference.pdf

  8. [8]

    Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. arXiv:2205.14135 [cs.LG] https://arxiv.org/abs/2205.14135

  9. [9]

    Qin Ding, Kevin Course, Linjian Ma, Jianhui Sun, Ruochen Liu, Zhao Zhu, Chunx- ing Yin, Wei Li, Dai Li, Yu Shi, Xuan Cao, Ze Yang, Han Li, Xing Liu, Bi Xue, Hongwei Li, Rui Jian, Daisy Shi He, Jing Qian, Matt Ma, Qunshu Zhang, and Rui Li. 2026. Bending the Scaling Law Curve in Large-Scale Recommendation Systems. arXiv:2602.16986 [cs.IR] https://arxiv.org/...

  10. [10]

    Chongming Gao, Shijun Li, Yuan Zhang, Jiawei Chen, Biao Li, Wenqiang Lei, Peng Jiang, and Xiangnan He. 2022. Kuairand: An unbiased sequential recom- mendation dataset with randomly exposed videos. InProceedings of the 31st ACM international conference on information & knowledge management. 3953–3957

  11. [11]

    Xin Gao, Sibasish Acharya, Sihui Han, Yongxiong Ren, Yanli Zhao, Liang Luo, Chucheng Wang, Pradeep Fernando, Saurabh Mishra, Siqi Yan, et al. 2025. DECK: Experiences on Delta Checkpointing for Industrial Recommendation Systems. Proceedings of the VLDB Endowment18, 12 (2025), 4978–4990

  12. [12]

    Yue Guan, Hongtao Yu, Peng Chen, Daohang Shi, Karthik Manivannan, Nicholas J Riasanovsky, Manman Ren, Lei Wang, Shane Nay, Partha Kanuparthy, Zaifeng Pan, Zhengding Hu, and Yufei Ding. 2026. TLX: Hardware-Native, Evolv- able MIMW GPU Compiler for Large-scale Production Environments. (2026). arXiv:2605.10905 [cs.AR] https://arxiv.org/abs/2605.10905

  13. [13]

    Liang Guo, Wei Li, Lucy Liao, Huihui Cheng, Rui Zhang, Yu Shi, Yueming Wang, Yanzun Huang, Keke Zhai, Pengchao Wang, Timothy Shi, Xuan Cao, Shengzhi Wang, Renqin Cai, Zhaojie Gong, Omkar Vichare, Rui Jian, Leon Gao, Shiyan Deng, Xingyu Liu, Xiong Zhang, Fu Li, Wenlei Xie, Bin Wen, Rui Li, Lu Fang, Xing Liu, and Jiaqi Zhai. 2025. Request-Only Optimization ...

  14. [14]

    Xingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen, Jie Jiang, and Mingsheng Long. 2023. On the Embedding Collapse when Scaling up Recommendation Models.arXiv preprint arXiv:2310.04400(2023)

  15. [15]

    Bojian Hou, Xiaolong Liu, Xiaoyi Liu, Jiaqi Xu, Yasmine Badr, Mengyue Hang, Sudhanshu Chanpuriya, Junqing Zhou, Yuhang Yang, Han Xu, et al. 2026. Kunlun: Establishing scaling laws for massive-scale recommendation systems through unified architecture design.arXiv preprint arXiv:2602.10016(2026)

  16. [16]

    Xu Huang, Hao Zhang, Zhifang Fan, Yunwen Huang, Zhuoxing Wei, Zheng Chai, Jinan Ni, Yuchao Zheng, and Qiwei Chen. 2026. MixFormer: Co-Scaling Up Dense and Sequence in Industrial Recommenders. arXiv:2602.14110 [cs.IR] https://arxiv.org/abs/2602.14110

  17. [17]

    Yunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin, Xuanyuan Luo, Zhe Wang, Zheng Chai, Shikang Wu, Yuchao Zheng, and Jingjian Lin. 2026. HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction. arXiv:2601.12681 [cs.IR] https://arxiv.org/abs/2601.12681

  18. [18]

    Yuchen Jiang, Jie Zhu, Xintian Han, Hui Lu, Kunmin Bai, Mingyu Yang, Shikang Wu, Ruihao Zhang, Wenlin Zhao, Shipeng Bai, Sijin Zhou, Huizhi Yang, Tianyi Liu, Wenda Liu, Ziyan Gong, Haoran Ding, Zheng Chai, Deping Xie, Zhe Chen, Yuchao Zheng, and Peng Xu. 2026. TokenMixer-Large: Scaling Up Large Ranking Models in Industrial Recommenders. arXiv:2602.06563 [...

  19. [19]

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361(2020)

  20. [20]

    Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic opti- mization.arXiv preprint arXiv:1412.6980(2014)

  21. [21]

    Zikun Liu, Liang Luo, Qianru Li, Zhengyu Zhang, Wei Ling, Jingyi Shen, Zeliang Chen, Yaning Huang, Jingxian Huang, Abdallah Aboelela, et al. 2026. SOLARIS: Speculative Offloading of Latent-bAsed Representation for Inference Scaling. arXiv preprint arXiv:2604.12110(2026)

  22. [22]

    Hui Lu, Zheng Chai, Shipeng Bai, Hao Zhang, Zhifang Fan, Kunmin Bai, Ke Sun, Yingwen Wu, Bingzheng Wei, Xiang Sun, Ziyan Gong, Tianyi Liu, Hua Chen, Deping Xie, Zhongkai Chen, Zhiliang Guo, Qiwei Chen, and Yuchao Zheng

  23. [23]

    Liang Luo, Yuxin Chen, Zhengyu Zhang, Mengyue Hang, Andrew Gu, Buyun Zhang, Boyang Liu, Chen Chen, Fan Yang, Feifan Gu, et al. 2026. Meta lattice: Model space redesign for cost-effective industry-scale ads recommendations. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1. 2335–2346

  24. [24]

    Liang Luo, Yinbin Ma, Quanyu Zhu, Vasiliy Kuznetsov, Yuxin Chen, Neng Shi, Jian Jiao, Jiecao Yu, Buyun Zhang, Tongyi Tang, Xiaohan Wei, Yanli Zhao, Zeliang Chen, Yuchen Hao, Venkatesh Ranganathan, Sandeep Parab, Yantao Yao, Maxim Naumov, Chunzhi Yang, Shen Li, Ellie Wen, Wenlin Chen, San- tanu Kolay, and Chunqiang Tang. 2026. LoKA: Low-precision Kernel Ap...

  25. [25]

    Liang Luo, Buyun Zhang, Michael Tsang, Yinbin Ma, Yuxin Chen, Ching-Hsiang Chu, Yanli Zhao, Shen Li, Yuchen Hao, Guna Lakshminarayanan, et al . 2024. Disaggregated multi-tower: Topology-aware modeling technique for efficient large scale recommendation.Proceedings of Machine Learning and Systems6 (2024), 266–278

  26. [26]

    Kelong Mao, Jieming Zhu, Liangcai Su, Guohao Cai, Yuru Li, and Zhenhua Dong

  27. [27]

    Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole- Jean Wu, Alisson G Azzolini, et al. 2019. Deep learning recommendation model for personalization and recommendation systems.arXiv preprint arXiv:1906.00091 (2019)

  28. [28]

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library.Advances in neural information processing systems32 (2019)

  29. [29]

    Qi Pi, Guorui Zhou, Yujing Zhang, Zhe Wang, Lejian Ren, Ying Fan, Xiaoqiang Zhu, and Kun Gai. 2020. Search-based user interest modeling with lifelong sequential behavior data for click-through rate prediction. InProceedings of the 29th ACM International Conference on Information & Knowledge Management. 2685–2692

  30. [30]

    Ads Recommendation. 2025. External Large Foundation Model: How to Efficiently Serve Trillions of Parameters for Online Ads Recommendation.arXiv preprint arXiv:2502.17494(2025)

  31. [31]

    Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. 2024. FlashAttention-3: fast and accurate attention with asynchrony and low-precision. InProceedings of the 38th International Conference on Neu- ral Information Processing Systems(Vancouver, BC, Canada)(NIPS ’24). Curran Associates Inc., Red Hook, NY, USA, Article 2193, 28 pages

  32. [32]

    Zihua Si, Lin Guan, ZhongXiang Sun, Xiaoxue Zang, Jing Lu, Yiqun Hui, Xingchao Cao, Zeyu Yang, Yichen Zheng, Dewei Leng, et al . 2024. Twin v2: Scaling ultra-long user behavior sequence modeling for enhanced ctr predic- tion at kuaishou. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management. 4890–4897

  33. [33]

    Philippe Tillet, H. T. Kung, and David Cox. 2019. Triton: an intermediate lan- guage and compiler for tiled neural network computations. InProceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Program- ming Languages(Phoenix, AZ, USA)(MAPL 2019). Association for Computing Machinery, New York, NY, USA, 10–19. doi:10.1145/3315508.3329973

  34. [34]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.Advances in neural information processing systems30 (2017)

  35. [35]

    Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed Chi. 2021. Dcn v2: Improved deep & cross network and practical lessons for web-scale learning to rank systems. InProceedings of the web conference 2021. 1785–1797

  36. [36]

    Bencheng Yan, Pengjie Wang, Kai Zhang, Feng Li, Hongbo Deng, Jian Xu, and Bo Zheng. 2022. Apg: Adaptive parameter generation network for click-through rate prediction.Advances in Neural Information Processing Systems35 (2022), 24740–24752

  37. [37]

    Ted Zadouri, Markus Hoehnerbach, Jay Shah, Timmy Liu, Vijay Thakkar, and Tri Dao. 2026. FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling. arXiv:2603.05451 [cs.CL] https://arxiv.org/abs/ 2603.05451

  38. [38]

    Zhichen Zeng, Xiaolong Liu, Mengyue Hang, Xiaoyi Liu, Qinghai Zhou, Chaofei Yang, Yiqun Liu, Yichen Ruan, Laming Chen, Yuxin Chen, et al. 2025. InterFormer: Effective Heterogeneous Interaction Learning for Click-Through Rate Prediction. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 6225–6233

  39. [39]

    Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhao- jie Gong, Fangda Gu, Michael He, Yinghai Lu, and Yu Shi. 2024. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations. arXiv:2402.17152 [cs.LG] https://arxiv.org/abs/2402.17152

  40. [40]

    Buyun Zhang, Liang Luo, Yuxin Chen, Jade Nie, Xi Liu, Shen Li, Yanli Zhao, Yuchen Hao, Yantao Yao, Ellie Dingqiao Wen, et al. 2024. Wukong: towards a scal- ing law for large-scale recommendation. InProceedings of the 41st International Conference on Machine Learning. 59421–59434

  41. [41]

    Buyun Zhang, Liang Luo, Xi Liu, Jay Li, Zeliang Chen, Weilin Zhang, Xiaohan Wei, Yuchen Hao, Michael Tsang, Wenjun Wang, et al. 2022. DHEN: A deep and hierarchical ensemble network for large-scale click-through rate prediction. arXiv preprint arXiv:2203.11014(2022)

  42. [42]

    Zhaoqi Zhang, Haolei Pei, Jun Guo, Tianyu Wang, Yufei Feng, Hui Sun, Shaowei Liu, and Aixin Sun. 2026. OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender. arXiv:2510.26104 [cs.IR] https://arxiv.org/abs/2510.26104

  43. [43]

    Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai. 2018. Deep interest network for click-through rate prediction. InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining. 1059–1068

  44. [44]

    Jieming Zhu, Quanyu Dai, Liangcai Su, Rong Ma, Jinyang Liu, Guohao Cai, Xi Xiao, and Rui Zhang. 2022. Bars: Towards open benchmarking for recommender systems. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2912–2923

  45. [45]

    Jie Zhu, Zhifang Fan, Xiaoxie Zhu, Yuchen Jiang, Hangyu Wang, Xintian Han, Haoran Ding, Xinmin Wang, Wenlin Zhao, Zhen Gong, et al. 2025. Rankmixer: Scaling up ranking models in industrial recommenders. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 6309–6316

  46. [46]

    Jieming Zhu, Jinyang Liu, Shuai Yang, Qi Zhang, and Xiuqiang He. 2021. Open benchmarking for click-through rate prediction. InProceedings of the 30th ACM international conference on information & knowledge management. 2759–2769. A Experiment Details A.1 Public Datasets A.1.1 Dataset Details.The meta-data of selected public datasets are shown in Table. 6. ...

  47. [2023]

    In Proceedings of the AAAI conference on artificial intelligence, Vol

    FinalMLP: an enhanced two-stream MLP model for CTR prediction. In Proceedings of the AAAI conference on artificial intelligence, Vol. 37. 4552–4560

  48. [2026]

    arXiv:2602.10455 [cs.IR] https://arxiv.org/abs/2602.10455 ROCS Team

    Compute Only Once: UG-Separation for Efficient Large Recommendation Models. arXiv:2602.10455 [cs.IR] https://arxiv.org/abs/2602.10455 ROCS Team