Pith. sign in

REVIEW 4 major objections 4 minor 2 cited by

InterFormer: Effective Heterogeneous Interaction Learning for Click-Through Rate Prediction

T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read InterFormer claims that bidirectional interleaving of static and sequence features improves click-through rate prediction, with up to 0.14% AUC gain on benchmarks and 0.15% NE gain in an industrial deployment.

desk verdict A sensible compositional CTR architecture with a clean ablation story, but the headline SOTA margins come from single runs and one table contradicts the text's gAUC claim. read the letter →

arxiv 2411.09852 v4 pith:RJOUTXND submitted 2024-11-15 cs.IR cs.AIcs.LG

classification cs.IRcs.AIcs.LG
keywords CTRpredictionheterogeneousinformationbidirectionalinteractionsequencemodelingfeatureinterleavingarchitectureselectiveaggregationindustrialrecommendation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that click-through rate prediction is held back by two design choices in existing models: information flows only one way, from static features to behavior sequences, and sequence information is aggressively summarized early on. InterFormer is proposed as a module that interleaves non-sequence and sequence processing across stacked layers, letting each mode inform the other, while a separate Cross Arch selects and compresses information before exchange. The authors report state-of-the-art results on three public benchmarks, with up to 0.14% AUC improvement over the strongest baseline, and an industrial-scale evaluation on a 70-billion-sample dataset showing a 0.15% normalized-entropy gain and 24% higher query throughput. If these margins hold, the design offers a reusable building block for large-scale advertising and recommendation systems.

What carries the argument

The load-bearing object is the InterFormer block, defined by three cooperating arches. The Interaction Arch takes non-sequence features plus sequence summarization as input and returns behavior-aware non-sequence embeddings; the Sequence Arch combines a Personalized FeedForward Network, which projects sequence embeddings using non-sequence summarization as a query, with multi-head attention and a prepended CLS token; the Cross Arch gates and summarizes both modes before exchange. Because each arch preserves the shape of its input, the paper can stack multiple InterFormer layers without aggressive pooling, and the same block is compatible with different interaction backbones such as dot product, DCNv2, and DHEN.

What would settle it

Re-run the benchmark comparisons with multiple random seeds (for example, ten) on one public dataset such as Amazon-Electronics and compute confidence intervals for the AUC difference between InterFormer and the strongest baseline; if the interval straddles zero, the SOTA claim collapses. For the industrial claim, an A/B test that fails to reproduce the 0.15% normalized-entropy gain or the 0.6% topline improvement at the same scale would falsify it.

Watch

Extended reading notes

Core claim

The central claim is that heterogeneous information in CTR prediction is best integrated by bidirectional, interleaved interaction rather than by unidirectional conditioning or early concatenation. InterFormer alternates an Interaction Arch, which models feature interactions among non-sequence embeddings enriched with summarized behavior, and a Sequence Arch, which models behavior sequences with multi-head attention conditioned on summarized non-sequence context. The Cross Arch keeps each mode's full representation intact while extracting gated low-dimensional summaries—CLS tokens, pooling-by-multihead-attention tokens, and recent items for sequences; gated MLP compression for non-sequences—so that neither mode is pooled prematurely. The authors assert that this design yields mutually beneficial learning, and they support it with ablations in which bidirectional flow consistently beats both unidirectional directions and selective aggregation beats average pooling, MLP, and multi-head-attention early summarization.

Load-bearing premise

The load-bearing premise is that the reported differences—roughly 0.0005 to 0.001 in AUC and 0.15% in normalized entropy—reflect genuine improvement rather than run-to-run variance, because the key benchmark and industrial tables report single runs without error bars or significance tests.

Editorial extensions

If this is right

  • InterFormer is claimed to beat 11 state-of-the-art baselines on Amazon-Electronics, TaobaoAds, and KuaiVideo, with gains up to 0.9% in gAUC, 0.14% in AUC, and 0.54% in LogLoss.
  • A three-layer InterFormer is claimed to improve normalized entropy by 0.15% over the internal state-of-the-art model at similar FLOPs, while delivering a 24% queries-per-second gain from overlapping communication and computation.
  • The bidirectional interleaving style is claimed to be universally beneficial across dot-product, DCNv2, and DHEN backbones, with performance ordered from weakest to strongest across scenarios: no exchange, separate arches, single-direction flows, and full bidirectional flow.
  • The results claim that selective aggregation improves CTR quality compared to aggressive early summarization by average pooling, MLP, or multi-head attention.
  • Pilot launches of the model in the paper's industrial advertising system are claimed to have produced a 0.6% improvement in topline metrics.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the core claim is to run the public-benchmark comparisons across multiple seeds and report confidence intervals; this would separate the small AUC and NE margins from training noise.
  • The separate-summarization principle could transfer to other multi-modal ranking tasks where token counts differ, since the Cross Arch decouples selection from interaction.
  • If the 24% query-throughput gain is robust, it suggests that co-scheduling communication-bound interaction modules and computation-bound sequence modules could speed up other large-scale ranking workloads.
  • The paper's observation that adding long sequences improves NE by 0.14% hints that interleaving may increase the benefit of sequence feature scaling, a hypothesis testable on public long-sequence datasets.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes InterFormer, a module for click-through rate prediction that learns heterogeneous information interactions between non-sequence features and user behavior sequences in an interleaving style. The architecture combines an Interaction Arch for behavior-aware non-sequence feature interactions, a Sequence Arch with personalized feed-forward networks and multi-head attention for context-aware sequence modeling, and a Cross Arch that selectively summarizes gated information for exchange between the two modes. The authors report state-of-the-art AUC, gAUC, LogLoss, and NE results on three public benchmarks and an internal Meta dataset, an industrial deployment with 0.15% NE gain and 24% QPS gain, and ablations supporting the bidirectional interleaving design.

Significance. If the empirical claims hold, InterFormer offers a useful and broadly applicable architectural pattern for CTR prediction, extending earlier unidirectional sequence-to-nonsequence designs with a bidirectional, interleaved information flow and a separate summarization arch. The architecture equations are internally coherent, the method is compatible with several interaction backbones, and the public-benchmark experiments are conducted within the BARS framework, which is a reproducible setup. The ablations in Figure 2 show a consistent ordering best-to-worst of int, n2s approximately s2n, sep, sole across three backbones, which directionally supports the central mechanism. However, the headline state-of-the-art claims rest on very small differences from single runs, and one stated gAUC claim is contradicted by the paper's own table, so the empirical support currently lags behind the strength of the claims.

major comments (4)
  1. [§5.2.1, Table 3] The text states that InterFormer outperforms the best competitor by up to 0.9% in gAUC, but Table 3 shows that InterFormer's gAUC on KuaiVideo is 0.6637, which is lower than DIEN's 0.6651. The largest gAUC gain in Table 3 is 0.0008 on AmazonElectronics (0.8843 vs. 0.8835), which is about 0.09% relative, not 0.9%. This internally inconsistent claim must be corrected, and the paper should either identify the dataset and metric on which each percentage gain is achieved or qualify that gAUC is not a primary metric.
  2. [§5.2.1, Table 3] The state-of-the-art claim is based on single-run benchmark scores with no error bars, confidence intervals, or significance tests. The decisive margins are very small: 0.0009 AUC on TaobaoAds, 0.0014 AUC on AmazonElectronics, and 0.0002 AUC on KuaiVideo, all of which could plausibly be within run-to-run variance for these datasets and model families. Without multiple seeds or a significance test, the central claim that InterFormer consistently achieves state-of-the-art AUC on public benchmarks is not yet established.
  3. [§5.2.2, Figures 2 and 3] The ablations that attribute performance gains to the interleaving learning style and to selective information aggregation are not parameter-matched. The int, n2s, s2n, sep, and sole scenarios differ in the number and size of modules, and the selective aggregation comparison in Figure 3 varies the aggregation mechanism without controlling for capacity. As a result, part of the observed improvement could be due to added model capacity rather than to the bidirectional interleaving or selective aggregation mechanism itself. Parameter-matched ablations, or equivalent-capacity baselines, are needed to support the mechanism-level interpretation.
  4. [§5.3.1 and §5.3.2] The industrial results are reported as a 0.15% NE gain and a 24% QPS gain against an unnamed internal SOTA model, with no confidence intervals or run-to-run variability information. Additionally, the QPS gain appears to include the model-system co-design optimizations described in Section 5.3.2 (communication overlap, FLOP reallocation, and kernel fusion), so the architecture's own efficiency contribution is not separated from systems engineering. The paper should identify the baseline model, state how many runs or evaluation periods were used, and disentangle architectural from systems-level gains.
minor comments (4)
  1. [§5.1] The sentence beginning 'public benchmark datasets are carried out in Section 5.2' is ungrammatical; it should be revised to 'public benchmark experiments are carried out' or similar.
  2. [Appendix B.1] The AmazonElectronics dataset description says there are 1,689,188 samples, but the subsequent split mentions 2.60M training and 0.38M test samples, which total about 2.98M and are inconsistent with the stated sample count. This should be corrected or clarified.
  3. [Figure 1 and Algorithm 1] The figure caption says the CLS token appending happens only at the first layer, and step 3 of Algorithm 1 similarly performs the prepending before the main loop, but the main loop in Algorithm 1 does not explicitly mark this one-time step. Making the one-time nature explicit in the algorithm pseudocode would improve clarity.
  4. [§2] The related work section cites many papers on graph learning and time-series forecasting that are not directly relevant to CTR prediction or heterogeneous interaction learning; tightening the citation list to the most relevant prior work would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: InterFormer's contribution is architectural and its SOTA claims are externally benchmarked, not derived from fitted parameters or self-referential definitions.

full rationale

InterFormer does not present a formal derivation whose output is assumed in its input. The architecture in Eqs. (7) through (11) composes standard, external building blocks (MHA, PMA, gating, PFFN, DCNv2/DHEN-style interaction modules), and none of these is defined in terms of the reported metrics AUC, gAUC, LogLoss, or NE. The performance claims in Sections 5.2 and 5.3 are measurements against public BARS benchmarks and an internal deployment baseline, not re-statements of fitted coefficients; no parameter is fitted to a subset and then reported as a prediction of a closely related quantity. The citations to DHEN [79] and Wukong [78] are legitimate prior-work backbones and baselines, and DHEN is additionally used as a baseline so InterFormer is evaluated against it rather than being equivalent to it. The unpublished M3C citation [31] supports deployment context but is not the load-bearing derivation of the central SOTA claim. Potential concerns about sub-0.001 AUC margins without error bars, the KuaiVideo gAUC deficit in Table 3, and uncontrolled parameter counts in ablations are statistical and correctness risks, not circularity under the definitions used here.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

No physical entities are invented. The new components, Cross Arch, PFFN, and gating, are software modules whose behavior is specified by equations and whose hyperparameters are tunable. The central empirical claim depends on the architecture choices and on the listed assumptions. There is no parameter-free derivation; the method is validated empirically.

free parameters (6)
  • Layer count L = 3 for internal; 1 to 4 searched on benchmarks
    Depth of the stacked Interaction, Sequence, and Cross Arch blocks; controlled in the scaling study, not derived.
  • Number of CLS tokens = 4
    Set in Appendix B.2; determines how much non-sequence context is prepended to the sequence.
  • Number of PMA tokens = 2
    Set in Appendix B.2; controls the learnable-query summarization of the behavior sequence.
  • Number of recent tokens = 2
    Set in Appendix B.2; controls the strength of the recency signal in the Cross Arch.
  • Learning rate = tuned in {1e-1, 1e-2, 1e-3}
    Appendix B reports a tuning range but not the per-dataset final values.
  • Attention and classifier MLP sizes and number of heads = MLP sizes in {512,256,128,64} or {1024,512,256,128}; heads in {1,2,4,8}
    Taken from BARS best configurations or defaults; not all per-dataset values are disclosed.
assumptions (5)
  • domain assumption Click probability is fully determined by user, item, and historical sequence features (Eq. 1).
    The model optimizes P(y=1 | u, i, S; theta) and ignores unobserved causes; standard in CTR modeling but an assumption.
  • standard math MHA, PMA, MaskNet, DCNv2, and DHEN are valid building blocks with the properties stated in Section 3.
    The architecture is assembled from these prior modules; their correctness is imported from the cited literature.
  • domain assumption Ablation scenarios differ only in the enabled information flows.
    The ordering sole, sep, n2s, s2n, int is interpreted as causal evidence for bidirectional flow, but equal capacity and compute across variants are not demonstrated.
  • domain assumption Offline NE gains extrapolate to online topline metrics.
    Section 5.3.3 reports 0.6% topline gains from pilot launches without deriving the mapping from offline NE to topline.
  • domain assumption Baseline hyperparameters from BARS are competitive and fairly tuned.
    Appendix B.2 states that best searched BARS configurations are used; the SOTA comparisons depend on this fairness assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of InterFormer: Effective Heterogeneous Interaction Learning for Click-Through Rate Prediction." pith.science (2026). https://pith.science/paper/RJOUTXND

@misc{pith2026241109852,
  author       = {Pith},
  title        = {Pith review of: InterFormer: Effective Heterogeneous Interaction Learning for Click-Through Rate Prediction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RJOUTXND}},
  note         = {Machine review of arXiv:2411.09852}
}
read the original abstract

Click-through rate (CTR) prediction, which predicts the probability of a user clicking an ad, is a fundamental task in recommender systems. The emergence of heterogeneous information, such as user profile and behavior sequences, depicts user interests from different aspects. A mutually beneficial integration of heterogeneous information is the cornerstone towards the success of CTR prediction. However, most of the existing methods suffer from two fundamental limitations, including (1) insufficient inter-mode interaction due to the unidirectional information flow between modes, and (2) aggressive information aggregation caused by early summarization, resulting in excessive information loss. To address the above limitations, we propose a novel module named InterFormer to learn heterogeneous information interaction in an interleaving style. To achieve better interaction learning, InterFormer enables bidirectional information flow for mutually beneficial learning across different modes. To avoid aggressive information aggregation, we retain complete information in each data mode and use a separate bridging arch for effective information selection and summarization. Our proposed InterFormer achieves state-of-the-art performance on three public datasets and a large-scale industrial dataset.

Figures

Figures reproduced from arXiv: 2411.09852 by the authors.

Figure 1
Figure 1. An overview of the InterFormer model architecture. Each block consists of three parts, including (1) Interaction Arch (orange) to learn behavior-aware non-sequence embeddings given sequence queries; (2) Sequence Arch (blue) to learn context-aware sequence embeddings given non-sequence queries; (3) Cross Arch (green) to exchange information between Interaction and Sequence Arch. Note that the dashed green line for CL… view at source ↗
Figure 2
Figure 2. Study on the interleaving learning style. We consider different non-sequence backbones (DOT, DCNv2, DHEN) and [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 4
Figure 4. Attention map on TaobaoAds. The first 4 tokens are [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: Ablation study on InterFormer, where ’-’ indicates ablating different information or modules. up to 0.003 AUC improvement. The sequence modeling modules (PFFN and MHA) also improves the performance to some extent. 5.3 Evaluation on Industrial Datasets To evaluate Inter…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design

    cs.IR 2026-02 reject novelty 6.0 of 10

    Kunlun reports a 2× scaling-efficiency gain and predictable power-law scaling for massive recommender systems, but the evidence is undermined by non-comparable cross-scale baselines and fitted slopes.

  2. A Collaborative Ensemble Framework for CTR Prediction

    cs.IR 2024-11 conditional novelty 4.0 of 10

    CETNet combines two CTR models with separate embeddings, KL collaboration, and entropy-based confidence fusion, yielding small AUC gains on public benchmarks but not in the internal deployment.

Reference graph

Works this paper leans on

91 extracted references · 31 canonical work pages · cited by 2 Pith papers

  1. [1]

    Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural ma- chine translation by jointly learning to align and translate.arXiv preprint arXiv:1409.0473(2014)

  2. [2]

    Wenxuan Bao, Zhichen Zeng, Zhining Liu, Hanghang Tong, and Jingrui He. 2024. Matcha: Mitigating Graph Structure Shifts with Test-Time Adaptation.arXiv preprint arXiv:2410.06976(2024)

  3. [3]

    Fedor Borisyuk, Mingzhou Zhou, Qingquan Song, Siyu Zhu, Birjodh Tiwana, Ganesh Parameswaran, Siddharth Dangi, Lars Hertel, Qiang Xiao, Xiaochen Hou, et al. 2024. LiRank: Industrial Large Scale Ranking Models at LinkedIn.arXiv preprint arXiv:2402.06859(2024)

  4. [4]

    Yekun Chai, Shuo Jin, and Xinwen Hou. 2020. Highway transformer: Self-gating enhanced self-attentive networks.arXiv preprint arXiv:2004.08178(2020)

  5. [5]

    Qiwei Chen, Huan Zhao, Wei Li, Pipei Huang, and Wenwu Ou. 2019. Behavior sequence transformer for e-commerce recommendation in alibaba. InProceedings of the 1st international workshop on deep learning practice for high-dimensional sparse data. 1–4

  6. [6]

    Xiaoshuang Chen, Gengrui Zhang, Yao Wang, Yulin Wu, Shuo Su, Kaiqiao Zhan, and Ben Wang. 2024. Cache-Aware Reinforcement Learning in Large-Scale Recommender Systems. InCompanion Proceedings of the ACM on Web Conference

  7. [7]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding.arXiv preprint arXiv:1810.04805(2018)

  8. [8]

    Ali Mamdouh Elkahky, Yang Song, and Xiaodong He. 2015. A multi-view deep learning approach for cross domain user modeling in recommendation systems. InProceedings of the 24th international conference on world wide web. 278–288

Show all 91 references
  1. [9]

    Yongqiang Han, Hao Wang, Kefan Wang, Likang Wu, Zhi Li, Wei Guo, Yong Liu, Defu Lian, and Enhong Chen. 2024. Efficient Noise-Decoupling for Multi-Behavior Sequential Recommendation. InProceedings of the ACM on Web Conference 2024. 3297–3306

  2. [10]

    Ruining He and Julian McAuley. 2016. Fusing similarity models with markov chains for sparse sequential recommendation. In2016 IEEE 16th international conference on data mining (ICDM). IEEE, 191–200

  3. [11]

    Ruining He and Julian McAuley. 2016. Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering. Inproceedings of the 25th international conference on world wide web. 507–517

  4. [12]

    Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, et al. 2014. Practical lessons from predicting clicks on ads at facebook. InProceedings of the eighth international workshop on data mining for online advert...

  5. [13]

    Po-Sen Huang, Xiaodong He, Jianfeng Gao, Li Deng, Alex Acero, and Larry Heck. 2013. Learning deep structured semantic models for web search using clickthrough data. InProceedings of the 22nd ACM international conference on Information & Knowledge Management. 2333–2338

  6. [14]

    Baoyu Jing, Yuchen Yan, Kaize Ding, Chanyoung Park, Yada Zhu, Huan Liu, and Hanghang Tong. 2024. Sterling: Synergistic representation learning on bipartite graphs. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 12976–12984

  7. [15]

    Baoyu Jing, Yuchen Yan, Yada Zhu, and Hanghang Tong. 2022. Coin: Co-cluster infomax for bipartite graphs.arXiv preprint arXiv:2206.00006(2022)

  8. [16]

    Diederik P Kingma. 2014. Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980(2014)

  9. [17]

    Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. 2019. Set transformer: A framework for attention-based permutation-invariant neural networks. InInternational conference on machine learning. PMLR, 3744–3753

  10. [18]

    Jinning Li, Ruipeng Han, Chenkai Sun, Dachun Sun, Ruijie Wang, Jingying Zeng, Yuchen Yan, Hanghang Tong, and Tarek Abdelzaher. 2024. Large language model- guided disentangled belief representation learning on polarized social graphs. In 2024 33rd International Conference on Co...

  11. [19]

    Jinning Li, Huajie Shao, Dachun Sun, Ruijie Wang, Yuchen Yan, Jinyang Li, Shengzhong Liu, Hanghang Tong, and Tarek Abdelzaher. 2022. Unsupervised belief representation learning with information-theoretic variational graph auto- encoders. InProceedings of the 45th International...

  12. [20]

    Yongqi Li, Meng Liu, Jianhua Yin, Chaoran Cui, Xin-Shun Xu, and Liqiang Nie

  13. [21]

    Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang, Zhongxia Chen, Xing Xie, and Guangzhong Sun. 2018. xdeepfm: Combining explicit and implicit feature in- teractions for recommender systems. InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data m...

  14. [22]

    Mingfu Liang, Xi Liu, Rong Jin, Boyang Liu, Qiuling Suo, Qinghai Zhou, Song Zhou, Laming Chen, Hua Zheng, Zhiyuan Li, Shali Jiang, Jiyan Yang, Xiaozhen Xia, Fan Yang, Yasmine Badr, Ellie Wen, Shuyu Xu, Hansey Chen, Zhengyu Zhang, Jade Nie, Chunzhi Yang, Zhichen Zeng, Weilin Zh...

  15. [23]

    Haokun Lin, Teng Wang, Yixiao Ge, Yuying Ge, Zhichao Lu, Ying Wei, Qingfu Zhang, Zhenan Sun, and Ying Shan. 2025. Toklip: Marry visual tokens to clip for multimodal comprehension and generation.arXiv preprint arXiv:2505.05422 (2025)

  16. [24]

    Haokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui, Yingtao Zhang, Linzhan Mou, Linqi Song, Zhenan Sun, and Ying Wei. 2024. Duquant: Distributing outliers via dual transformation makes stronger quantized llms.Advances in Neural Information Processing Systems37 (2024), 87766–87800

  17. [25]

    Haokun Lin, Haobo Xu, Yichen Wu, Ziyu Guo, Renrui Zhang, Zhichao Lu, Ying Wei, Qingfu Zhang, and Zhenan Sun. 2025. Quantization Meets dLLMs: A Sys- tematic Study of Post-training Quantization for Diffusion LLMs.arXiv preprint arXiv:2508.14896(2025)

  18. [26]

    Xiao Lin, Zhichen Zeng, Tianxin Wei, Zhining Liu, Hanghang Tong, et al. 2025. Cats: Mitigating correlation shift for multivariate time series classification.arXiv preprint arXiv:2504.04283(2025)

  19. [27]

    Xiaolong Liu, Zhichen Zeng, Xiaoyi Liu, Siyang Yuan, Weinan Song, Mengyue Hang, Yiqun Liu, Chaofei Yang, Donghyun Kim, Wen-Yen Chen, et al. 2024. A col- laborative ensemble framework for ctr prediction.arXiv preprint arXiv:2411.13700 (2024)

  20. [28]

    Zhining Liu, Ruizhong Qiu, Zhichen Zeng, Hyunsik Yoo, David Zhou, Zhe Xu, Yada Zhu, Kommy Weldemariam, Jingrui He, and Hanghang Tong. 2023. Class-imbalanced graph learning without class rebalancing.arXiv preprint arXiv:2308.14181(2023)

  21. [29]

    Zhining Liu, Ruizhong Qiu, Zhichen Zeng, Yada Zhu, Hendrik Hamann, and Hanghang Tong. 2024. AIM: Attributing, interpreting, mitigating data unfairness. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2014–2025

  22. [30]

    Zhining Liu, Ze Yang, Xiao Lin, Ruizhong Qiu, Tianxin Wei, Yada Zhu, Hendrik Hamann, Jingrui He, and Hanghang Tong. 2025. Breaking Silos: Adaptive Model Fusion Unlocks Better Time Series Forecasting.arXiv preprint arXiv:2505.18442 (2025)

  23. [31]

    Liang Luo, Mengyue Hang, Zhengyu Zhang, Andrew Gu, Buyun Zhang, Boyang Liu, Chen Chen, Fan Yang, Huayu Li, Jade Nie, et al. [n. d.]. M3C: a Multi-Domain Multi-Objective, Mixed-Modality Framework for Cost-Effective, Industry Scale Recommendation. ([n. d.])

  24. [32]

    Ze Lyu, Yu Dong, Chengfu Huo, and Weijun Ren. 2020. Deep match to rank model for personalized click-through rate prediction. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 156–163

  25. [33]

    Massimo Quadrana, Paolo Cremonesi, and Dietmar Jannach. 2018. Sequence- aware recommender systems.ACM computing surveys (CSUR)51, 4 (2018), 1–36

  26. [34]

    Prajit Ramachandran, Barret Zoph, and Quoc V Le. 2017. Searching for activation functions.arXiv preprint arXiv:1710.05941(2017)

  27. [35]

    Steffen Rendle. 2010. Factorization machines. In2010 IEEE International conference on data mining. IEEE, 995–1000

  28. [36]

    Guy Shani, David Heckerman, Ronen I Brafman, and Craig Boutilier. 2005. An MDP-based recommender system.Journal of machine Learning research6, 9 (2005)

  29. [37]

    Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. Autoint: Automatic feature interaction learning via self- attentive neural networks. InProceedings of the 28th ACM international conference on information and knowledge management....

  30. [38]

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing568 (2024), 127063

  31. [39]

    Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang

  32. [40]

    Yang Sun, Junwei Pan, Alex Zhang, and Aaron Flores. 2021. FM2: Field-matrixed factorization machines for recommender systems. InProceedings of the web conference 2021. 2828–2837

  33. [41]

    InProceedings of the 28th ACM international conference on information and knowledge management

    BERT4Rec: Sequential recommendation with bidirectional encoder rep- resentations from transformer. InProceedings of the 28th ACM international conference on information and knowledge management. 1441–1450

  34. [42]

    Tianchi. [n. d.]. Ad display/click data on taobao.com, 2018. https://tianchi.aliyun. com/dataset/56

  35. [43]

    Jiaxi Tang and Ke Wang. 2018. Personalized top-n sequential recommenda- tion via convolutional sequence embedding. InProceedings of the eleventh ACM international conference on web search and data mining. 565–573

  36. [44]

    Dingsu Wang, Yuchen Yan, Ruizhong Qiu, Yada Zhu, Kaiyu Guan, Andrew Margenot, and Hanghang Tong. 2023. Networked time series imputation via Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Zhichen Zeng et al. position-aware graph enhanced variational autoencoders. InPro...

  37. [45]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.Advances in neural information processing systems30 (2017)

  38. [46]

    Ruijie Wang, Baoyu Li, Yichen Lu, Dachun Sun, Jinning Li, Yuchen Yan, Shengzhong Liu, Hanghang Tong, and Tarek F Abdelzaher. 2023. Noisy positive- unlabeled learning with self-training for speculative knowledge graph reasoning. arXiv preprint arXiv:2306.07512(2023)

  39. [47]

    Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & cross network for ad click predictions. InProceedings of the ADKDD’17. 1–7

  40. [48]

    Ruijie Wang, Yuchen Yan, Jialu Wang, Yuting Jia, Ye Zhang, Weinan Zhang, and Xinbing Wang. 2018. Acekg: A large-scale knowledge graph for academic data mining. InProceedings of the 27th ACM international conference on information and knowledge management. 1487–1490

  41. [49]

    Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed Chi. 2021. Dcn v2: Improved deep & cross network and practical lessons for web-scale learning to rank systems. InProceedings of the web conference 2021. 1785–1797

  42. [50]

    Zhiqiang Wang, Qingyun She, and Junlin Zhang. 2021. Masknet: Introducing feature-wise multiplication to CTR ranking models by instance-guided mask. arXiv preprint arXiv:2102.07619(2021)

  43. [51]

    Shoujin Wang, Liang Hu, Yan Wang, Longbing Cao, Quan Z Sheng, and Mehmet Orgun. 2019. Sequential recommender systems: challenges, progress and prospects.arXiv preprint arXiv:2001.04830(2019)

  44. [52]

    Jun Xiao, Hao Ye, Xiangnan He, Hanwang Zhang, Fei Wu, and Tat-Seng Chua

  45. [53]

    Xue Xia, Pong Eksombatchai, Nikil Pancha, Dhruvil Deven Badani, Po-Wei Wang, Neng Gu, Saurabh Vishwas Joshi, Nazanin Farahpour, Zhiyuan Zhang, and An- drew Zhai. 2023. Transact: Transformer-based realtime user action model for recommendation at pinterest. InProceedings of the ...

  46. [54]

    Xin Xin, Bo Chen, Xiangnan He, Dong Wang, Yue Ding, and Joemon M Jose. 2019. CFM: Convolutional factorization machines for context-aware recommendation.. InIJCAI, Vol. 19. 3926–3932

  47. [55]

    Haobo Xu, Yuchen Yan, Dingsu Wang, Zhe Xu, Zhichen Zeng, Tarek F Abdelzaher, Jiawei Han, and Hanghang Tong. [n. d.]. Slog: An inductive spectral graph neural network beyond polynomial filter. InForty-first International Conference on Machine Learning

  48. [56]

    Zhibo Xiao, Luwei Yang, Wen Jiang, Yi Wei, Yi Hu, and Hao Wang. 2020. Deep multi-interest network for click-through rate prediction. InProceedings of the 29th ACM International Conference on Information & Knowledge Management. 2265–2268

  49. [57]

    Yuchen Yan, Yuzhong Chen, Huiyuan Chen, Xiaoting Li, Zhe Xu, Zhichen Zeng, Lihui Liu, Zhining Liu, and Hanghang Tong. 2024. Thegcn: Temporal heterophilic graph convolutional network.arXiv preprint arXiv:2412.16435(2024)

  50. [58]

    Yuchen Yan, Yuzhong Chen, Huiyuan Chen, Minghua Xu, Mahashweta Das, Hao Yang, and Hanghang Tong. 2023. From trainable negative depth to edge heterophily in graphs.Advances in Neural Information Processing Systems36 (2023), 70162–70178

  51. [59]

    Zhe Xu, Ruizhong Qiu, Yuzhong Chen, Huiyuan Chen, Xiran Fan, Menghai Pan, Zhichen Zeng, Mahashweta Das, and Hanghang Tong. 2024. Discrete-state continuous-time diffusion for graph generation.Advances in Neural Information Processing Systems37 (2024), 79704–79740

  52. [60]

    Yuchen Yan, Yongyi Hu, Qinghai Zhou, Lihui Liu, Zhichen Zeng, Yuzhong Chen, Menghai Pan, Huiyuan Chen, Mahashweta Das, and Hanghang Tong. 2024. Pacer: Network embedding from positional to structural. InProceedings of the ACM Web Conference 2024. 2485–2496

  53. [61]

    Yuchen Yan, Yongyi Hu, Qinghai Zhou, Shurang Wu, Dingsu Wang, and Hang- hang Tong. 2024. Topological anonymous walk embedding: A new structural node embedding approach. InProceedings of the 33rd ACM International Confer- ence on Information and Knowledge Management. 2796–2806

  54. [62]

    Yuchen Yan, Yuzhong Chen, Mahashweta Das, Hao Yang, and Hanghang Tong. [n. d.]. ReD-GCN: Revisit the Depth of Graph Convolutional Network. ([n. d.])

  55. [63]

    Yuchen Yan, Lihui Liu, Yikun Ban, Baoyu Jing, and Hanghang Tong. 2021. Dy- namic knowledge graph alignment. InProceedings of the AAAI conference on artificial intelligence, Vol. 35. 4564–4572

  56. [64]

    Yuchen Yan, Si Zhang, and Hanghang Tong. 2021. Bright: A bridging algorithm for network alignment. InProceedings of the web conference 2021. 3907–3917

  57. [65]

    Yuchen Yan, Baoyu Jing, Lihui Liu, Ruijie Wang, Jinning Li, Tarek Abdelzaher, and Hanghang Tong. 2023. Reconciling competing sampling strategies of network embedding.Advances in Neural Information Processing Systems36 (2023), 6844– 6861

  58. [66]

    Carl Yang, Lanxiao Bai, Chao Zhang, Quan Yuan, and Jiawei Han. 2017. Bridging collaborative filtering and semi-supervised learning: a neural approach for poi recommendation. InProceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining. 1245–1254

  59. [67]

    Xiaodong Yang, Huiyuan Chen, Yuchen Yan, Yuxin Tang, Yuying Zhao, Eric Xu, Yiwei Cai, and Hanghang Tong. 2024. SimCE: Simplifying Cross-Entropy Loss for Collaborative Filtering.arXiv preprint arXiv:2406.16170(2024)

  60. [68]

    Yuchen Yan, Qinghai Zhou, Jinning Li, Tarek Abdelzaher, and Hanghang Tong

  61. [69]

    Hyunsik Yoo, Zhichen Zeng, Jian Kang, Ruizhong Qiu, David Zhou, Zhining Liu, Fei Wang, Charlie Xu, Eunice Chan, and Hanghang Tong. 2024. Ensuring user-side fairness in dynamic recommender systems. InProceedings of the ACM Web Conference 2024. 3667–3678

  62. [70]

    Qi Yu, Zhichen Zeng, Yuchen Yan, Zhining Liu, Baoyu Jing, Ruizhong Qiu, Ariful Azad, and Hanghang Tong. 2025. PLANETALIGN: A Comprehensive Python Library for Benchmarking Network Alignment.arXiv preprint arXiv:2505.21366 (2025)

  63. [71]

    Qi Yu, Zhichen Zeng, Yuchen Yan, Lei Ying, R Srikant, and Hanghang Tong. 2025. Joint optimal transport and embedding for network alignment. InProceedings of the ACM on Web Conference 2025. 2064–2075

  64. [72]

    Yeongwook Yang, Hong-Jun Jang, and Byoungwook Kim. 2020. A hybrid rec- ommender system for sequential recommendation: combining similarity models with markov chains.IEEE Access8 (2020), 190136–190146

  65. [73]

    Zhichen Zeng, Ruizhong Qiu, Wenxuan Bao, Tianxin Wei, Xiao Lin, Yuchen Yan, Tarek F Abdelzaher, Jiawei Han, and Hanghang Tong. 2025. Pave Your Own Path: Graph Gradual Domain Adaptation on Fused Gromov-Wasserstein Geodesics. arXiv preprint arXiv:2505.12709(2025)

  66. [74]

    Zhichen Zeng, Ruizhong Qiu, Zhe Xu, Zhining Liu, Yuchen Yan, Tianxin Wei, Lei Ying, Jingrui He, and Hanghang Tong. [n. d.]. Graph mixup on approximate gromov–wasserstein geodesics. InForty-first International Conference on Machine Learning

  67. [75]

    Zhichen Zeng, Si Zhang, Yinglong Xia, and Hanghang Tong. 2023. Parrot: Position-aware regularized optimal transport for network alignment. InPro- ceedings of the ACM web conference 2023. 372–382

  68. [76]

    Zhichen Zeng, Boxin Du, Si Zhang, Yinglong Xia, Zhining Liu, and Hanghang Tong. 2024. Hierarchical multi-marginal optimal transport for network alignment. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 16660– 16668

  69. [77]

    Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhao- jie Gong, Fangda Gu, Michael He, et al. 2024. Actions speak louder than words: Trillion-parameter sequential transducers for generative recommendations.arXiv preprint arXiv:2402.17152(2024)

  70. [78]

    Buyun Zhang, Liang Luo, Yuxin Chen, Jade Nie, Xi Liu, Daifeng Guo, Yanli Zhao, Shen Li, Yuchen Hao, Yantao Yao, et al. 2024. Wukong: Towards a Scaling Law for Large-Scale Recommendation.arXiv preprint arXiv:2403.02545(2024)

  71. [79]

    Buyun Zhang, Liang Luo, Xi Liu, Jay Li, Zeliang Chen, Weilin Zhang, Xiaohan Wei, Yuchen Hao, Michael Tsang, Wenjun Wang, et al. 2022. DHEN: A deep and hierarchical ensemble network for large-scale click-through rate prediction. arXiv preprint arXiv:2203.11014(2022)

  72. [80]

    Zhichen Zeng, Ruike Zhu, Yinglong Xia, Hanqing Zeng, and Hanghang Tong

  73. [81]

    Yongfeng Zhang, Qingyao Ai, Xu Chen, and W Bruce Croft. 2017. Joint repre- sentation learning for top-n recommendation with heterogeneous information sources. InProceedings of the 2017 ACM on Conference on Information and Knowl- edge Management. 1449–1458

  74. [82]

    Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al. 2023. Pytorch fsdp: experiences on scaling fully sharded data parallel.arXiv preprint arXiv:2304.11277 (2023)

  75. [83]

    Guorui Zhou, Weijie Bian, Kailun Wu, Lejian Ren, Qi Pi, Yujing Zhang, Can Xiao, Xiang-Rong Sheng, Na Mou, Xinchen Luo, et al. 2020. CAN: revisiting feature co- action for click-through rate prediction.arXiv preprint arXiv:2011.05625(2020)

  76. [84]

    Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Deep interest evolution network for click-through rate prediction. InProceedings of the AAAI conference on artificial intelligence, Vol. 33. 5941–5948

  77. [85]

    Shuai Zhang, Lina Yao, Aixin Sun, and Yi Tay. 2019. Deep learning based recom- mender system: A survey and new perspectives.ACM computing surveys (CSUR) 52, 1 (2019), 1–38

  78. [86]

    Jieming Zhu, Quanyu Dai, Liangcai Su, Rong Ma, Jinyang Liu, Guohao Cai, Xi Xiao, and Rui Zhang. 2022. Bars: Towards open benchmarking for recommender systems. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. 291...

  79. [90]

    Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai. 2018. Deep interest network for click-through rate prediction. InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining. ...

  80. [2017]

    InProceedings of the 26th International Joint Conference on Artificial Intelligence

    Attentional factorization machines: learning the weight of feature in- teractions via attention networks. InProceedings of the 26th International Joint Conference on Artificial Intelligence. 3119–3125

  81. [2019]

    InProceedings of the 27th ACM international conference on multimedia

    Routing micro-videos via a temporal graph-guided recommendation system. InProceedings of the 27th ACM international conference on multimedia. 1464–1472

  82. [2022]

    InProceedings of the 31st ACM International Conference on Information & Knowledge Management

    Dissecting cross-layer dependency inference on multi-layered inter- dependent networks. InProceedings of the 31st ACM International Conference on Information & Knowledge Management. 2341–2351

  83. [2023]

    InInternational Conference on Ma- chine Learning

    Generative graph dictionary learning. InInternational Conference on Ma- chine Learning. PMLR, 40749–40769

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.