Pith. sign in

REVIEW 4 major objections 4 minor 41 references

The paper claims that a CPU-only local filter, LO-FAR, can rank 475 sparse ad features in about two CPU-hours while preserving downstream NE gains competitive with interaction-aware baselines across budgets of 100-400 retained features.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 09:05 UTC pith:ZTI2ABXO

load-bearing objection Plausible industrial feature-pruning system, but the headline 'competitive NE' claim rests on a figure with no numbers — send to review, condition acceptance on adding the missing numbers and details. the 4 major comments →

arxiv 2607.20873 v1 pith:ZTI2ABXO submitted 2026-07-23 cs.IR

LO-FAR: A Cost-Aware Local Filter for Sparse Feature Ranking in Industrial Ad Recommendation

classification cs.IR
keywords sparse ID-list featuresfeature rankingCTR predictionCVR predictionNormalized Entropyembedding table pruningCPU-only rankingfilter method feature selection
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

LO-FAR is a CPU-only, model-agnostic workflow for ranking sparse ID-list features in industrial ad recommendation. The paper tries to establish that a feature's stand-alone held-out predictive signal, measured by a lightweight local estimator on exploded per-ID rows, is enough to order features for first-stage pruning: retraining the downstream ranker on LO-FAR's top-100 to top-400 features preserves Normalized Entropy gains on CTR and CVR tasks comparable to shuffle-based importance and Binary Stochastic Neurons, while ranking completes in about two CPU-hours instead of multiple GPU-days. Because sparse embeddings dominate the model's parameter count, retaining 100-300 of 475 features deterministically removes roughly 40-75% of sparse-embedding storage. The contribution is operational: when ranking must be rerun frequently under compute and turnaround constraints, a cheap local filter is a viable production choice, with interaction-aware methods reserved for the harder near-boundary decisions on a pre-filtered pool.

Core claim

The central claim is that first-stage sparse feature ranking in industrial ad recommenders does not need interaction-aware, retraining-coupled procedures to be useful: a per-feature local estimator, fit on labeled rows exploded into one row per identifier, produces rankings whose downstream NE gains after retraining are competitive with shuffle-based importance and BSN at budgets of 100-400 out of 475 features on a production dataset of over one million logged interactions. LO-FAR ranks features by held-out log loss of the local estimator, parallelizes linearly over CPU workers, and completes in ~2 CPU-hours, versus multi-day GPU runs for the baselines. The paper also makes a systems argumen

What carries the argument

The carrying mechanism is feature explosion plus a local ID-level predictor. Each variable-length list of categorical IDs is unrolled into one row per identifier with the example's label replicated; a per-ID estimator is fit on the exploded training rows, using the empirical positive rate for identifiers occurring at least K times and a k-nearest-neighbor backoff over frequency-weighted token distance for rare IDs. Example-level scores are the mean of their ID-level scores, and features are ranked by held-out log loss of this local predictor. This per-feature independence yields O(p·n·ℓ·log(nℓ)) total complexity and embarrassingly parallel execution across CPU workers, eliminating the need t

Load-bearing premise

The whole method relies on the assumption that a feature's stand-alone held-out predictive power is a faithful guide to how much it will help the retrained downstream ranker once the chosen subset is in place; features whose value appears only through interactions or redundancy can be ranked too low and dropped.

What would settle it

Construct a synthetic sparse feature that is individually uninformative (near-random local log loss) but whose cross-product with an existing feature is strongly predictive, add it to the 475-feature pool, and compare the downstream NE of the retrained ranker using LO-FAR's top-B subset against a subset that deliberately includes the synthetic feature at the same budget; if the latter consistently beats LO-FAR's subset, the local-signal proxy is falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Feature pruning can be re-evaluated on a near-daily cadence on commodity CPUs, since ranking no longer contends for accelerator capacity.
  • Retaining 100-300 of 475 features removes roughly 40-75% of sparse embedding tables, directly reducing training memory, serving footprint, and retraining time.
  • At a budget of around 200 features, the retrained ranker preserves a large share of the full-pool NE gain while removing more than half of the sparse tables, giving operators a quality-cost frontier.
  • Because LO-FAR does not model interactions, its rankings diverge from BSN (Jaccard ~60% at N=300), yet still identify an informative subset; the paper argues ranking similarity is not the success criterion.
  • The workflow is positioned as a first-stage filter, so teams can stage it ahead of interaction-aware methods, applying the expensive selectors only to the reduced pool where their marginal cost is justified.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper's load-bearing proxy—stand-alone local signal is a sufficient proxy for downstream NE after retraining—is untested for interaction-only features; a natural extension is to inject a synthetic feature that is useless alone but strongly predictive in conjunction with an existing feature and measure whether LO-FAR mis-orders it at the studied budgets.
  • The ~60% Jaccard overlap between LO-FAR and BSN at N=300 suggests that multiple near-optimal subsets exist in this feature space; a testable extension is whether a hybrid that adds pairwise redundancy checks to LO-FAR's scores can capture BSN's near-boundary decisions at a fraction of its cost.
  • The recurring-cost decomposition C(m,B,R) implies a break-even calculus: even a modest per-run quality penalty can be outweighed by cheaper reruns when the refresh cadence R is high; computing this break-even on a given stack would help operators choose between LO-FAR and training-coupled selectors.
  • The explosion-plus-local-estimator recipe is not tied to ad domains; it may transfer to session-based recommendation or any high-cardinality categorical setting where list-valued features dominate and rankings must be refreshed often.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes LO-FAR, a CPU-only, model-agnostic workflow for ranking sparse ID-list features in industrial ad recommendation. LO-FAR scores each feature by fitting a lightweight ID-level predictor on the feature's exploded values and labels, then ranks features by held-out log loss. The authors claim that on a production dataset with 475 sparse features, LO-FAR completes ranking in about two CPU-hours and that the downstream Normalized Entropy (NE) gain after retraining with the top-100–400 features is competitive with shuffle-based importance, Binary Stochastic Neurons, and a coverage-based heuristic for both CTR and CVR tasks. They also claim a deterministic 40–75% reduction in sparse embedding storage at these budgets. The paper frames the contribution as a systems-level operating point: cheap, repeatable, first-stage filtering before more expensive interaction-aware methods are applied.

Significance. If the central claim is supported, the contribution is practically useful: a simple local filter that matches interaction-aware baselines in downstream NE while costing roughly two CPU-hours could be a viable first-stage prune in production settings with tight compute and turnaround constraints. The paper also contributes a useful framing of feature ranking as a recurring operational problem (Section 3.5) and is unusually explicit about its scope and limitations (Section 4.4). These are real strengths. However, the manuscript as provided does not make the headline experimental claim checkable: Figure 1 is described only qualitatively, no numeric ΔNE values, error bars, confidence intervals, or statistical tests are reported, and no artifact is released. This evidentiary gap is load-bearing because 'competitive' is a comparative magnitude claim. The storage-reduction percentage also needs reconciliation with the paper's own per-budget arithmetic.

major comments (4)
  1. [§6.1, Figure 1] The abstract's central claim — that LO-FAR 'preserves downstream NE gains ... competitive with shuffle-based importance, BSN, and a coverage-based heuristic across budgets of 100–400' — rests entirely on Figure 1, but the manuscript reports no numeric ΔNE values, no per-budget table, no error bars or confidence intervals, and no statistical tests. The phrase 'competitive' is a magnitude comparison; without the underlying numbers, the reader cannot verify whether differences are material or within noise. Relatedly, §4.3 states that mean, median, and max aggregation produce 'statistically indistinguishable downstream NE' but gives no supporting statistics. Please provide exact NE gains (with uncertainties) for all four budgets, both tasks, and all four methods, and specify the number of retraining runs used to assess variance.
  2. [§3.4 and Abstract] The reported storage reduction range is inconsistent with the paper's own arithmetic. Retaining B of 475 features removes 79%, 58%, 37%, and 16% of sparse embedding tables for B = 100, 200, 300, 400, respectively. The abstract and Section 3.4 report a '40–75%' reduction, and Section 5.5 says retaining 100–300 features removes 'approximately 40–75%' — but 37% is outside 40–75 and 79% is above it. Please reconcile the range or report per-budget reduction explicitly.
  3. [§4.1, stage 4] The rare-identifier backoff is a core component of the local estimator, but its mechanism is underspecified. For rare IDs, s_j(id) is said to back off to a k-nearest-neighbor estimate 'over the K closest training identifiers' where the neighborhood is induced by 'frequency-weighted token distance.' Neither the token distance nor the frequency weighting is defined. Since the backoff is the only mechanism that assigns scores to rare and unseen IDs, an undefined distance leaves the method incompletely specified and prevents replication. Please define the distance and the exact kNN procedure, or remove the dependency by using a simpler, fully specified fallback.
  4. [§5.2, Eq. (7)] The sign convention of ΔNE appears to conflict with the 'gain' terminology. Equation (7) defines ΔNE = (NE_method − NE_baseline)/NE_baseline. Since lower NE is better, a method that improves over the dense-only baseline will have a negative ΔNE by this definition. The text repeatedly refers to positive 'NE gains' and Figure 1 is described as showing 'NE gain,' which suggests the intended definition is (NE_baseline − NE_method)/NE_baseline or that the plotted quantity is the negative of Eq. (7). Please clarify the sign convention or correct Eq. (7). Without numeric values, the reader cannot determine which convention was used in the figure.
minor comments (4)
  1. [§4.2, Eq. (5)] The stated complexity O(p·n·ℓ·log(nℓ)) covers construction and lookup of the local estimator but does not account for the kNN backoff. If the backoff requires computing pairwise distances among frequent IDs, the cost may be higher. Please state the complexity of the backoff or justify its omission.
  2. [§4.4] The interaction-only limitation is acknowledged, but the paper does not test on this dataset whether low-ranked, interaction-only features would add downstream NE if retained. The claimed empirical competitiveness may hold regardless, but a short experiment (e.g., adding LO-FAR low-ranked but BSN/Shuffle high-ranked features and measuring NE after retraining) would directly address the assumption that stand-alone utility is a sufficient proxy at the studied budgets.
  3. [§5.4 and Table 2] The operational comparison (two CPU-hours vs. multi-day GPU runs; under $100 vs. over $4,000) is reported as observed in one environment, which is appropriately hedged. Still, the manuscript gives no hardware configuration, number of CPU workers, or measurement procedure, making these numbers difficult to interpret or reproduce. A short appendix with the environment description would help.
  4. [§6.2] The Jaccard overlap of ~60% between LO-FAR and BSN at N=300 is reported as a single number. It would be useful to show the overlap at all N values and budgets, and to report overlap with Shuffle as well, since the interpretability claim depends on the full curve.

Circularity Check

0 steps flagged

No circularity: LO-FAR is an empirically evaluated filter, not a derivation that reduces to its inputs.

full rationale

LO-FAR's ranking is produced from per-feature held-out log loss on exploded ID rows (Eqs. 3–4), and downstream quality is assessed by retraining the ranker on the selected subset (Eqs. 6–7). The ranking is not fitted to the downstream NE evaluation; the held-out split is reserved for end-to-end NE. The storage-reduction numbers (§3.4, §5.5) are arithmetic consequences of the budget, explicitly labeled 'deterministic,' not fitted predictions. The paper's own §4.4 flags the central limitation (interaction-only features may be missed) and states that pruning 'should be validated through end-to-end retraining,' which shows the authors treat the local proxy as an assumption to be tested, not as a definition of downstream quality. No load-bearing self-citation or author-imported uniqueness theorem appears; the relevant citations (e.g., Fan & Lv [4], Fisher et al. [5]) are independent prior work. The absence of numeric ΔNE values in §6.1 is an evidentiary weakness that makes the headline claim hard to verify, but unverifiability is not circularity: the claimed result is not equivalent to the method's inputs by construction. Therefore no circular step is exhibited.

Axiom & Free-Parameter Ledger

7 free parameters · 6 axioms · 0 invented entities

LO-FAR inherits the standard filter/screening paradigm (SIS) and adds per-ID rate estimation with a kNN back-off. No new physical entities are introduced. The empirical claim rests on several domain assumptions and unstated hyperparameters, most notably the proxy assumption linking local log loss to downstream NE.

free parameters (7)
  • Sample size n = ≈10^6 rows
    Stage (1) chooses n≈10^6 'to balance statistical fidelity against turnaround time'; not derived.
  • Train/test split ratio = not stated
    Stage (2) partitions sampled rows but the ratio is not given; affects estimator stability.
  • Neighbor count K for rare-ID back-off = not stated
    Stage (4) uses K nearest identifiers; authors claim marginal effect but no actual K or sensitivity curve is reported.
  • Back-off distance metric = not defined
    'Frequency-weighted token distance' in §4.1 is never specified; the rare-ID estimator is not reproducible.
  • Aggregation function = mean
    Mean chosen after preliminary studies; median/max said indistinguishable but no evidence is shown.
  • Feature budgets B = 100, 200, 300, 400
    Practitioner-chosen evaluation points; not optimized or justified beyond covering aggressive to conservative pruning.
  • Short-list inclusion threshold = 99th-percentile list length < 5
    Dataset filter in §5.1; restricts scope to short lists and shapes the feature pool.
axioms (6)
  • domain assumption Stand-alone held-out predictive utility orders features by downstream NE contribution after retraining.
    Core screening assumption invoked in §4.1 stages 4–6 and acknowledged as a limitation in §4.4 for interaction-only features.
  • domain assumption Interaction-only features are dispensable at the evaluated budgets on this dataset.
    §4.4 admits LO-FAR can miss such features; the paper does not establish they are absent or unimportant here.
  • domain assumption The downstream ranker is a DLRM-family model with sparse embeddings >97% of parameters.
    §3.1; the storage-reduction estimates and evaluation setup depend on this architecture.
  • domain assumption Stratified sampling of n≈10^6 rows preserves the ranking produced on the full corpus.
    §4.1 stage 1 assumes the sample is representative; no stability analysis across samples is provided.
  • domain assumption Retained and removed features have approximately comparable per-table embedding size.
    §3.4 storage reduction ≈1−B/p uses this 'approximately satisfied' assumption.
  • ad hoc to paper The kNN back-off on 'frequency-weighted token distance' gives useful estimates for rare IDs.
    Introduced in §4.1 stage 4 without definition or validation; no independent evidence is given.

pith-pipeline@v1.3.0-alltime-deepseek · 14070 in / 11180 out tokens · 107581 ms · 2026-08-01T09:05:13.767225+00:00 · methodology

0 comments
read the original abstract

Industrial ad recommendation models rely heavily on sparse, high-cardinality ID-list features that encode user histories and contextual identifiers. Each is backed by a dedicated embedding table, so these features dominate storage, training, and serving cost and must be revisited as traffic and downstream models evolve. Therefore, sparse feature ranking is not just an offline modeling problem but also a recurrent systems decision limited by compute budgets and iteration cadence. We present Localized Feature Ranking (LO-FAR), a CPU-only, model-agnostic workflow that ranks each candidate feature from its stand-alone held-out predictive signal using lightweight local estimators rather than the GPU-bound retraining loops of permutation- and stochastic-gate-based methods. On a production dataset of more than one million logged interactions and 475 sparse ID-list features, LO-FAR completes ranking in approximately two CPU-hours and preserves downstream Normalized Entropy gains on CTR and CVR tasks that are competitive with shuffle-based importance, Binary Stochastic Neurons, and a coverage-based heuristic across budgets of 100--400 retained features. The contribution is a deployable workflow showing that, when cost and turnaround constraints are binding, a simple local filter can be a practical production choice over heavier interaction-aware alternatives.

Figures

Figures reproduced from arXiv: 2607.20873 by Egemen Erbayat, Luis Duque, Mohammad Amin, Sohini Roychowdhury, Srihari Reddy.

Figure 1
Figure 1. Figure 1: Downstream NE gain (ΔNE) relative to a dense-only reference model trained without sparse features. Each curve corresponds to a ranking method, and the horizontal axis sweeps the retained feature budget 𝐵. LO-FAR remains competitive with shuffle-based importance and BSN across realistic feature budgets while requiring substantially less ranking compute [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Jaccard similarity over the top-𝑁 retained features for each pair of ranking methods. Similarity is reported as a diagnostic only; end-to-end NE gain ( [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

41 extracted references · 2 canonical work pages

  1. [1]

    Bilge Acun, Matthew Murphy, Xiaodong Wang, Jade Nie, Carole-Jean Wu, and Kim Hazelwood. 2021. Understanding Training Efficiency of Deep Learning Recommendation Models at Scale. InIEEE International Symposium on High- Performance Computer Architecture (HPCA)

  2. [2]

    Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, Rohan Anil, Zakaria Haque, Lichan Hong, Vihan Jain, Xiaobing Liu, and Hemal Shah

  3. [3]

    Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep neural networks for YouTube recommendations. InProceedings of the 10th ACM Conference on Recommender Systems. ACM, 191–198

  4. [4]

    Jianqing Fan and Jinchi Lv. 2008. Sure Independence Screening for Ultrahigh Dimensional Feature Space.Journal of the Royal Statistical Society: Series B (Statistical Methodology)70, 5 (2008), 849–911

  5. [5]

    Aaron Fisher, Cynthia Rudin, and Francesca Dominici. 2019. All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously.Journal of Machine Learning Research20, 177 (2019), 1–81

  6. [6]

    Huan Gui, Ruoxi Wang, Ke Yin, Long Jin, Maciej Kula, Taibai Xu, Lichan Hong, and Ed H Chi. 2023. Hiformer: Heterogeneous feature interactions learning with transformers for recommender systems.arXiv preprint arXiv:2311.05884(2023)

  7. [7]

    Huifeng Guo, Wei Guo, Yong Gao, Ruiming Tang, Xiuqiang He, and Wenzhi Liu. 2021. ScaleFreeCTR: MixCache-based Distributed Training System for CTR Models with Huge Embedding Table. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1269–1278. doi:10.1145/3404835.3462976

  8. [8]

    Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: A Factorization-Machine based Neural Network for CTR Prediction. InProceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence (IJCAI-17). 1725–1731

  9. [9]

    Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, and Joaquin Quiñonero-Candela. 2014. Practical Lessons from Predicting Clicks on Ads at Facebook. InProceedings of the Eighth International Workshop on Data Mining for Online Advertising(New York, NY, USA)(ADKDD’14). Association for Comp...

  10. [10]

    Pengyue Jia, Yejing Wang, Zhaocheng Du, Xiangyu Zhao, Yichao Wang, Bo Chen, Wanyu Wang, Huifeng Guo, and Ruiming Tang. 2024. ERASE: Benchmarking Feature Selection Methods for Deep Recommender Systems. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Barcelona, Spain)(KDD ’24). Association for Computing Machinery, New...

  11. [11]

    Wang-Cheng Kang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi, Ting Chen, Lichan Hong, and Ed H. Chi. 2021. Learning to Embed Categorical Features without Embedding Tables for Recommendation. InProceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. 840–850. doi:10. 1145/3447548.3467304

  12. [12]

    Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recom- mendation. In2018 IEEE international conference on data mining (ICDM). IEEE, 197–206

  13. [13]

    Ron Kohavi and George H. John. 1997. Wrappers for Feature Subset Selection. Artificial Intelligence97, 1–2 (1997), 273–324

  14. [14]

    Arjun Krishnan et al. 2023. Mutual information-based filter hybrid feature selec- tion method for medical dataset classification.Multimedia Tools and Applications 82, 15 (2023), 22193–22216. doi:10.1007/s11042-023-15143-0

  15. [15]

    Shiwei Li, Huifeng Guo, Xing Tang, Ruiming Tang, Lu Hou, Ruixuan Li, and Rui Zhang. 2024. Embedding Compression in Recommender Systems: A Survey. Comput. Surveys56, 5 (2024), 130:1–130:21. doi:10.1145/3637841

  16. [16]

    Yu Li, Chiawei Chu, and Xingyu Liu. 2024. Improved DeepFM-Based Model with LSTM on Click-Through-Rate. InProceedings of the 5th International Conference on Computer Information and Big Data Applications(Wuhan, China)(CIBDA ’24). Association for Computing Machinery, New York, NY, USA, 52–56. doi:10.1145/ 3671151.3671161

  17. [17]

    Weilin Lin, Xiangyu Zhao, Yejing Wang, Tong Xu, and Xian Wu. 2022. AdaFS: Adaptive feature selection in deep recommender system. InProceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining. 3309–3317

  18. [18]

    Bin Liu, Chenxu Zhu, Guilin Li, Weinan Zhang, Jincai Lai, Ruiming Tang, Xi- uqiang He, Zhenguo Li, and Yong Yu. 2020. AutoFIS: Automatic Feature In- teraction Selection in Factorization Models for Click-Through Rate Prediction. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2636–2645

  19. [19]

    Siyi Liu, Chen Gao, Yihong Chen, Depeng Jin, and Yong Li. 2021. Learnable Em- bedding Sizes for Recommender Systems. InInternational Conference on Learning Representations (ICLR)

  20. [20]

    Zhuoran Liu, Leqi Zou, Xuan Zou, Caihua Wang, Biao Zhang, Da Tang, Bolin Zhu, Yijie Zhu, Peng Wu, Ke Wang, and Youlong Cheng. 2022. Monolith: Real Time Recommendation System With Collisionless Embedding Table.arXiv preprint arXiv:2209.07663(2022)

  21. [21]

    Christos Louizos, Max Welling, and Diederik P. Kingma. 2018. Learning Sparse Neural Networks through𝐿0 Regularization. InInternational Conference on Learn- ing Representations (ICLR)

  22. [22]

    Liang Luo, Yuxin Chen, Zhengyu Zhang, Mengyue Hang, Andrew Gu, Buyun Zhang, Boyang Liu, Chen Chen, et al. 2026. Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads Recommendations. InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’26). Association for Computing Machinery. arXiv preprint arXiv...

  23. [23]

    Fuyuan Lyu, Xing Tang, Dugang Liu, Liang Chen, Xiuqiang He, and Xue Liu

  24. [24]

    Fuyuan Lyu, Xing Tang, Hong Zhu, Huifeng Guo, Yingxue Zhang, Ruiming Tang, and Xue Liu. 2022. OptEmbed: Learning Optimal Embedding Table for Click- through Rate Prediction. InProceedings of the 31st ACM International Conference on Information & Knowledge Management. 1399–1409

  25. [25]

    Singh, Maxime Ransan, and Sagar Jain

    Wenhan Lyu, Devashish Tyagi, Yihang Yang, Ziwei Li, Ajay Somani, Karthikeyan Shanmugasundaram, Nikola Andrejevic, Ferdi Adeputra, Curtis Zeng, Arun K. Singh, Maxime Ransan, and Sagar Jain. 2025. DV365: Extremely Long User History Modeling at Instagram. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining

  26. [26]

    Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherni- avskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, Volodymyr Kon- dratenko, Stephanie Pereira, Xianjie Chen, Wenlin Chen, Vijay Rao,...

  27. [27]

    Hanchuan Peng, Fuhui Long, and Chris Ding. 2005. Feature selection based on mutual information: criteria of max-dependency, max-relevance, and min- redundancy.IEEE Transactions on pattern analysis and machine intelligence27, 8 (2005), 1226–1238

  28. [28]

    Matthew Richardson, Ewa Dominowska, and Robert Ragno. 2007. Predicting clicks: estimating the click-through rate for new ads. InProceedings of the 16th international conference on World Wide Web. 521–530

  29. [29]

    Geet Sethi, Bilge Acun, Niket Agarwal, Christos Kober, Carole-Jean Wu, and Kim Hazelwood. 2022. RecShard: Statistical Feature-Based Memory Optimization for Industry-Scale Neural Recommendation. InProceedings of Machine Learning and Systems (MLSys)

  30. [30]

    Ruoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed H. Chi. 2021. DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems. InProceedings of the Web Conference 2021. 1785–1797. doi:10.1145/3442381.3450078

  31. [31]

    Yejing Wang, Xiangyu Zhao, Tong Xu, and Xian Wu. 2022. AutoField: Automating Feature Selection in Deep Recommender Systems. InProceedings of the ACM Web Conference 2022. Association for Computing Machinery, New York, NY, USA, 1977–1986. doi:10.1145/3485447.3512071

  32. [32]

    Zhikai Wang, Kangyi Lin, Yanyan Shen, and Zibin Zhang. 2023. Feature Staleness Aware Incremental Learning for CTR Prediction. InProceedings of the Thirty- Second International Joint Conference on Artificial Intelligence (IJCAI). 2348–2356. doi:10.24963/ijcai.2023/261

  33. [33]

    Zhiqiang Xu, Dong Li, Weijie Zhao, Xing Shen, Tianbo Huang, Xiaoyun Li, and Ping Li. 2021. Agile and Accurate CTR Prediction Model Training for Massive- Scale Online Advertising Systems. InProceedings of the 2021 International Confer- ence on Management of Data. 2404–2409. doi:10.1145/3448016.3457236

  34. [34]

    Bing Xue, Mengjie Zhang, Will N Browne, and Xin Yao. 2015. A survey on evolutionary computation approaches to feature selection.IEEE Transactions on evolutionary computation20, 4 (2015), 606–626

  35. [35]

    Yutaro Yamada, Ofir Lindenbaum, Sahand Negahban, and Yuval Kluger. 2020. Feature selection using stochastic gates. InInternational conference on machine learning. PMLR, 10648–10659. RecSys ’26, September 27-October 02, 2026, Minneapolis, MN, USA Erbayat et al

  36. [36]

    Xuanhua Yang, Xiaoyu Peng, Penghui Wei, Shaoguo Liu, Liang Wang, and Bo Zheng. 2022. Adasparse: Learning adaptively sparse structures for multi-domain click-through rate prediction. InProceedings of the 31st ACM International Con- ference on Information & Knowledge Management. 4635–4639

  37. [37]

    Hang Yin, Kuang-Hung Liu, Mengying Sun, Yuxin Chen, Buyun Zhang, Jiang Liu, Vivek Sehgal, Rudresh Rajnikant Panchal, Eugen Hotaj, Xi Liu, et al. 2024. AutoML for Large Capacity Modeling of Meta’s Ranking Systems. InCompanion Proceedings of the ACM Web Conference 2024. 374–382

  38. [38]

    Junqi Zhang, Yiqun Liu, Jiaxin Mao, Xiaohui Xie, Min Zhang, Shaoping Ma, and Qi Tian. 2022. Global or Local: Constructing Personalized Click Models for Web Search. InProceedings of the ACM Web Conference 2022(Virtual Event, Lyon, France)(WWW ’22). Association for Computing Machinery, New York, NY, USA, 213–223. doi:10.1145/3485447.3511950

  39. [39]

    Xiangyu Zhao, Maolin Wang, Xinjian Zhao, Jiansheng Li, Shucheng Zhou, Dawei Yin, Qing Li, Jiliang Tang, and Ruocheng Guo. 2023. Embedding in Recom- mender Systems: A Survey.arXiv preprint arXiv:2310.18608(2023). https: //arxiv.org/html/2310.18608v2 Comprehensive survey on embedding techniques for recommendation systems, covering high-dimensional sparse fe...

  40. [2016]

    InProceedings of the 1st Workshop on Deep Learning for Recommender Systems(Boston, MA, USA) (DLRS 2016)

    Wide & Deep Learning for Recommender Systems. InProceedings of the 1st Workshop on Deep Learning for Recommender Systems(Boston, MA, USA) (DLRS 2016). Association for Computing Machinery, New York, NY, USA, 7–10. doi:10.1145/2988450.2988454

  41. [2023]

    InProceedings of the ACM Web Conference 2023

    Optimizing Feature Set for Click-Through Rate Prediction. InProceedings of the ACM Web Conference 2023. 3386–3395