REVIEW 3 major objections 10 minor 41 references
Memory Layer: Train the In-Model Cache for Recommendation Models
T0 review · 3 major / 10 minor · reviewed 2026-07-31 · grok-4.5
Pith's one-line read Co-training the in-model item embedding cache with the ranking model removes the training–serving representation gap by construction.
desk verdict Solid industrial co-design: one writable in-model item cache plus always-on misses and RES, with real production wins—and a soft causal claim on how much of the NE-gap drop is representation sharing versus coverage. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The memory layer: a co-trained sparse key-value embedding table (item ID → item-tower embedding) with Writeback exact assignment, multi-table training to populate the full candidate pool, raw embedding streaming, and always-on embeddings for cache misses.
What would settle it
On a live early-ranking surface, measure training versus serving Normalized Entropy on the same window before and after replacing the external cache with the co-trained memory layer; if the gap does not shrink by a large fraction (paper reports up to 86% on pselect) while coverage stays below 100% or freshness stays at multi-minute scale, the central claim fails.
Extended reading notes
Core claim
The structural training–serving discrepancy in early-stage ranking is caused by an item embedding cache that exists only outside the training loop. Making that cache a co-trained in-model memory layer—written by the item tower in training, read at serving, and streamed as raw embeddings—creates one source of truth for item representations by construction, closes most of the measured NE gap, and makes full prediction coverage and near-real-time freshness properties of the system rather than add-on jobs.
Load-bearing premise
The system must run continuous online training that regularly cycles the full candidate pool through the item tower; without that, write-behind cache rows go stale and the claimed single-source freshness and gap reduction do not hold.
Editorial extensions
If this is right
- Early-stage ranking can own scoring for every candidate, including brand-new items, without fallback paths or separate bulk-evaluation clusters.
- Trainer-to-predictor updates collapse to one self-contained pipeline, cutting operational dependencies and full-snapshot override artifacts.
- Training-and-publish compute drops materially because embeddings are produced in the training loop rather than recomputed at publish.
- The same co-trained cache pattern can span early- to late-stage ranking and can store auxiliary item features beside embeddings.
- Foundation-model item towers become deployable under tight latency if they only write the cache in training and serving reads cached vectors only.
Reading between the lines
- Any multi-stage ranker that still keeps a serving-only item store is leaving a similar representation debt on the table, not only two-tower early rankers.
- Explicitly simulating cache-miss masks in training (noted as future work) is likely the next lever if residual NE gap must go to zero.
- Surfaces stuck on periodic batch retrain cannot buy the freshness claim without first adopting online training, so adoption is gated by training regime more than by model architecture.
- Treating the cache as a general model-colocated key-value store suggests the same write-during-train / read-at-serve pattern for user history or cross features, not only item towers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the memory layer, an in-model key-value item-embedding cache (an MPZCH sparse table) that is co-trained with the ranking model: during training, item-tower outputs are written into the table via "Writeback" (an exact assignment implemented as η=1 SGD through the existing TorchRec TBE backward path), and at serving the model reads the same table, refreshed at ~15 s cadence by Raw Embedding Streaming (RES). Multi-table training ingests the full candidate pool to populate the cache; always-on embeddings (author ID) score cache-missed items so every item receives a prediction. Deployed on Instagram Reels early-stage ranking since late 2025, the system reportedly raises prediction coverage from 96% to 100%, improves embedding freshness from O(5 min) to O(20 s) (P99 TTS ~21 s), narrows the training-serving NE gap by 86% on pselect (12.11%→1.64%) and 35% on reshare (5.24%→3.42%), yields >2× recall for sub-5-minute content and 5–6% cold-start engagement lifts, and cuts training-and-publish cost by 30% at neutral serving cost, with 2.42× serving throughput reinvested into +50% table capacity.
Significance. If the numbers hold, this is a substantial industrial-systems contribution. It identifies a real, under-measured failure mode (training–serving representation skew in cached early-stage ranking), offers a simple and verifiable mechanism (Writeback repurposes η=1 SGD as an exact assignment through the existing TBE backward path, avoiding custom kernels), and validates it with production A/B evidence at scale rather than offline proxies. The claims are empirical measurements against a strong internal baseline (SilverTorch), not constructions that reduce by definition, and the limitations (write-behind staleness, implicit cache-miss fusion, int8 residual) are disclosed with unusual candor. Key building blocks (MPZCH, Writeback, RES) are released via TorchRec/FBGEMM, giving the community reproducible components even though the end-to-end system is internal. The co-design principle — make the serving cache a trained component on a single update path — is likely to see broad adoption in industry ranking stacks.
major comments (3)
- [§6.4, Table 2 (also §1.2, §6.1, §6.2)] Table 2 / §6.4: the headline 86% gap reduction is a compound number and the serving-NE population is never defined. §1.2 lists cache-miss fallback as one of three discrepancy sources, and the baseline (§6.1) handles misses with a fixed negative score on ~4% of items (96%→100% coverage, §6.2), leaving them 'effectively unranked.' If baseline serving NE includes those fallback-scored items, the 12.11% conflates catastrophic miss-handling error with representation staleness, and the share attributable to the 'one source of truth' mechanism — the paper's central claim — is unquantified: always-on embeddings, an architecturally independent coverage fix, could account for much of the reduction. That pselect's baseline gap is 2.3× reshare's is consistent with task-dependent fallback exposure. §6.6 shows the authors can reason about scored-population effects for A/B metrics; the same analysis is
- [Abstract, §6.2, §7.1] The freshness claim 'O(5 min) → O(20 s)' measures time-to-serve — propagation latency for embeddings just written (P99 ~21 s, dominated by the 15 s RES interval). The staleness that matters at serving is per-item embedding age: M[i]=f_item(x_i;θ_t) read at t+k (§7.1), where k is bounded only by how fast multi-table training (§4.3) cycles the full candidate pool (hundreds of millions of items, §1.1). That cycle time is never reported, so for long-tail items the embedding age could be far larger than 20 s and the abstract invites the stronger reading. Please report the distribution (P50/P99) of time-since-last-Writeback for served items, or the implied full-pool cycle time from the pool-batch rate. This would also directly quantify residual-gap source (ii) in §6.4, currently unattributed in magnitude.
- [§3.2, §4.2, Algorithm 1, §3.4] The gradient flow is under-specified to the point of apparent inconsistency. Algorithm 1 (line 5) computes the loss on e_cached = M.lookup(I), and CacheLayer's backward returns the saved (cached − target) as the TBE gradient — i.e., it replaces the task gradient at the table. As written, no gradient path connects the ranking loss to the item tower, contradicting §4.2's claim that 'the training-loss gradient independently updates all model parameters θ (including the item tower).' Meanwhile §3.4 ('during evaluation steps the model reads ê_i = M[i] ... rather than the fresh item tower output') implies the training loss normally consumes the fresh tower output, contradicting Algorithm 1 line 5 and §3.2's 'the scoring network is trained on exactly what serving reads.' A coherent interpretation exists (CacheLayer forwards e_target, passes ∂L/∂output through to the tower, and substitutes the W
minor comments (10)
- [Table 1] 'All A/B improvements are statistically significant' is asserted without test statistics, confidence intervals, experiment duration, or traffic share; the '5–6%' and '6–7%' ranges are unexplained (across tasks? periods? variants?). Please add minimal detail.
- [Abstract] The abstract reports 'up to 86%' for the NE-gap reduction without the reshare figure (35%). Reporting both would avoid selection-on-max; contribution (1) already does this correctly.
- [§6.5, Table 3] The ablation is qualitative ('A/B metrics improved') and disclaimed as an early prototype not comparable to §6. Also note in Tables 1–2 captions that the treatment bundles +50% table capacity and always-on embeddings (§6.1), so single-number attributions are system-level.
- [§5.2] Clarify the 30% cost accounting boundary (does it net out added pool-batch item-tower compute from §4.3 and RES capture overhead?), and reconcile in one sentence how 2.42× serving throughput yields neutral serving cost (reinvested into +50% capacity — currently scattered across §6.1/A.6).
- [§3.2 notation] e_i vs ê_i vs E_item,i are easily confused, and 'ensemble embeddings' is used without definition. found_mask appears in §4.4 without a first-use definition. A small notation table would help.
- [Figure 3] State the denominator of cache hit rate (all requested items? all pool items?) and how baseline misses were handled in this specific measurement.
- [§6.3] 'These improvements translate to attributed topline impact' is garbled. Also give one line on what the A/B pretest ablation varied to support the attribution claim.
- [§3.3 vs §4.4] §3.3 lists author ID, topic ID, etc., but §4.4 says the deployed always-on set is author ID only (plus hit-gated media ID). Say so in §3.3, and briefly justify why the hit-gated per-item media-ID embedding does not reintroduce train/serve skew on misses.
- [§A.2] Dummy batches skip Writeback for pool items; please quantify the observed dummy-batch frequency (or state it is negligible), since it directly affects cache population rate.
- [§4, refs [20,21]] The open-sourced components (MPZCH, Writeback, RES in TorchRec/FBGEMM) would be easier to locate with a version/commit pointer and a sentence mapping paper components to released APIs.
Circularity Check
Empirical systems paper: NE-gap and A/B claims are measured against a baseline, not quantities forced equal to their inputs by construction.
full rationale
The paper's central claim is architectural and empirical: co-training an in-model item cache (Writeback + multi-table training + RES + always-on embeddings) removes the training-serving representation discrepancy and yields measured gains (coverage 96%→100%, TTS O(5 min)→O(20 s), pselect NE gap 12.11%→1.64%, cold-start A/B lifts, 30% cost cut). None of these quantities is defined to equal its inputs. Writeback is an exact η=1 SGD assignment of the item-tower output into M[i], not a fitted identity presented as a prediction. Cache-based evaluation deliberately reads M[i] so offline NE tracks serving NE; that is intentional alignment, not circular derivation. Self-citations (SilverTorch, MPZCH, TorchRec, SlimPer) supply infrastructure building blocks with independent purposes; none is a uniqueness theorem that forces the result. The skeptic's concern that the 86% NE-gap reduction is undecomposed (fallback scoring vs. representation staleness) is a causal-attribution / correctness issue, not circularity: the paper still reports an externally measured gap change rather than a quantity true by definition. Score 1 for ordinary non-load-bearing self-citation of prior infra; no circular steps.
Assumptions & free parameters
free parameters (4)
- RES streaming interval (~15 s; P99 TTS ~21 s) =
~15 s batch push; ~21 s P99 TTS
- MPZCH table capacity and 512-d embeddings (+50% vs baseline) =
512-d; +50% capacity vs baseline
- Always-on feature set (production: author ID; optional media-ID on hit) =
author ID (deployed)
- Multi-table train/pool batch-size ratio
assumptions (5)
- domain assumption Early-stage ranking must score O(1K) candidates under double-digit-ms latency, motivating precomputed item embeddings rather than full item-tower inference at request time.
- domain assumption The surface runs continuous online training so the item tower and memory layer refresh in near real time; batch-only retraining cannot achieve the claimed freshness.
- domain assumption A small set of item-agnostic features (e.g., author ID) yields a useful score when the per-item cache entry is missing or zeroed.
- standard math EXACT_SGD with η=1 implements exact row assignment M[i]←e_i via gradient g=M[i]−e_i without custom CUDA kernels.
- domain assumption MPZCH provides zero-collision ID→row mapping with bounded LRU capacity suitable as the shared train/serve table.
invented entities (4)
-
Memory layer (co-trained in-model KV item embedding cache)
-
CacheLayer Writeback
-
Raw Embedding Streaming (RES)
-
Always-on embeddings (hybrid cached + live item-agnostic features)
Cite this review
Pith. "Pith review of Memory Layer: Train the In-Model Cache for Recommendation Models." pith.science (2026). https://pith.science/paper/WGKU7TPI
@misc{pith2026260725110,
author = {Pith},
title = {Pith review of: Memory Layer: Train the In-Model Cache for Recommendation Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/WGKU7TPI}},
note = {Machine review of arXiv:2607.25110}
}
abstract
Early ranking stages in recommendation systems precompute item embeddings and cache them in-model for scoring within strict latency constraints. Because this cache exists only at serving time, outside the training loop, training and serving use different item representations, a structural discrepancy that limits quality and adds operational fragility. We show that co-designing the training and serving paths removes this representation discrepancy at its source. We introduce the memory layer, an in-model key-value embedding cache co-trained with the model: the item tower writes embeddings during training and the model reads them at serving, one source of truth for item representations by construction. Always-on embeddings cover items not yet cached, so every item receives a prediction, and the design consolidates three separate trainer-to-predictor update paths into a single self-contained pipeline. Deployed in production on Instagram Reels, the memory layer raises prediction coverage from 96% to 100%, improves embedding freshness from $O(5\text{ min})$ to $O(20\text{ s})$, and narrows the training-serving Normalized Entropy (NE) gap by up to 86%, yielding over $2\times$ recall for the freshest content and a 5-6% cold start engagement lift. Because embeddings are produced during training, the system needs no separate bulk-evaluation or publish-time recomputation, cutting training-and-publish computational cost by 30% at neutral serving computational cost.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc, et al. 2017. TFX: A TensorFlow-Based Production-Scale Machine Learning Platform. InProceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD). ACM
2017
-
[2]
Vincent-Pierre Berges, Barlas Oğuz, Daniel Haziza, Wen-tau Yih, Luke Zettle- moyer, and Gargi Ghosh. 2025. Memory Layers at Scale. InProceedings of the 42nd International Conference on Machine Learning. PMLR
2025
-
[3]
Xiaoshuang Chen, Gengrui Zhang, Yao Wang, Yulin Wu, Shuo Su, Kaiqiao Zhan, and Ben Wang. 2024. Cache-Aware Reinforcement Learning in Large-Scale Recommender Systems. InCompanion Proceedings of the ACM Web Conference
2024
-
[4]
Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep Neural Networks for YouTube Recommendations. InProceedings of the 10th ACM Conference on Recommender Systems (RecSys). ACM
2016
-
[5]
Lu Fang, Shiyan Deng, Hongyi Jia, Huamin Li, Ilina Mitra, Sheng Qin, Zhengkai Zhang, Zhuoran Zhao, and Zinnia Zheng. 2026. Building Highly Efficient Inference System for Recommenders Using PyTorch. https://pytorch.org/blog/building-highly-efficient-inference-system-for- recommenders-using-pytorch/. PyTorch Blog
2026
-
[6]
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014. Neural Turing Machines. arXiv preprint arXiv:1410.5401(2014)
arXiv 2014
-
[7]
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Ag- nieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al. 2016. Hybrid Computing Using a Neural Network with Dynamic External Memory.Nature538, 7626 (2016)
2016
-
[8]
Udit Gupta, Carole-Jean Wu, Xiaodong Wang, Maxim Naumov, Brandon Reagen, David Brooks, Bradford Cottel, Kim Hazelwood, Mark Hempstead, Bill Jia, et al
Show all 41 references
-
[9]
Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, et al. 2014. Practical Lessons from Predict- ing Clicks on Ads at Facebook. InProceedings of the 8th International Workshop on Data Mining for Online Adverti...
2014
-
[10]
Juhee Hong, Meng Liu, Shengzhi Wang, Xiaoheng Mao, Huihui Cheng, Leon Gao, Christopher Leung, Jin Zhou, Chandra Mouli Sekar, Zhao Zhu, Ruochen Liu, Tuan Trieu, Dawei Sun, Jeet Kanjani, Rui Li, Jing Qian, Xuan Cao, Minjie Fan, and Mingze Gao. 2025. Generative Early Stage Rankin...
2025
-
[11]
Jui-Ting Huang, Ashish Sharma, Shuying Sun, Li Xia, David Zhang, Philip Pronin, Janani Padmanabhan, Giuseppe Ottaviano, and Linjun Yang. 2020. Embedding- based Retrieval in Facebook Search. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & ...
2020
-
[12]
Dmytro Ivchenko, Dennis van der Staay, Colin Taylor, Xing Liu, Will Feng, Rahul Kindi, Anirudh Sudarshan, and Shahin Sefati. 2022. TorchRec: A PyTorch Domain Library for Recommendation Systems. InProceedings of the 16th ACM Conference on Recommender Systems (RecSys). ACM
2022
-
[13]
Manos Karpathiotakis, Vlassios Rizopoulos, Basri Kahveci, Tiziano Carotti, Artem Gelum, Hazem Nada, and Yuri Dolgov. 2025. Scribe: How Meta Transports Terabytes per Second in Real Time.Proceedings of the VLDB Endowment(2025)
2025
-
[14]
Xiangyang Li, Bo Chen, HuiFeng Guo, Jingjie Li, Chenxu Zhu, Xiang Long, Sujian Li, Yichao Wang, Wei Guo, Longxia Mao, et al. 2022. IntTower: The Next Generation of Two-Tower Model for Pre-Ranking System. InProceedings of the 31st ACM International Conference on Information and...
2022
-
[15]
Zhuoran Liu, Leqi Zou, Xuan Zou, Caihua Wang, Biao Zhang, Da Tang, Bolin Zhu, Yijie Zhu, Peng Wu, Ke Wang, et al. 2022. Monolith: Real Time Recommendation System With Collisionless Embedding Table. InProceedings of 5th Workshop on Online Recommender Systems and User Modeling, ...
2022
-
[16]
Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine, and Xi Peng. 2022. Are Multimodal Transformers Robust to Missing Modality?. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE
2022
-
[17]
Mengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov, Cathy Wu, and Xi Peng
-
[18]
Kiran Kumar Matam, Hani Ramezani, Fan Wang, Zeliang Chen, Yue Dong, Mao- mao Ding, Zhiwei Zhao, Zhengyu Zhang, Ellie Wen, and Assaf Eisenman. 2024. QuickUpdate: A Real-Time Personalization System for Large-Scale Recommen- dation Models. In21st USENIX Symposium on Networked Sys...
2024
-
[19]
H Brendan McMahan, Gary Holt, David Sculley, Michael Young, Dietmar Ebner, Julian Grady, Lan Nie, Todd Phillips, Eugene Davydov, Daniel Golovin, et al. 2013. Ad Click Prediction: A View from the Trenches. InProceedings of the 19th ACM SIGKDD International Conference on Knowled...
2013
-
[20]
Meta Platforms, Inc. 2026. FBGEMM. https://github.com/pytorch/FBGEMM. Software repository, accessed 2026
2026
-
[21]
Meta Platforms, Inc. 2026. TorchRec. https://github.com/meta-pytorch/torchrec/. Software repository, accessed 2026
2026
-
[22]
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G Azzolini, et al. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems.arXiv preprint...
2019 arXiv
-
[23]
Neoklis Polyzotis, Sudip Roy, Steven Euijong Whang, and Martin Zinkevich. 2018. Data Lifecycle Challenges in Production Machine Learning: A Survey.SIGMOD Record(2018)
2018
-
[24]
Schein, Alexandrin Popescul, Lyle H
Andrew I. Schein, Alexandrin Popescul, Lyle H. Ungar, and David M. Pennock
-
[25]
Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Denni- son
D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Denni- son. 2015. Hidden Technical Debt in Machine Learning Systems. InAdvances in Neural Information Processing Systems (NeurIPS)...
2015
-
[26]
Chijun Sima, Yao Fu, Man-Kit Sit, Liyi Guo, Xuri Gong, Feng Lin, Junyu Wu, Yongsheng Li, Haidong Rong, Pierre-Louis Aublin, and Luo Mai. 2022. Ekko: A Large-Scale Deep Learning Recommender System with Low-Latency Model Up- date. In16th USENIX Symposium on Operating Systems Des...
2022
-
[27]
Ashish Thusoo, Joydeep Sen Sarma, Namit Jain, Zheng Shao, Prasad Chakka, Suresh Anthony, Hao Liu, Pete Wyckoff, and Raghotham Murthy. 2009. Hive: A Warehousing Solution over a Map-Reduce Framework.Proceedings of the VLDB Endowment2, 2 (2009), 1626–1629
2009
-
[28]
Li Wang, Binbin Jin, Zhenya Huang, Hongke Zhao, Defu Lian, Qi Liu, and Enhong Chen. 2021. Preference-Adaptive Meta-Learning for Cold-Start Recommendation. InProceedings of the 30th International Joint Conference on Artificial Intelligence (IJCAI). International Joint Conferenc...
2021
-
[29]
Siqi Wang, Xianjie Chen, Shaofeng Deng, Albert Chen, Romil Shah, Jiawei Huang, Zhaoqin Wang, Zhang Zhang, Yiqun Liu, Meilei Jiang, et al. 2026. SlimPer: Make Personalization Model Slim and Smart.arXiv preprint arXiv:2607.12281(2026)
2026 arXiv
-
[30]
Zhikai Wang, Yanyan Shen, Zibin Zhang, and Kangyi Lin. 2023. Feature Stal- eness Aware Incremental Learning for CTR Prediction. InProceedings of the Thirty-Second International Joint Conference on Artificial Intelligence (IJCAI). In- ternational Joint Conferences on Artificial...
2023
-
[31]
Zhe Wang, Liqin Zhao, Biye Jiang, Guorui Zhou, Xiaoqiang Zhu, and Kun Gai
-
[32]
Bi Xue, Hong Wu, Lei Chen, Chao Yang, Yiming Ma, Fei Ding, Zhen Wang, Liang Wang, Xiaoheng Mao, Ke Huang, et al. 2026. SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs. InProceedings of the 49th International ACM SIGIR Conference on R...
2026
-
[33]
YaChen Yan and Liubo Li. 2024. RankTower: A Synergistic Framework for En- hancing Two-Tower Pre-Ranking Model. InProceedings of the AdKDD Workshop at KDD. ACM
2024
-
[34]
Xinyang Yi, Ji Yang, Lichan Hong, Derek Zhiyuan Cheng, Lukasz Heldt, Aditee Kumthekar, Zhe Zhao, Li Wei, and Ed Chi. 2019. Sampling-Bias-Corrected Neural Modeling for Large Corpus Item Recommendations. InProceedings of the 13th ACM Conference on Recommender Systems (RecSys). ACM
2019
-
[35]
InProceedings of the DLP-KDD Workshop at KDD
COLD: Towards the Next Generation of Pre-Ranking System. InProceedings of the DLP-KDD Workshop at KDD. ACM
-
[36]
Ziliang Zhao, Bi Xue, Emma Lin, Tianqi Lu, Mengjiao Zhou, Kaustubh Vartak, Shakhzod Ali-Zade, Tao Li, Bin Kuang, Rui Jian, et al. 2026. Multi-Probe Zero Collision Hash (MPZCH): Mitigating Embedding Collisions and Enhancing Model Freshness in Large-Scale Recommenders.arXiv prep...
2026 arXiv
-
[39]
Mark Zhao, Dhruv Choudhary, Devashish Tyagi, Ajay Somani, Max Kaplan, Sung-Han Lin, Sarunya Pumma, Jongsoo Park, Aarti Basant, Niket Agarwal, et al
-
[2002]
InProceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR)
Methods and Metrics for Cold-Start Recommendations. InProceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). ACM
-
[2020]
InIEEE International Symposium on High-Performance Computer Architecture (HPCA)
The Architectural Implications of Facebook’s DNN-based Personalized Rec- ommendation. InIEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE
-
[2021]
InProceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI)
SMIL: Multimodal Learning with Severely Missing Modality. InProceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI). AAAI Press
-
[2023]
InProceedings of Machine Learning and Systems (MLSys)
RecD: Deduplication for End-to-End Deep Learning Recommendation Model Training Infrastructure. InProceedings of Machine Learning and Systems (MLSys). mlsys.org
Reviewed July 31, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.