REVIEW 2 major objections 6 minor 99 references
Sparse content embeddings beat dense vectors for cold-start recommendations, improving NDCG@20 by 16.6%–75.5% at comparable storage, through sharpening and denoising item-item similarities.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 18:46 UTC pith:POVLUYIS
load-bearing objection Sparse content embeddings for cold-start show large, consistent gains, but the headline comparison is partly confounded by model width and a missing SAE baseline; still a serious paper worth refereeing. the 2 major comments →
Learning Sparse Representations of Multimodal Content for Enhanced Cold Item Recommendation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that sparsification changes the geometry of item-item similarity in a way that suits content-based cold-start: zero dot products between dissimilar items act as built-in denoising, while a pre-sparsification α-entmax activation (applied to a two-sided concatenation [y;−y]) induces sharpness in the surviving similarities. Under the standard random-split cold-start protocol, sparse SEMCo/sparsemax with 32 active dimensions of 1024 beats dense SEMCo at both equal storage and at up to 16× smaller storage (a 64-cost sparse model outperforming 1024-d dense), and the advantage concentrates among users with several interests. The paper also claims the sparse dimensions a
What carries the argument
The machinery is the top-k sparsification layer applied to content encoder outputs, plus pre-sparsification activation functions (PSAFs), specifically α-entmax with two-sided concatenation. Top-k enforces explicit sparsity before L2 normalization; the PSAF sharpens and denoises similarity distributions without breaking dot-product inference, in the spirit of linear attention where a nonlinearity on keys and queries approximates softmax. The α-entmax family interpolates between softmax (α=1) and sparsemax (α=2), allowing the model to zero out low logits and calibrate sharpness, while preserving negatives through two-siding retains useful signal. This is what carries the argument: sparsity doe
Load-bearing premise
The load-bearing premise, flagged by the authors in Section 4.1, is that a randomly held-out item behaves like a genuinely new item; if production cold-start items differ systematically in content distribution from old catalog items, the reported gains may not transfer.
What would settle it
A temporal-split evaluation on the same four datasets, where the cold set is the most recent 20% of items by timestamp instead of a random sample: if sparse SEMCo's advantage over dense SEMCo shrinks or reverses when items are genuinely new, the central claim fails.
If this is right
- Sparse content embeddings can replace dense content embeddings in cold-start pipelines with accuracy gains and storage savings.
- Users with multiple interests benefit disproportionately, making sparse representations a partial answer to multi-interest modeling in content-based cold-start.
- The sparse latent dimensions are semantically interpretable and cover ground-truth categories, enabling segment-level analysis without metadata supervision.
- The modest training overhead (about 15–20%) and fast dot-product retrieval with few active dimensions make the approach practical for industrial-scale inverted-index retrieval.
- Accuracy is robust to the choice of sparsity level, so practitioners can trade storage against accuracy without retraining from scratch.
Where Pith is reading between the lines
- Because the gains are largest for multi-interest users, a direct test is to measure the same models on a dataset where user interest diversity is explicitly controlled; the gap should widen with the number of interest clusters.
- The linear-attention analogy suggests the PSAF could be swapped for other attention-inspired kernels, and a learnable α per layer might calibrate sharpening more flexibly than the fixed α=2 sparsemax used here.
- The paper's own stated limitation is the random-split cold-start protocol; a temporal split with the most recent 20% of items as the cold set is the natural stress test for production transfer.
- If sparsity is acting as a regularizer rather than only as compression, similar accuracy gains might appear in warm content-based recommenders, not just cold-start—a testable extension beyond the paper.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a pipeline for learning sparse content embeddings for cold-start item recommendation. It adapts the SEMCo and ELSA content-based training regimes to produce top-k sparsified embeddings, and introduces pre-sparsification activation functions (PSAFs) based on alpha-entmax and two-sided concatenation. The central claim, stated in the abstract and Section 5.1, is that sparse embeddings improve cold-start NDCG@20/R@20 by 16.6%–75.5% over the best dense content-based LAE at comparable storage budgets, with larger gains for users with multiple interests, and with additional storage and interpretability benefits. The evaluation covers four multimodal datasets, storage ablations, PSAF ablations, training efficiency, and a PMI-based interpretability case study.
Significance. If the central claim is robust, the paper makes a practically valuable contribution: it shows that sparse content embeddings can improve cold-start ranking accuracy while reducing storage, and it provides a plausible mechanism (sharpness and denoising of item-item similarities) plus an interpretability analysis. The manuscript has clear strengths: the implementation is released, the training regimes for sparse and dense methods are otherwise matched, four datasets are used, and the multi-interest analysis in Figures 4–5 is a useful attempt at explanation. The main obstacle is that the headline sparse-vs-dense comparison is confounded by model width and capacity, and a post-hoc sparsification baseline is missing, so the paper does not currently isolate the effect of sparsity itself.
major comments (2)
- [Table 2; §4.2.1; §5.1.1] The headline comparison of sparse vs. dense embeddings is confounded by model capacity. Sparse methods use a 1024-dimensional content-encoder output (2048 for two-sided PSAFs), while the dense baselines use dimension 64. Thus the projection layers of the sparse models have roughly 16x–32x more parameters than the dense models they are compared with. Storage is matched by active dimensions (32 active ≈ 64 dense), but parameter count and forward computation are not. The reported 16.6%–75.5% gains could therefore be due to the wider representation space rather than to sparsity per se. Section 5.1.1 concludes that sparsification 'universally improves' cold-start outcomes, but this is only demonstrated under unequal model sizes. Please add matched-capacity comparisons (e.g., dense 1024d content LAEs with re-tuned hyperparameters, or sparse models with width 64/128), and a post-hoc sparsificat
- [§4.1; §5.1.1; Conclusion] The evaluation protocol randomly selects 20% of items as the cold set, and the footnote acknowledges that temporal splitting is more realistic but is not used. This assumption is load-bearing for the real-world claim: newly added items in production may differ systematically in content distribution, popularity, and modality composition from randomly held-out catalog items. If so, the reported gains may not transfer to actual cold-start scenarios. The paper should provide evidence on distribution shift (e.g., compare content features of random vs. temporal cold sets) or run a temporal split where timestamps are available, and state whether the main conclusions hold. At minimum, the abstract and conclusion should qualify the claim as applying to random-split offline evaluation, or explicitly identify this as a limitation in the conclusion rather than only in a footnote.
minor comments (6)
- [Tables 2 and 3] All metrics are averaged over five runs, but no standard deviations or confidence intervals are reported. Given that many method–dataset–metric combinations are tested, please report error bars or at least worst/best over runs, and consider a multiple-comparison correction or a note that the main conclusions are robust without correction.
- [Table 2 caption and §5.1.1] The improvement percentage is computed by comparing the best sparse method against the best dense content LAE. Since the best sparse method is selected after inspecting results, this can overstate the expected improvement of a single configuration. Report the performance of a pre-specified representative configuration, or average over the sparse variants.
- [§5.4] The interpretability analysis uses top-8 active dimensions, while the deployed model uses 32 active dimensions. It would be useful to state whether the PMI structure persists at 32 active dimensions, or to justify the choice of top-8 as an interpretability-only setting.
- [Figures 4–5 and §3.2.2] The user-interest groups are derived from K-means on SEMCo embeddings, which are the same type of model being evaluated. This is understandable as an analytic tool, but a robustness check with different numbers of clusters or with an independent interest proxy (e.g., genre labels) would strengthen the claim that the multi-interest gains are not an artifact of the clustering choice.
- [Throughout] There are a few typos, e.g., 'interpetability' in the contribution list, and the reference to 'As seen in Figure 8' in §5.4 is misleading because Figure 8 is a grid of user/item active-dimension results, not a genre-robustness plot. Please renumber or rephrase.
- [Eq. (7)] The notation for alpha-entmax is slightly compressed; explicitly define eta as the threshold and clarify the domain of y (post two-siding) so that the formula is unambiguous.
Circularity Check
No significant circularity: empirical study with design motivated by a case study, not a derivation that reduces to its inputs.
full rationale
The paper is an empirical study rather than a derivation, and its central claims are tested on held-out cold items across four datasets against multiple dense baselines. The sparse pipeline is built by adapting the authors' prior SEMCo method, but this is a legitimate base-model choice: dense SEMCo serves as a baseline, and the sparse variant changes only the transformation pipeline (top-k sparsification and PSAF) while keeping the objective unchanged. The case study in Section 3.2 identifies sharpening and denoising as desirable properties by applying non-linear functions to dense SEMCo similarities, and the PSAF is then designed to induce these properties; however, the resulting method is evaluated independently and the properties are measured post hoc in Figure 6, so no prediction reduces to a fitted input by construction. The storage comparisons are matched by design (32 active dims vs 64 dense dims), but the accuracy improvements are empirical outcomes, not algebraic consequences. Self-citations to SEMCo [43], popularity-bias work [42], and artist-catalog splitting [59] are used as framing or experimental-protocol justifications, not as load-bearing proofs; the authors also explicitly flag the temporal-split limitation, which is an external-validity concern rather than circularity. No uniqueness theorem is imported from the authors, no ansatz is smuggled in solely via self-citation, and no known result is merely renamed. The only notable issue is the skeptic's confound of model width versus sparsity, which is a fairness-of-comparison concern, not a circularity one.
Axiom & Free-Parameter Ledger
free parameters (8)
- PSAF alpha (entmax order) =
2 (sparsemax); 1.5 in ablations
- PSAF temperature omega =
tuned per dataset; ranges on GitHub
- Active dimensions k =
32 (default)
- Embedding width =
1024 (2048 for two-sided methods)
- Number of interest clusters (K-means) =
64
- Top-200 similarity cutoff =
200
- Top-8 active dims for interpretability =
8
- Exponential pruning schedule =
decays active dims from full width to 32 over training
axioms (8)
- domain assumption SEMCo's contrastive regime yields content embeddings whose dot-product similarities correlate with user preference (Eqs. 1-2, Section 3.1).
- domain assumption Users can be represented as the sum of the embeddings of their interacted items (RY in Eq. 2) with acceptable loss of preference signal.
- domain assumption Top-k sparsification by absolute value, followed by L2 normalization, preserves the semantic ordering needed for recommendation (Section 3.4.1).
- domain assumption The linear-attention analogy (Section 3.3) implies that sparsification and PSAFs induce beneficial spikiness in item-item similarities.
- domain assumption Storage cost equivalence: n active sparse dimensions cost about the same as 2n dense dimensions (float32 values + int16 indices, Section 5.1).
- domain assumption Random splitting of items into warm and cold sets is a valid proxy for production cold-start (Section 4.1).
- ad hoc to paper K-means clusters over SEMCo embeddings identify meaningful 'user interests' (Sections 3.2.2, 5.1.3).
- standard math Paired t-test across five runs, with no multiple-comparison correction, is valid evidence of statistical significance (Section 5.1.1).
read the original abstract
The scale and rapid growth of item catalogs in modern digital platforms present significant challenges to recommender system (RS) practitioners. Most RSs use embedding similarity to predict user-item preferences, but embedding storage and low-latency retrieval are challenging in industry-scale catalogs. Furthermore, newly added items do not have corresponding embeddings and cannot be recommended effectively; previous works often tackle this item cold-start problem by generating cold item representations from auxiliary content, such as images or descriptive text, so that user preferences can be predicted without historical interactions. In this paper, we argue that sparse embeddings have notable advantages over standard dense vectors in this content-based cold-start paradigm. We describe how existing cold-start training regimes can be adapted for sparse representation learning, and build on insights from linear attention to design a pre-sparsification activation technique that induces sharpness and denoising effects in learned item-item similarities. We show that the resulting sparse embeddings achieve significant improvements in cold-start recommendation accuracy over dense embeddings at considerably lower storage costs, especially for users with multiple interests. Through comprehensive experiments on four multimodal RS datasets, we also demonstrate the interpretability of sparse content embeddings and their robustness in the trade-off between size and accuracy.
Figures
Reference graph
Works this paper leans on
-
[1]
Haoyue Bai, Min Hou, Le Wu, Yonghui Yang, Kun Zhang, Richang Hong, and Meng Wang. 2023. Gorec: a generative cold-start recommendation framework. In Proceedings of the 31st ACM International Conference on Multimedia. 1004–1012
2023
-
[2]
Haoyue Bai, Le Wu, Min Hou, Miaomiao Cai, Zhuangzhuang He, Yuyang Zhou, Richang Hong, and Meng Wang. 2024. Multimodality invariant learning for multimedia-based new item recommendation. InProceedings of the 47th Inter- national ACM SIGIR Conference on Research and Development in Information Retrieval. 677–686
2024
-
[3]
2018.Annoy: Approximate Nearest Neighbors in C++/Python
Erik Bernhardsson. 2018.Annoy: Approximate Nearest Neighbors in C++/Python. https://pypi.org/project/annoy/
2018
-
[5]
Sebastian Bruch, Franco Maria Nardini, Cosimo Rulli, and Rossano Venturini
-
[6]
Yukuo Cen, Jianwei Zhang, Xu Zou, Chang Zhou, Hongxia Yang, and Jie Tang
-
[7]
Beidi Chen, Tri Dao, Eric Winsor, Zhao Song, Atri Rudra, and Christopher Ré
-
[8]
Hao Chen, Zefan Wang, Feiran Huang, Xiao Huang, Yue Xu, Yishi Lin, Peng He, and Zhoujun Li. 2022. Generative adversarial framework for cold-start item recommendation. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2565–2571
2022
-
[9]
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. 2020. Rethinking attention with performers.arXiv preprint arXiv:2009.14794(2020)
Pith/arXiv arXiv 2020
-
[10]
Yuhong Chou, Man Yao, Kexin Wang, Yuqi Pan, Ruijie Zhu, Yiran Zhong, Yu Qiao, Jibin Wu, Bo Xu, and Guoqi Li. 2024. Metala: Unified optimal linear approximation to softmax attention map.Advances in Neural Information Processing Systems37 (2024), 71034–71067
2024
-
[11]
Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep neural networks for youtube recommendations. InProceedings of the 10th ACM Conference on Recommender Systems. 191–198
2016
-
[12]
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey
-
[13]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. CoRRabs/1810.04805 (2018). arXiv:1810.04805 http://arxiv.org/abs/1810.04805
Pith/arXiv arXiv 2018
-
[14]
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024. The faiss library.arXiv preprint arXiv:2401.08281(2024)
Pith/arXiv arXiv 2024
-
[15]
Xiaoyu Du, Xiang Wang, Xiangnan He, Zechao Li, Jinhui Tang, and Tat-Seng Chua. 2020. How to learn item representation for cold-start multimedia recom- mendation?. InProceedings of the 28th ACM International Conference on Multime- dia
2020
-
[16]
Stanley C Eisenstat, MC Gursky, Martin H Schultz, and Andrew H Sherman. 1977. Yale sparse matrix package. i. the symmetric codes. Technical Report
1977
-
[17]
Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. Splade: Sparse lexical and expansion model for first stage ranking. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2288–2292
2021
-
[18]
Christian Ganhör, Marta Moscati, Anna Hausberger, Shah Nawaz, and Markus Schedl. 2024. A Multimodal Single-Branch Embedding Network for Recommen- dation in Cold-Start and Missing Modality Scenarios. InProceedings of the 18th ACM Conference on Recommender Systems. 380–390
2024
-
[19]
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. 2024. Scaling and evaluating sparse autoencoders.arXiv preprint arXiv:2406.04093(2024)
Pith/arXiv arXiv 2024
-
[20]
Nuno Gonçalves, Marcos V Treviso, and Andre Martins. 2025. AdaSplash: Adap- tive Sparse Flash Attention. InInternational Conference on Machine Learning. PMLR, 19878–19896
2025
-
[21]
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets.Advances in Neural Information Processing Systems27 (2014)
2014
-
[22]
Huifeng Guo, Wei Guo, Yong Gao, Ruiming Tang, Xiuqiang He, and Wenzhi Liu
-
[23]
Jordan, and Nikhil Garg
Wenshuo Guo, Karl Krauth, Michael I. Jordan, and Nikhil Garg. 2021. The Stereotyping Problem in Collaboratively Filtered Recommender Systems.Equity and Access in Algorithms, Mechanisms, and Optimization(2021)
2021
-
[24]
Feiran Huang, Yuanchen Bei, Zhenghang Yang, Junyi Jiang, Hao Chen, Qijie Shen, Senzhang Wang, Fakhri Karray, and Philip S Yu. 2025. Large Language Model Simulator for Cold-Start Recommendation. InProceedings of the Eighteenth ACM International Conference on Web Search and Data Mining. 261–270
2025
-
[25]
Feiran Huang, Zefan Wang, Xiao Huang, Yufeng Qian, Zhetao Li, and Hao Chen
-
[26]
Petr Kasalick`y, Martin Spišák, Vojtěch Vančura, Daniel Bohuněk, Rodrigo Alves, and Pavel Kordík. 2025. The Future is Sparse: Embedding Compression for Scalable Retrieval in Recommender Systems. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 1099–1103
2025
-
[27]
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret
-
[28]
InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval
Scalefreectr: Mixcache-based distributed training system for ctr models with huge embedding table. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1269–1278
-
[29]
Diederik P. Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Opti- mization.CoRRabs/1412.6980 (2014)
Pith/arXiv arXiv 2014
-
[30]
Diederik P. Kingma and Max Welling. 2014. Auto-Encoding Variational Bayes. In2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings. arXiv:http://arxiv.org/abs/1312.6114v10
Pith/arXiv arXiv 2014
-
[31]
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, et al. 2022. Matryoshka representation learning.Advances in Neural Information Processing Systems35 (2022), 30233–30249
2022
-
[32]
InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’23)
Aligning Distillation For Cold-start Item Recommendation. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’23). Association for Computing Machinery, New York, NY, USA
-
[33]
Chao Li, Zhiyuan Liu, Mengmeng Wu, Yuchi Xu, Huan Zhao, Pipei Huang, Guoliang Kang, Qiwei Chen, Wei Li, and Dik Lun Lee. 2019. Multi-interest network with dynamic routing for recommendation at Tmall. InProceedings of the 28th ACM International Conference on Information and Knowledge Management. 2615–2623
2019
-
[34]
Guohui Li, Li Zou, Zhiying Deng, and Qi Chen. 2024. Neighborhood-Enhanced Multimodal Collaborative Filtering for Item Cold Start Recommendation. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 7815–7819
2024
-
[35]
InInternational Conference on machine learning
Transformers are rnns: Fast autoregressive transformers with linear atten- tion. InInternational Conference on machine learning. PMLR, 5156–5165
-
[36]
Jinri Kim, Eungi Kim, Kwangeun Yeo, Yujin Jeon, Chanwoo Kim, Sewon Lee, and Joonseok Lee. 2024. Content-based Graph Reconstruction for Cold-start Item Recommendation. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1263–1273
2024
-
[37]
James MacQueen. 1967. Some methods for classification and analysis of multivari- ate observations. InProceedings of the fifth Berkeley symposium on mathematical statistics and probability, Vol. 1. 281–297
1967
-
[38]
Daniele Malitesta, Emanuele Rossi, Claudio Pomo, Tommaso Di Noia, and Fragkiskos D Malliaros. 2026. Training-free Graph-based Imputation of Missing Modalities in Multimodal Recommendation.IEEE Transactions on Knowledge and Data Engineering(2026)
2026
-
[39]
Andre Martins and Ramon Astudillo. 2016. From softmax to sparsemax: A sparse model of attention and multi-label classification. InInternational Conference on machine learning. PMLR, 1614–1623
2016
-
[40]
Honglak Lee, Alexis Battle, Rajat Raina, and Andrew Ng. 2006. Efficient sparse coding algorithms.Advances in Neural Information Processing Systems19 (2006)
2006
-
[41]
Gregor Meehan and Johan Pauwels. 2025. Artist Considerations in Offline Eval- uation of Music Recommender Systems. MuRS 2025: 3rd Music Recommender Systems Workshop, September 22nd, 2025
2025
-
[42]
Gregor Meehan and Johan Pauwels. 2025. On Inherited Popularity Bias in Cold- Start Item Recommendation. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 649–654
2025
-
[43]
Shiwei Li, Huifeng Guo, Xing Tang, Ruiming Tang, Lu Hou, Ruixuan Li, and Rui Zhang. 2024. Embedding compression in recommender systems: A survey. Comput. Surveys56, 5 (2024), 1–21
2024
-
[44]
Ilya Loshchilov and Frank Hutter. 2016. Sgdr: Stochastic gradient descent with warm restarts.arXiv preprint arXiv:1608.03983(2016)
Pith/arXiv arXiv 2016
-
[45]
Marta Moscati, Emilia Parada-Cabaleiro, Yashar Deldjoo, Eva Zangerle, and Markus Schedl. 2022. Music4All-Onion – A Large-Scale Multi-Faceted Content- Centric Music Recommendation Dataset. InProceedings of the 31st ACM In- ternational Conference on Information & Knowledge Management(Atlanta, GA, USA)(CIKM ’22). Association for Computing Machinery, New York...
arXiv 2022
-
[46]
Yongxin Ni, Yu Cheng, Xiangyan Liu, Junchen Fu, Youhua Li, Xiangnan He, Yongfeng Zhang, and Fajie Yuan. 2023. A Content-Driven Micro-Video Recom- mendation Dataset at Scale.arXiv preprint arXiv:2309.15379(2023)
Pith/arXiv arXiv 2023
-
[47]
Xia Ning and George Karypis. 2011. Slim: Sparse linear methods for top-n recommender systems. In2011 IEEE 11th International Conference on data mining. IEEE, 497–506
2011
-
[48]
Julian McAuley, Christopher Targett, Qinfeng Shi, and Anton Van Den Hengel
-
[49]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library.Advances in Neural Information Processing Systems32 (2019)
2019
-
[50]
Ben Peters, Vlad Niculae, and André FT Martins. 2019. Sparse sequence-to- sequence models.arXiv preprint arXiv:1905.05702(2019)
Pith/arXiv arXiv 2019
-
[51]
Igor André Pegoraro Santana, Fabio Pinhelli, Juliano Donini, Leonardo Catharin, Rafael Biazus Mangolin, Valéria Delisandra Feltrim, Marcos Aurélio Domingues, et al. 2020. Music4all: A new music database and its applications. In2020 In- ternational Conference on Systems, Signals and Image Processing (IWSSIP). IEEE, 399–404
2020
-
[52]
Gregor Meehan and Johan Pauwels. 2026. Sparse Contrastive Learning for Content-Based Cold Item Recommendation. InProceedings of the 49th Interna- tional ACM SIGIR Conference on Research and Development in Information Retrieval. 3994–3999
2026
-
[53]
Zaiqiao Meng, Richard McCreadie, Craig Macdonald, and Iadh Ounis. 2020. Ex- ploring data splitting strategies for the evaluation of recommendation models. In Proceedings of the 14th ACM Conference on Recommender Systems. 681–686. RecSys ’26, September 27-October 02, 2026, Minneapolis, MN, USA Gregor Meehan and Johan Pauwels
2020
-
[54]
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.arXiv preprint arXiv:1701.06538(2017)
Pith/arXiv arXiv 2017
-
[55]
Kihyuk Sohn, Honglak Lee, and Xinchen Yan. 2015. Learning structured output representation using deep conditional generative models.Advances in Neural Information Processing Systems28 (2015)
2015
-
[56]
Mohamed Sordo, Oscar Celma, Martin Blech, and Enric Guaus. 2008. The quest for musical genres: Do the experts and the wisdom of crowds agree?. InISMIR. 255–260
2008
-
[57]
Biswajit Paria, Chih-Kuan Yeh, Ian EH Yen, Ning Xu, Pradeep Ravikumar, and Barnabás Póczos. 2020. Minimizing flops to learn efficient sparse representations. arXiv preprint arXiv:2004.05665(2020)
Pith/arXiv arXiv 2020
-
[58]
Changfeng Sun, Han Liu, Meng Liu, Zhaochun Ren, Tian Gan, and Liqiang Nie
-
[59]
Yan-Martin Tamm, Gregor Meehan, Vojtěch Nekl, Vojtech Vancura, Rodrigo Alves, Johan Pauwels, and Anna Aljanaki. 2026. Leveraging Artist Catalogs for Cold-Start Music Recommendation. InProceedings of the 34th ACM Conference on User Modeling, Adaptation and Personalization. 137–146
2026
-
[60]
Aaron Van den Oord, Sander Dieleman, and Benjamin Schrauwen. 2013. Deep content-based music recommendation.Advances in Neural Information Processing Systems26 (2013)
2013
-
[61]
Andrew I Schein, Alexandrin Popescul, Lyle H Ungar, and David M Pennock
-
[62]
Vojtěch Vančura, Petr Kasalick`y, Rodrigo Alves, and Pavel Kordík. 2025. Evaluat- ing Linear Shallow Autoencoders on Large Scale Datasets.ACM Transactions on Recommender Systems(2025)
2025
-
[63]
Wenling Shang, Kihyuk Sohn, Diogo Almeida, and Honglak Lee. 2016. Under- standing and improving convolutional neural networks via concatenated rectified linear units. InInternational Conference on machine learning. PMLR, 2217–2225
2016
-
[64]
Vojtěch Vančura, Martin Spišák, Rodrigo Alves, and Ladislav Peška. 2026. Ef- ficient Learning of Sparse Representations from Interactions.arXiv preprint arXiv:2602.09935(2026)
arXiv 2026
-
[65]
Maksims Volkovs, Guangwei Yu, and Tomi Poutanen. 2017. Dropoutnet: Ad- dressing cold start in recommender systems.Advances in Neural Information Processing Systems30 (2017)
2017
-
[66]
Jianling Wang, Haokai Lu, and Minmin Chen. 2024. Fresh content recommenda- tion at scale: A multi-funnel solution and the potential of LLMs. InProceedings of the 17th ACM International Conference on Web Search and Data Mining. 1186– 1187
2024
-
[67]
Harald Steck. 2019. Embarrassingly shallow autoencoders for sparse data. InThe World Wide Web Conference. 3251–3257
2019
-
[68]
Wenbo Wang, Bingquan Liu, Lili Shan, Chengjie Sun, Ben Chen, and Jian Guan
-
[69]
InProceedings of the 13th International Conference on Web Search and Data Mining
LARA: Attribute-to-feature adversarial learning for new-item recommen- dation. InProceedings of the 13th International Conference on Web Search and Data Mining
-
[70]
Tiansheng Wen, Yifei Wang, Zequn Zeng, Zhong Peng, Yudi Su, Xinyang Liu, Bo Chen, Hongwei Liu, Stefanie Jegelka, and Chenyu You. 2025. Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation. InInternational Conference on Machine Learning. PMLR, 66520–66538
2025
-
[71]
Minz Won, Yun-Ning Hung, and Duc Le. 2024. A foundation model for music informatics. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1226–1230
2024
-
[72]
Vojtěch Vančura, Rodrigo Alves, Petr Kasalick`y, and Pavel Kordík. 2022. Scalable linear shallow autoencoder for collaborative filtering. InProceedings of the 16th ACM Conference on Recommender Systems. 604–609
2022
-
[73]
Junda Wu, Cheng-Chun Chang, Tong Yu, Zhankui He, Jianing Wang, Yupeng Hou, and Julian McAuley. 2024. Coral: collaborative retrieval-augmented large language models improve long-tail recommendation. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3391–3401
2024
-
[74]
Vojtěch Vančura, Pavel Kordík, and Milan Straka. 2024. beeFormer: Bridging the Gap Between Semantic and Interaction Similarity in Recommender Systems. In Proceedings of the 18th ACM Conference on Recommender Systems. 1102–1107
2024
-
[75]
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. 2021. Nyströmformer: A nyström-based algorithm for approximating self-attention. InProceedings of the AAAI conference on artificial intelligence, Vol. 35. 14138–14148
2021
-
[76]
Zhiqiang Xu, Dong Li, Weijie Zhao, Xing Shen, Tianbo Huang, Xiaoyun Li, and Ping Li. 2021. Agile and accurate CTR prediction model training for massive-scale online advertising systems. InProceedings of the 2021 International Conference on management of data. 2404–2409
2021
-
[77]
Hamed Zamani, Mostafa Dehghani, W Bruce Croft, Erik Learned-Miller, and Jaap Kamps. 2018. From neural re-ranking to neural ranking: Learning a sparse representation for inverted indexing. InProceedings of the 27th ACM International Conference on Information and Knowledge Management. 497–506
2018
-
[78]
Lei Wang and Ee-Peng Lim. 2023. Zero-shot next-item recommendation using large pretrained language models.arXiv preprint arXiv:2304.03153(2023)
Pith/arXiv arXiv 2023
-
[79]
Michael Zhang, Kush Bhatia, Hermann Kumbong, and Christopher Ré. 2024. The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry. arXiv preprint arXiv:2402.04347(2024)
Pith/arXiv arXiv 2024
-
[80]
InProceedings of the AAAI Conference on Artificial Intelligence, Vol
Preference Aware Dual Contrastive Learning for Item Cold-Start Recom- mendation. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 9125–9132
-
[81]
Yinwei Wei, Xiang Wang, Qi Li, Liqiang Nie, Yan Li, Xuanping Li, and Tat-Seng Chua. 2021. Contrastive learning for cold-start recommendation. InProceedings of the 29th ACM International Conference on Multimedia. 5382–5390
2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.