REVIEW 4 major objections 5 minor 1 cited by
Progressive Semantic Residual Quantization for Multimodal-Joint Interest Modeling in Music Recommendation
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that residual quantization can be made semantics-preserving by clustering each residual together with the prefix reconstruction from earlier layers, and that the resulting modal-specific and modal-joint IDs, read by a…
desk verdict Real deployment and a plausible mechanism, but Eq. (2) does not type-check beyond layer 2 and the offline gains are tiny; still worth a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the PSRQ codebook-construction identity in Eq. (2): every quantization layer $l \ge 2$ runs K-means on $\mathbf{X}^m_l \oplus (\mathbf{X}^m - \mathbf{X}^m_l)$, the current residual concatenated with the prefix semantic reconstruction from previous layers, instead of on the bare residual. This is what claims to stop semantic drift, by forcing each deeper centroid to be chosen with the original content in view. The second mechanism is MCCA's shared-query cross-attention: the modal-joint embedding of the target item acts as the query over the modal-specific and joint embedding sequences of the user's history, with a separate collaborative attention stream, so the modal-joint codebook carries cross-modal correlation while the modal-specific codebooks preserve fine-grained interests.
What would settle it
On the paper's own datasets, compute the conditional entropy of PSRQ's second- and third-layer cluster assignments given the first-layer assignment; if that entropy is near zero, deeper codes are essentially copying the first layer and PSRQ is not adding semantic detail. A complementary check is to compare, in the original embedding space, the average pairwise similarity of items sharing a full PSRQ ID with items sharing an RQ ID of the same depth.
Extended reading notes
Core claim
The central claim is that conventional residual quantization's layer-by-layer loss of original semantics is avoidable. Where RQ clusters the residual $\mathbf{X}^m_l$ at each layer, PSRQ builds each codebook layer $l\ge2$ on the concatenation $\mathbf{X}^m_l \oplus (\mathbf{X}^m - \mathbf{X}^m_l)$, so the cluster centers at depth $l$ are determined jointly by what is still left to explain and by the semantic prefix already reconstructed above. The paper claims this prefix feature anchors the deeper codes to the item's original content without giving up RQ's hierarchical approximation. On the modeling side, MCCA allocates separate embedding tables to textual, audio, and modal-joint semantic IDs and uses the target item's modal-joint embedding as a shared query in cross-attention over the user's history sequences, which the paper argues captures per-modality taste and cross-modal associations in one pass. The experimental claim is that PSRQ+MCCA reports higher or equal AUC and lower logloss than VBPR, DIN, SimTier+MAKE, and QARM on the three datasets, with cold-start gains up to +1.76% relative, and that deployed A/B tests show 2.81% higher collect and 0.95% higher full-play rates.
Load-bearing premise
The load-bearing premise is that appending the earlier layers' reconstruction to each residual anchors the next clustering without swamping the residual; if that prefix dominates, items sharing a first-layer cluster would receive nearly identical deeper codes and the claimed semantic-preservation benefit would collapse.
Editorial extensions
If this is right
- A PSRQ codebook remains semantically interpretable at depth, so the same IDs could serve downstream tasks beyond the ranking classifier, such as retrieval or generative ranking, without re-quantizing content.
- Cold-start items benefit most in the reported tables because any item with content can be assigned a semantic ID immediately and inherit trainable codebook embeddings.
- The separation of modal-specific and modal-joint codebooks makes the framework modular: a new modality can be added by quantizing its content and training its own attention stream while keeping the shared query.
- Because the ID embeddings are randomly initialized and trained end-to-end rather than frozen centroids, the recommendation stage adapts the content representation to observed user behavior.
- The online A/B results imply the offline AUC gains carry into live commercial metrics, with larger relative gains on tracks released within the last 30 days.
Reading between the lines
- Inference: the PSRQ prefix concatenation is a modality-agnostic correction, so it could be dropped into any residual-quantization pipeline for short-video or e-commerce content IDs; the paper demonstrates it only for text and audio, but Eq. (2) does not depend on the modality.
- Inference: given the modest All-AUC gains (+0.24% to +1.04%) alongside larger cold-start gains, the mechanism's practical value is probably concentrated where content semantics matter most; slicing the reported data by interaction sparsity would test this directly.
- Inference: MCCA is a late-fusion design, so an obvious untested combination is to keep PSRQ's semantic IDs and add a light contrastive alignment between modal and joint embeddings, capturing the alignment signal used by QARM without surrendering modal-specific codebooks.
- Inference: a direct way to check whether the deeper codes are doing real work is to measure the entropy of layer-$l$ cluster assignments conditioned on the first-layer cluster; if that entropy is near zero, the appended prefix is dominating and the semantic benefit is an artifact of shallower structure.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a two-stage multimodal recommendation framework for music. Stage 1 introduces Progressive Semantic Residual Quantization (PSRQ), a modification of residual quantization that concatenates each layer's residual with the original feature vector to preserve prefix semantic information, producing modal-specific and modal-joint semantic IDs. Stage 2 introduces a Multi-Codebook Cross-Attention (MCCA) network that uses a modal-joint codebook as a shared query to attend over modal-specific ID sequences, with the goal of capturing both modal-specific and cross-modal user interests. The framework is evaluated on Amazon Baby, an industrial dataset, and Music4All in offline experiments, plus an online A/B test on a music streaming platform, and compared against DIN, VBPR, SimTier+MAKE, and QARM. The paper claims state-of-the-art performance and significant improvements in commercial metrics.
Significance. The problem addressed—semantic drift in hierarchical quantization and the difficulty of fusing modal-specific and cross-modal information—is relevant and timely for industrial multimodal recommendation. The proposed PSRQ idea of explicitly constraining residuals with the original feature is intuitive and could be a useful design pattern, and the MCCA attention mechanism is a reasonable approach. The paper reports experiments on three datasets and includes an online deployment, which gives ecological validity. However, the central claim of semantic preservation is never directly measured, and the definition of PSRQ in Eq. (2) is not dimensionally consistent for the number of layers used in the experiments. The offline performance gains are very small, with no statistical significance testing. The contribution would be strengthened by a corrected formal definition, direct semantic-fidelity metrics, and significance analysis; as written, the results do not yet substantiate the state-of-the-art claim.
major comments (4)
- [Section 4.1, Eq. (2)] The PSRQ recurrence is dimensionally inconsistent for l≥3. Since C2 is obtained from the concatenated input X^m_2 ⊕ (X^m - X^m_2), C2 ∈ R^{k×2d}. The recurrence then states X^m_3 = X^m_2 − NearestRep(X^m_2, C2), which requires comparing the d-dimensional residual X^m_2 to 2d-dimensional centers and subtracting a 2d vector from a d-dimensional vector. The manuscript never specifies whether the search is conducted on a projected d-dimensional subspace, whether the codebook vectors are split, or how this is resolved; the same issue propagates to all deeper layers. Since the experimental setting uses l={3,4,3,3} (Section 5.1.2), Eq. (2) as written cannot be implemented, and the reported results cannot be attributed to a well-defined algorithm.
- [Section 5.2, Tables 2–4] The offline improvements are within a range that cannot be distinguished from noise. In Table 2, the All AUC improvement over the best baseline is +0.24% on Amazon Baby, +1.04% on Industrial, and +0.00% on Music4All (where PSRQ+MCCA and QARM both report 0.7347). No standard deviations, confidence intervals, significance tests, or repeated runs are provided for any metric. Given effect sizes at the third decimal place, the statement that PSRQ+MCCA 'outperforms state-of-the-art baselines' is not supported by the reported evidence.
- [Section 4.1 and Section 5.2] The paper's motivating claim that PSRQ 'explicitly preserves the prefix semantic feature' is never directly tested. The experiments only report recommendation AUC/Logloss; they do not report per-layer codebook diversity, residual norms, reconstruction error of the original embedding from the concatenated representation, or any metric of semantic similarity between the original multimodal features and the generated semantic IDs. The observed downstream gains are therefore not evidence for the semantic-preservation mechanism; an alternative explanation, such as codebook redundancy or increased embedding capacity, cannot be ruled out.
- [Section 5.3] The online A/B test is reported with only four relative lift numbers (2.81% collect, 0.95% full_play, 5.98%/2.2% for new tracks, 3.05% listening hours) and no details on the number of users in each arm, the duration of the test, the choice of metrics, or any statistical significance measure. The baseline is DLRM, which is not one of the offline baselines, making it hard to compare with the offline results. Given the small offline effect sizes, the online claims of significant improvement require a more rigorous presentation to be acceptable.
minor comments (5)
- [Section 3, Eq. (1)] The RQ recurrence uses 'X^t_2' in the second line where 'X^m_2' is intended; please fix the subscript typo.
- [Tables 2–4] The '%Improv.' sign convention for Logloss is confusing because lower Logloss is better, but the reported improvements are negative. Please define the formula or use a sign that makes 'improvement' positive for better metrics.
- [Section 1, Introduction] The bullet item 'intra-semantic of multimodal' is ungrammatical and appears to be a fragment of 'intra-modal semantic degradation'; please revise.
- [Section 4.2.2, Eq. (4)] The notation suggests a standard cross-attention operation with Q, K, V, but the formula Attention(e^z_j ⊕ e^o_t) · e^z_j uses a scalar attention weight from a feed-forward network; the equivalence between the two formulations should be spelled out.
- [Section 4.1, Figure 2] The text says 'The different of semantic IDs retrieve process between conventional RQ and PSRQ is shown in Figure 2,' but Figure 2 is referenced without a detailed description of the panels; please expand the caption and refer to it at the appropriate place.
Circularity Check
No significant circularity: PSRQ and MCCA are defined design choices validated on external datasets and baselines; the main weakness is a dimensional underspecification in Eq. (2), which is a correctness issue, not a circularity.
full rationale
The paper does not fit a parameter to a subset of data and then report the same quantity as a prediction. PSRQ is introduced in Eq. (2) as a modified residual-quantization codebook construction: each layer clusters the residual concatenated with the original modal vector. That is an explicit design choice rather than an outcome derived from the evaluation labels. The downstream MCCA model is trained with the same cross-entropy objective as the baselines, and all comparisons are against external datasets (Amazon Baby, an industrial dataset, Music4all) and external baselines (VBPR, SimTier+MAKE, DIN, QARM, PQ, VQ, RQ, RQ-VAE), with an additional online A/B test against DLRM. No self-citation is load-bearing: the author self-citations in the reference list are not used to justify the central claim, and no uniqueness theorem is imported from prior work by the same authors. The only notable weakness is that Eq. (2) is dimensionally underspecified for layers l≥3 because C_l is said to be in R^{k×2d} while the residual X^m_l used in the nearest-neighbor search remains d-dimensional; however, that is a reproducibility/correctness concern, explicitly outside the circularity definition. Since the derivation chain is a design-and-benchmark comparison rather than a fit-to-prediction loop, no circular step can be exhibited, so the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- Number of clusters k per dataset =
Amazon Baby: 64, Industrial: 256, Music4all: 128
- Number of quantization layers l =
Amazon Baby: 3, Industrial: 4, Music4all: 3
- Embedding dimension d' =
64
- Learning rates and batch sizes =
Amazon Baby: lr 5e-4, batch 64; Industrial and Music4all: lr 1e-4, batch 512
assumptions (3)
- standard math K-means clustering converges to a useful local optimum for the multimodal embeddings.
- domain assumption LLaMA3.2, Baichuan2, and MERT embeddings capture semantics relevant to user preference.
- domain assumption User behavior sequences of positive interactions (clicks, collects) contain enough signal to learn multimodal interests.
Cite this review
Pith. "Pith review of Progressive Semantic Residual Quantization for Multimodal-Joint Interest Modeling in Music Recommendation." pith.science (2026). https://pith.science/paper/6DLRPP6T
@misc{pith2026250820359,
author = {Pith},
title = {Pith review of: Progressive Semantic Residual Quantization for Multimodal-Joint Interest Modeling in Music Recommendation},
year = {2026},
howpublished = {\url{https://pith.science/paper/6DLRPP6T}},
note = {Machine review of arXiv:2508.20359}
}
read the original abstract
In music recommendation systems, multimodal interest learning is pivotal, which allows the model to capture nuanced preferences, including textual elements such as lyrics and various musical attributes such as different instruments and melodies. Recently, methods that incorporate multimodal content features through semantic IDs have achieved promising results. However, existing methods suffer from two critical limitations: 1) intra-modal semantic degradation, where residual-based quantization processes gradually decouple discrete IDs from original content semantics, leading to semantic drift; and 2) inter-modal modeling gaps, where traditional fusion strategies either overlook modal-specific details or fail to capture cross-modal correlations, hindering comprehensive user interest modeling. To address these challenges, we propose a novel multimodal recommendation framework with two stages. In the first stage, our Progressive Semantic Residual Quantization (PSRQ) method generates modal-specific and modal-joint semantic IDs by explicitly preserving the prefix semantic feature. In the second stage, to model multimodal interest of users, a Multi-Codebook Cross-Attention (MCCA) network is designed to enable the model to simultaneously capture modal-specific interests and perceive cross-modal correlations. Extensive experiments on multiple real-world datasets demonstrate that our framework outperforms state-of-the-art baselines. This framework has been deployed on one of China's largest music streaming platforms, and online A/B tests confirm significant improvements in commercial metrics, underscoring its practical value for industrial-scale recommendation systems.
Figures
Forward citations
Cited by 1 Pith paper
-
The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation
A dual-branch recommender that merges hash-ID and semantic-ID representations outperforms baselines while improving tail-item accuracy without losing head-item accuracy.
Reference graph
Works this paper leans on
-
[1]
Artem Babenko and Victor Lempitsky. 2014. Additive Quantization for Extreme Vector Compression. In 2014 IEEE Conference on Computer Vision and Pattern Recognition. 931–938. doi:10.1109/CVPR.2014.124
-
[2]
Feiyu Chen, Junjie Wang, Yinwei Wei, Hai-Tao Zheng, and Jie Shao. 2022. Break- ing isolation: Multimodal graph fusion for multimedia recommendation by edge- wise modulation. In Proceedings of the 30th ACM International Conference on Multimedia. 385–394
2022
-
[3]
Gaode Chen, Ruina Sun, Yuezihan Jiang, Jiangxia Cao, Qi Zhang, Jingjian Lin, Han Li, Kun Gai, and Xinghua Zhang. 2024. A Multi-modal Modeling Framework for Cold-start Short-video Recommendation. In Proceedings of the 18th ACM Confer- ence on Recommender Systems (Bari, Italy) (RecSys ’24). Association for Computing Machinery, New York, NY, USA, 391–400. do...
arXiv 2024
-
[4]
Jiaxin Deng, Shiyao Wang, Kuo Cai, Lejian Ren, Qigen Hu, Weifeng Ding, Qiang Luo, and Guorui Zhou. 2025. OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment. arXiv:2502.18965 [cs.IR] https://arxiv.org/abs/2502.18965
arXiv 2025
-
[5]
Sohrab Ferdowsi, Slava Voloshynovskiy, and Dimche Kostadinov. 2017. Regular- ized Residual Quantization: a multi-layer sparse dictionary learning approach. arXiv:1705.00522 [cs.LG] https://arxiv.org/abs/1705.00522
work page Pith review arXiv 2017
-
[6]
Jennifer Fiore. 2016. Analysis of Lyrics from Group Songwriting with Bereaved Children and Adolescents. Journal of Music Ther- apy 53, 3 (05 2016), 207–231. arXiv:https://academic.oup.com/jmt/article- pdf/53/3/207/7953134/thw005.pdf doi:10.1093/jmt/thw005
-
[7]
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. ImageBind One Embedding Space to Bind Them All. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 15180–15190. doi:10.1109/CVPR52729.2023.01457
arXiv 2023
-
[8]
James A Hanley and Barbara J McNeil. 1982. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology 143, 1 (1982), 29–36
1982
Show all 50 references
-
[9]
Ruining He and Julian McAuley. 2016. VBPR: visual Bayesian Personalized Ranking from implicit feedback. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (Phoenix, Arizona) (AAAI’16). AAAI Press, 144–150
2016
-
[10]
Yupeng Hou, Zhankui He, Julian McAuley, and Wayne Xin Zhao. 2023. Learn- ing Vector-Quantized Item Representation for Transferable Sequential Rec- ommenders. In Proceedings of the ACM Web Conference 2023 (Austin, TX, USA) (WWW ’23). Association for Computing Machinery, New Yor...
2023
-
[11]
Yupeng Hou, Jiacheng Li, Zhankui He, An Yan, Xiusi Chen, and Julian McAuley
-
[12]
Hengchang Hu, Wei Guo, Yong Liu, and Min-Yen Kan. 2023. Adaptive Multi- Modalities Fusion in Sequential Recommendation Systems. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Manage- ment (Birmingham, United Kingdom) (CIKM ’23). Associatio...
2023
-
[13]
Herve Jégou, Matthijs Douze, and Cordelia Schmid. 2011. Product Quantization for Nearest Neighbor Search. IEEE Transactions on Pattern Analysis and Machine Intelligence 33, 1 (2011), 117–128. doi:10.1109/TPAMI.2010.57
2011 doi
-
[14]
Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic opti- mization. arXiv preprint arXiv:1412.6980 (2014)
2014 arXiv
-
[15]
Zhirui Kuai, Zuxu Chen, Huimu Wang, Mingming Li, Dadong Miao, Binbin Wang, Xusong Chen, Li Kuang, Yuxing Han, Jiaxing Wang, Guoyu Tang, Lin Liu, Songlin Wang, and Jingwei Zhuo. 2024. Breaking the Hourglass Phenomenon of Residual Quantization: Enhancing the Upper Bound of Gener...
2024 doi
-
[16]
Taku Kudo. 2018. Subword regularization: Improving neural network translation models with multiple subword candidates.arXiv preprint arXiv:1804.10959 (2018)
2018 arXiv
-
[17]
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. 2022. Autoregressive Image Generation Using Residual Quantization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 11523–11532
2022
-
[18]
Guanghan Li, Xun Zhang, Yufei Zhang, Yifan Yin, Guojun Yin, and Wei Lin. 2025. Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic Tokenization. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. AAAI P...
2025 doi
-
[19]
Selvaraju, Akhilesh D
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh D. Gotmare, Shafiq Joty, Caim- ing Xiong, and Steven C.H. Hoi. 2021. Align before fuse: vision and language representation learning with momentum distillation. In Proceedings of the 35th International Conference on Neural Informati...
2021
-
[20]
Yue Li, Wenrui Ding, Chunlei Liu, Baochang Zhang, and Guodong Guo. 2021. TRQ: Ternary Neural Networks With Residual Quantization. In Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence , Vol. 35. AAAI Press, Palo Alto, California, USA, 8538–8546. doi:10....
2021 doi
-
[21]
Yizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma, Xingran Chen, Hanzhi Yin, Cheng- hao Xiao, Chenghua Lin, Anton Ragni, Emmanouil Benetos, et al. 2023. Mert: Acoustic music understanding model with large-scale self-supervised training. arXiv preprint arXiv:2306.00107 (2023)
2023 arXiv
-
[22]
Zefan Li, Bingbing Ni, Wenjun Zhang, Xiaokang Yang, and Wen Gao. 2017. Perfor- mance Guaranteed Network Acceleration via High-Order Residual Quantization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2017
-
[23]
Ji Lin, Hongxu Yin, Wei Ping, Yao Lu, Pavlo Molchanov, Andrew Tao, Huizi Mao, Jan Kautz, Mohammad Shoeybi, and Song Han. 2023. VILA: On Pre-training for Visual Language Models. arXiv:2312.07533 [cs.CV]
2023 arXiv
-
[24]
Linden, B
G. Linden, B. Smith, and J. York. 2003. Amazon.com recommendations: item-to- item collaborative filtering. IEEE Internet Computing 7, 1 (2003), 76–80. doi:10. 1109/MIC.2003.1167344
2003 arXiv
-
[25]
Yifan Liu, Kangning Zhang, Xiangyuan Ren, Yanhua Huang, Jiarui Jin, Yingjie Qin, Ruilong Su, Ruiwen Xu, Yong Yu, and Weinan Zhang. 2024. AlignRec: Aligning and Training in Multimodal Recommendations. In Proceedings of the 33rd ACM International Conference on Information and Kn...
2024
-
[26]
Xinchen Luo, Jiangxia Cao, Tianyu Sun, Jinkai Yu, Rui Huang, Wei Yuan, Hezheng Lin, Yichen Zheng, Shiyao Wang, Qigen Hu, Changqing Qiu, Jiaqi Zhang, Xu Zhang, Zhiheng Yan, Jingming Zhang, Simin Zhang, Mingxing Wen, Zhaojie Liu, Kun Gai, and Guorui Zhou. 2024. QARM: Quantitativ...
2024 arXiv
-
[27]
Yiqing Ma, David Baker, Katherine Vukovics, Connor Davis, and Emily Elliott
-
[28]
Vladimir Malinovskii, Andrei Panferov, Ivan Ilin, Han Guo, Peter Richtárik, and Dan Alistarh. 2024. Pushing the Limits of Large Language Model Quantization via the Linearity Theorem. arXiv:2411.17525 [cs.LG] https://arxiv.org/abs/2411.17525
2024 arXiv
-
[29]
Hoos, and James J
Julieta Martinez, Holger H. Hoos, and James J. Little. 2014. Stacked Quantizers for Compositional Vector Compression. arXiv:1411.2173 [cs.CV] https://arxiv. org/abs/1411.2173
2014 arXiv
-
[30]
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherni- avskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, V...
2019 arXiv
-
[31]
Rafailidis, P
D. Rafailidis, P. Kefalas, and Y. Manolopoulos. 2017. Preference dynamics with multimodal user-item interactions in social media recommendation. Expert Systems with Applications 74 (2017), 11–18. doi:10.1016/j.eswa.2017.01.005
2017 doi
-
[32]
Tran, Jonah Samost, Maciej Kula, Ed H
Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Keshavan, Trung Vu, Lukasz Heidt, Lichan Hong, Yi Tay, Vinh Q. Tran, Jonah Samost, Maciej Kula, Ed H. Chi, and Maheswaran Sathiamoorthy. 2023. Recommender systems with generative retrieval. In Proceedings of the 37th Inte...
2023
-
[33]
Soravitt Sangnark, Phairot Autthasan, Puntawat Ponglertnapakorn, Phudit Chalekarn, Thapanun Sudhawiyangkul, Manatsanan Trakulruangroj, Sarita Songsermsawad, Rawin Assabumrungrat, Supalak Amplod, Kajornvut Ounjai, and Theerawit Wilaiprasitporn. 2021. Revealing Preference in Pop...
2021
-
[34]
Mangolin, Yandre M
Igor André Pegoraro Santana, Fabio Pinhelli, Juliano Donini, Leonardo Gabiato Catharin, Rafael B. Mangolin, Yandre M. G. Costa, Valéria Delisandra Feltrim, and Marcos Aurélio Domingues. 2020. Music4All: A New Music Database and Its Applications. 2020 International Conference o...
2020
-
[35]
Xiang-Rong Sheng, Feifan Yang, Litong Gong, Biao Wang, Zhangming Chan, Yujing Zhang, Yueyao Cheng, Yong-Nan Zhu, Tiezheng Ge, Han Zhu, Yuning Jiang, Jian Xu, and Bo Zheng. 2024. Enhancing Taobao Display Advertising with Multimodal Representations: Challenges, Approaches and In...
2024
-
[36]
Anima Singh, Trung Vu, Nikhil Mehta, Raghunandan Keshavan, Maheswaran Sathiamoorthy, Yilin Zheng, Lichan Hong, Lukasz Heldt, Li Wei, Devansh Tandon, Ed Chi, and Xinyang Yi. 2024. Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations. In Proceedin...
2024 doi
-
[37]
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017. Neural discrete representation learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17). Curran Associates Inc., Red Hook, NY, ...
2017
-
[38]
Jiaqi Wang, Hanqi Jiang, Yiheng Liu, Chong Ma, Xu Zhang, Yi Pan, Mengyuan Liu, Peiran Gu, Sichen Xia, Wenjun Li, Yutong Zhang, Zihao Wu, Zhengliang Liu, Tianyang Zhong, Bao Ge, Tuo Zhang, Ning Qiang, Xintao Hu, Xi Jiang, Xin Zhang, Wei Zhang, Dinggang Shen, Tianming Liu, and S...
2024 arXiv
-
[39]
Shijia Wang, Tianpei Ouyang, Yunfan Zhou, Qiang Xiao, Yintao Ren, Yifei Pan, Fangjian Li, and Chuanjiang Luo. 2025. Enhanced Emotion-aware Music Recom- mendation via Large Language Models. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining ...
2025
-
[40]
Shijia Wang, Yi Zheng, Qiang Xiao, Yilong Zhao, Qimeng Yang, and Chuanjiang Luo. 2024. Sparsity-Aware Personalized Pattern Extractor Network for Music Multi-task Learning. In Database Systems for Advanced Applications , Makoto Onizuka, Jae-Gil Lee, Yongxin Tong, Chuan Xiao, Yo...
2024
-
[41]
Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, Richang Hong, and Tat-Seng Chua. 2019. MMGCN: Multi-modal Graph Convolution Network for Personalized Recommendation of Micro-video. In Proceedings of the 27th ACM International Conference on Multimedia (Nice, France) (MM ’19). ...
2019
-
[42]
Songpei Xu, Shijia Wang, Da Guo, Xianwen Guo, Qiang Xiao, Fangjian Li, and Chuanjiang Luo. 2025. An Efficient Large Recommendation Model: Towards a Resource-Optimal Scaling Law. arXiv:2502.09888 [cs.IR] https://arxiv.org/abs/ 2502.09888
2025 arXiv
-
[43]
Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, Fei Deng, Feng Wang, Feng Liu, Guangwei Ai, Guosheng Dong, Haizhou Zhao, Hang Xu, Haoze Sun, Hongda Zhang, Hui Liu, Jiaming Ji, Jian Xie, JunTao Dai, Kun Fa...
2023 arXiv
-
[44]
Qimeng Yang, Shijia Wang, Da Guo, Dongjin Yu, Qiang Xiao, Dongjing Wang, and Chuanjiang Luo. 2024. Cascading Multimodal Feature Enhanced Contrast Learning for Music Recommendation. In 2024 IEEE International Conference on Data Mining (ICDM). 905–910. doi:10.1109/ICDM59182.2024.00113
2024
-
[45]
Jiashuo Yu, Jinyu Liu, Ying Cheng, Rui Feng, and Yuejie Zhang. 2022. Modality- aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence Detection. In Proceedings of the 30th ACM International Conference on Multimedia (Lisboa, Portu...
2022
-
[46]
Carolina Zheng, Minhui Huang, Dmitrii Pedchenko, Kaushik Rangadurai, Siyu Wang, Gaby Nahum, Jie Lei, Yang Yang, Tao Liu, Zutian Luo, et al. 2025. Enhancing Embedding Representation Stability in Recommendation Systems with Semantic ID. arXiv preprint arXiv:2504.02137 (2025)
2025 arXiv
-
[47]
Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai. 2018. Deep Interest Network for Click- Through Rate Prediction. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining...
2018
-
[48]
Hang Zhou, Yucheng Wang, and Huijing Zhan. 2025. MDE: Modality Discrimi- nation Enhancement for Multi-modal Recommendation. arXiv:2502.18481 [cs.IR] https://arxiv.org/abs/2502.18481
2025 arXiv
-
[2021]
doi:10.31234/osf.io/ 5ku43
Generalizing the Effect of Lyrics on Emotion Rating. doi:10.31234/osf.io/ 5ku43
-
[2024]
arXiv:2403.03952 [cs.IR] https://arxiv.org/abs/2403.03952
Bridging Language and Items for Retrieval and Recommendation. arXiv:2403.03952 [cs.IR] https://arxiv.org/abs/2403.03952
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.