JAM models natural-language music queries as vector translations from user to item embeddings, and with cross-attention fusion of audio, lyrics, and collaborative-filtering signals it outperforms the tested baselines on a new 112k-triple session dataset.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Just Ask for Music (JAM): Multimodal and Personalized Natural Language Music Recommendation
JAM models natural-language music queries as vector translations from user to item embeddings, and with cross-attention fusion of audio, lyrics, and collaborative-filtering signals it outperforms the tested baselines on a new 112k-triple session dataset.