← back to paper
arxiv: 2506.15154 · 2 revisions
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning