REVIEW 5 cited by
LLark: A Multimodal Instruction-Following Language Model for Music
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Music has a unique and complex structure which is challenging for both expert humans and existing AI systems to understand, and presents unique challenges relative to other forms of audio. We present LLark, an instruction-tuned multimodal model for \emph{music} understanding. We detail our process for dataset creation, which involves augmenting the annotations of diverse open-source music datasets and converting them to a unified instruction-tuning format. We propose a multimodal architecture for LLark, integrating a pretrained generative model for music with a pretrained language model. In evaluations on three types of tasks (music understanding, captioning, reasoning), we show that LLark matches or outperforms existing baselines in music understanding, and that humans show a high degree of agreement with its responses in captioning and reasoning tasks. LLark is trained entirely from open-source music data and models, and we make our training code available along with the release of this paper. Additional results and audio examples are at https://bit.ly/llark, and our source code is available at https://github.com/spotify-research/llark .
Forward citations
Cited by 5 Pith papers
-
MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing
MixAssist is the first audio-grounded, multi-turn conversational dataset for co-creative music mixing instruction, and fine-tuning Qwen-Audio on it yields human-comparable mixing advice.
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
Assessing Factual Music Comprehension in Large Audio Language Models
Standard NLP metrics fail to capture factual music understanding in audio-language models; a CLAP-based metric and an LLM-parsed factual QA protocol measure it more directly.
-
Can Large Language Models Predict Audio Effects Parameters from Natural Language?
LLMs can predict equalizer and reverb parameters from natural language descriptions, and adding DSP features, DSP function code, and few-shot examples improves the predictions.
-
Exploring listeners' perceptions of AI-generated and human-composed music for functional emotional applications
Preference and perceived emotional efficacy dissociate for AI-generated versus human-composed music, with listeners preferring AI tracks but crediting human tracks with stronger functional emotion elicitation.
Discussion (0). Continue with ORCID to comment.