Pith. sign in

REVIEW 5 cited by

LLark: A Multimodal Instruction-Following Language Model for Music

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07160 v3 pith:GVNOLFFO submitted 2023-10-11 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords musicllarkmodelmultimodalunderstandingaudioavailablecaptioning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Music has a unique and complex structure which is challenging for both expert humans and existing AI systems to understand, and presents unique challenges relative to other forms of audio. We present LLark, an instruction-tuned multimodal model for \emph{music} understanding. We detail our process for dataset creation, which involves augmenting the annotations of diverse open-source music datasets and converting them to a unified instruction-tuning format. We propose a multimodal architecture for LLark, integrating a pretrained generative model for music with a pretrained language model. In evaluations on three types of tasks (music understanding, captioning, reasoning), we show that LLark matches or outperforms existing baselines in music understanding, and that humans show a high degree of agreement with its responses in captioning and reasoning tasks. LLark is trained entirely from open-source music data and models, and we make our training code available along with the release of this paper. Additional results and audio examples are at https://bit.ly/llark, and our source code is available at https://github.com/spotify-research/llark .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing

    cs.SD 2025-07 conditional novelty 7.0 of 10

    MixAssist is the first audio-grounded, multi-turn conversational dataset for co-creative music mixing instruction, and fine-tuning Qwen-Audio on it yields human-comparable mixing advice.

  2. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  3. Assessing Factual Music Comprehension in Large Audio Language Models

    cs.SD 2025-11 conditional novelty 6.0 of 10

    Standard NLP metrics fail to capture factual music understanding in audio-language models; a CLAP-based metric and an LLM-parsed factual QA protocol measure it more directly.

  4. Can Large Language Models Predict Audio Effects Parameters from Natural Language?

    cs.SD 2025-05 conditional novelty 6.0 of 10

    LLMs can predict equalizer and reverb parameters from natural language descriptions, and adding DSP features, DSP function code, and few-shot examples improves the predictions.

  5. Exploring listeners' perceptions of AI-generated and human-composed music for functional emotional applications

    cs.HC 2025-06 conditional novelty 5.0 of 10

    Preference and perceived emotional efficacy dissociate for AI-generated versus human-composed music, with listeners preferring AI tracks but crediting human tracks with stronger functional emotion elicitation.

Pith tools