REVIEW 9 cited by
ChatMusician: Understanding and Generating Music Intrinsically with LLM
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While Large Language Models (LLMs) demonstrate impressive capabilities in text generation, we find that their ability has yet to be generalized to music, humanity's creative language. We introduce ChatMusician, an open-source LLM that integrates intrinsic musical abilities. It is based on continual pre-training and finetuning LLaMA2 on a text-compatible music representation, ABC notation, and the music is treated as a second language. ChatMusician can understand and generate music with a pure text tokenizer without any external multi-modal neural structures or tokenizers. Interestingly, endowing musical abilities does not harm language abilities, even achieving a slightly higher MMLU score. Our model is capable of composing well-structured, full-length music, conditioned on texts, chords, melodies, motifs, musical forms, etc, surpassing GPT-4 baseline. On our meticulously curated college-level music understanding benchmark, MusicTheoryBench, ChatMusician surpasses LLaMA2 and GPT-3.5 on zero-shot setting by a noticeable margin. Our work reveals that LLMs can be an excellent compressor for music, but there remains significant territory to be conquered. We release our 4B token music-language corpora MusicPile, the collected MusicTheoryBench, code, model and demo in GitHub.
Forward citations
Cited by 9 Pith papers
-
Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music
A meta-benchmark that auto-generates multimodal music-perception multiple-choice tests from user symbolic music, demonstrated on ChoraleBricks with text-only and white-noise controls.
-
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models
A synthetic music sheet QA dataset and a LoRA-fine-tuned Phi-3 model show large accuracy gains on OMR and chord tasks, but only within the synthetic distribution.
-
Can Large Language Models Predict Audio Effects Parameters from Natural Language?
LLMs can predict equalizer and reverb parameters from natural language descriptions, and adding DSP features, DSP function code, and few-shot examples improves the predictions.
-
Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation
A bar-level symbolic-score song generator (BACH) is claimed to beat published systems and commercial Suno on human-rated quality, duration, and efficiency, but the supporting full text is corrupted and unverifiable.
-
Exploring GPT's Ability as a Judge in Music Understanding
GPT-3.5 detects synthetic annotation errors in beat, chord, and key tasks above random chance, with accuracy partly influenced by musical concepts in the prompt.
-
WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation
An open multi-agent system that orchestrates specialized music models for understanding, composition, and synthesis, with local or hosted deployment.
-
CoComposer: LLM Multi-agent Collaborative Music Composition
A five-agent LLM system for ABC-notation composition scores modestly higher than ComposerX and a single LLM on an automated aesthetic model, but no error bars or significance tests are reported.
-
Are Large Language Models Capable of Deep Relational Reasoning? Insights from DeepSeek-R1 and Benchmark Comparisons
DeepSeek-R1 outperforms GPT-4o and DeepSeek-V3 on family tree and graph reasoning benchmarks at sizes 10 and 20, but all models collapse at size 40.
-
Content filtering methods for music recommendation: A review
A survey of content-based music recommendation methods, including audio analysis, lyrics analysis, and context awareness, with no new experimental results.
Discussion (0). Continue with ORCID to comment.