REVIEW 14 cited by
MuseCoco: Generating Symbolic Music from Text
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Generating music from text descriptions is a user-friendly mode since the text is a relatively easy interface for user engagement. While some approaches utilize texts to control music audio generation, editing musical elements in generated audio is challenging for users. In contrast, symbolic music offers ease of editing, making it more accessible for users to manipulate specific musical elements. In this paper, we propose MuseCoco, which generates symbolic music from text descriptions with musical attributes as the bridge to break down the task into text-to-attribute understanding and attribute-to-music generation stages. MuseCoCo stands for Music Composition Copilot that empowers musicians to generate music directly from given text descriptions, offering a significant improvement in efficiency compared to creating music entirely from scratch. The system has two main advantages: Firstly, it is data efficient. In the attribute-to-music generation stage, the attributes can be directly extracted from music sequences, making the model training self-supervised. In the text-to-attribute understanding stage, the text is synthesized and refined by ChatGPT based on the defined attribute templates. Secondly, the system can achieve precise control with specific attributes in text descriptions and offers multiple control options through attribute-conditioned or text-conditioned approaches. MuseCoco outperforms baseline systems in terms of musicality, controllability, and overall score by at least 1.27, 1.08, and 1.32 respectively. Besides, there is a notable enhancement of about 20% in objective control accuracy. In addition, we have developed a robust large-scale model with 1.2 billion parameters, showcasing exceptional controllability and musicality.
Forward citations
Cited by 14 Pith papers
-
BeatEdit: Symbolic Music Generation as Explicit Editing
Explicit edit operations on Beat encoding outperform AR and diffusion on music error correction, accompaniment editing, and segment completion while running under 100 ms.
-
Text2Score: Generating Sheet Music From Textual Prompts
Text2Score turns text prompts into sheet music by having an LLM produce a bar-wise structural plan and a hierarchical decoder write ABC notation from that plan.
-
MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing
MixAssist is the first audio-grounded, multi-turn conversational dataset for co-creative music mixing instruction, and fine-tuning Qwen-Audio on it yields human-comparable mixing advice.
-
Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation
A neuro-symbolic generate-verify-repair harness raises audited twelve-tone delivery yield from 13.3% to 48.1% and improves expert preference without claiming whole-piece legality.
-
MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI
MusicScaffold reports that scaffolding generative AI with symbolic explanations and reflective refinement improves 12-14 year olds' structured music expression, strategic adjustments, and self-efficacy compared with d...
-
Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music
Amadeus generates symbolic music by autoregressively predicting note-level latents and decoding their attributes in parallel with a masked discrete diffusion model, yielding faster and more controllable generation tha...
-
Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.
-
TOMI: Transforming and Organizing Music Ideas for Multi-Track Compositions with Full-Song Structure
TOMI uses a four-node structure plus LLM in-context learning to turn sample clips into whole multi-track electronic songs with planned section structure.
-
Large Language Models' Internal Perception of Symbolic Music
LLM-generated MIDI data carries enough genre and style signal to train above-chance classifiers and melody predictors, but far less than real music data.
-
ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
DORA-MOS, a dual-branch MuQ/RoBERTa model with cross-attention and Gaussian label softening, achieved the top system-level SRCC of 0.991 for MI and 0.952 for TA on the AudioMOS 2025 Track 1 test set.
-
From Generality to Mastery: Composer-Style Symbolic Music Generation via Large-Scale Pre-training
A two-stage pre-train-then-fine-tune transformer with style adapters improves composer-style symbolic piano generation over from-scratch training and the NotaGen baseline.
-
Workflow-Based Evaluation of Music Generation Systems
A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.
-
FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration
FlowSonic combines deterministic rectified-flow inversion, cached cross-attention injection, and a 'seeded' third-order Adams-Bashforth solver to report better timbre and genre edits on small datasets.
-
Genre Controlled Music Generation via Activation Steering
Activation steering with linear probe weights on MusicGen's residual stream shifts generated music between genres at inference time, outperforming text prompting in CLAP and listener preference but with incomplete reporting.
Discussion (0). Sign in to comment.