Pith. sign in

REVIEW 14 cited by

MuseCoco: Generating Symbolic Music from Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00110 v1 pith:EZI4XFNM submitted 2023-05-31 cs.SD cs.AIcs.CLcs.LGcs.MMeess.AS

classification cs.SDcs.AIcs.CLcs.LGcs.MMeess.AS
keywords musictextcontroldescriptionsmusecocoattributesgenerationmusical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating music from text descriptions is a user-friendly mode since the text is a relatively easy interface for user engagement. While some approaches utilize texts to control music audio generation, editing musical elements in generated audio is challenging for users. In contrast, symbolic music offers ease of editing, making it more accessible for users to manipulate specific musical elements. In this paper, we propose MuseCoco, which generates symbolic music from text descriptions with musical attributes as the bridge to break down the task into text-to-attribute understanding and attribute-to-music generation stages. MuseCoCo stands for Music Composition Copilot that empowers musicians to generate music directly from given text descriptions, offering a significant improvement in efficiency compared to creating music entirely from scratch. The system has two main advantages: Firstly, it is data efficient. In the attribute-to-music generation stage, the attributes can be directly extracted from music sequences, making the model training self-supervised. In the text-to-attribute understanding stage, the text is synthesized and refined by ChatGPT based on the defined attribute templates. Secondly, the system can achieve precise control with specific attributes in text descriptions and offers multiple control options through attribute-conditioned or text-conditioned approaches. MuseCoco outperforms baseline systems in terms of musicality, controllability, and overall score by at least 1.27, 1.08, and 1.32 respectively. Besides, there is a notable enhancement of about 20% in objective control accuracy. In addition, we have developed a robust large-scale model with 1.2 billion parameters, showcasing exceptional controllability and musicality.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BeatEdit: Symbolic Music Generation as Explicit Editing

    cs.SD 2026-07 conditional novelty 7.0 of 10

    Explicit edit operations on Beat encoding outperform AR and diffusion on music error correction, accompaniment editing, and segment completion while running under 100 ms.

  2. Text2Score: Generating Sheet Music From Textual Prompts

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    Text2Score turns text prompts into sheet music by having an LLM produce a bar-wise structural plan and a hierarchical decoder write ABC notation from that plan.

  3. MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing

    cs.SD 2025-07 conditional novelty 7.0 of 10

    MixAssist is the first audio-grounded, multi-turn conversational dataset for co-creative music mixing instruction, and fine-tuning Qwen-Audio on it yields human-comparable mixing advice.

  4. Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A neuro-symbolic generate-verify-repair harness raises audited twelve-tone delivery yield from 13.3% to 48.1% and improves expert preference without claiming whole-piece legality.

  5. MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI

    cs.HC 2025-09 conditional novelty 6.0 of 10

    MusicScaffold reports that scaffolding generative AI with symbolic explanations and reflective refinement improves 12-14 year olds' structured music expression, strategic adjustments, and self-efficacy compared with d...

  6. Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music

    cs.SD 2025-08 conditional novelty 6.0 of 10

    Amadeus generates symbolic music by autoregressively predicting note-level latents and decoding their attributes in parallel with a masked discrete diffusion model, yielding faster and more controllable generation tha...

  7. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.

  8. TOMI: Transforming and Organizing Music Ideas for Multi-Track Compositions with Full-Song Structure

    cs.SD 2025-06 conditional novelty 6.0 of 10

    TOMI uses a four-node structure plus LLM in-context learning to turn sample clips into whole multi-track electronic songs with planned section structure.

  9. Large Language Models' Internal Perception of Symbolic Music

    cs.CL 2025-07 conditional novelty 5.0 of 10

    LLM-generated MIDI data carries enough genre and style signal to train above-chance classifiers and melody predictors, but far less than real music data.

  10. ASTAR-NTU solution to AudioMOS Challenge 2025 Track1

    cs.SD 2025-07 conditional novelty 5.0 of 10

    DORA-MOS, a dual-branch MuQ/RoBERTa model with cross-attention and Gaussian label softening, achieved the top system-level SRCC of 0.991 for MI and 0.952 for TA on the AudioMOS 2025 Track 1 test set.

  11. From Generality to Mastery: Composer-Style Symbolic Music Generation via Large-Scale Pre-training

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A two-stage pre-train-then-fine-tune transformer with style adapters improves composer-style symbolic piano generation over from-scratch training and the NotaGen baseline.

  12. Workflow-Based Evaluation of Music Generation Systems

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.

  13. FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration

    cs.SD 2026-07 reject novelty 4.0 of 10

    FlowSonic combines deterministic rectified-flow inversion, cached cross-attention injection, and a 'seeded' third-order Adams-Bashforth solver to report better timbre and genre edits on small datasets.

  14. Genre Controlled Music Generation via Activation Steering

    cs.SD 2025-06 conditional novelty 4.0 of 10

    Activation steering with linear probe weights on MusicGen's residual stream shifts generated music between genres at inference time, outperforming text prompting in CLAP and listener preference but with incomplete reporting.

Pith tools