Pith. sign in

REVIEW 26 cited by

MuseCoco: Generating Symbolic Music from Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00110 v1 pith:EZI4XFNM submitted 2023-05-31 cs.SD cs.AIcs.CLcs.LGcs.MMeess.AS

classification cs.SDcs.AIcs.CLcs.LGcs.MMeess.AS
keywords musictextcontroldescriptionsmusecocoattributesgenerationmusical
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Generating music from text descriptions is a user-friendly mode since the text is a relatively easy interface for user engagement. While some approaches utilize texts to control music audio generation, editing musical elements in generated audio is challenging for users. In contrast, symbolic music offers ease of editing, making it more accessible for users to manipulate specific musical elements. In this paper, we propose MuseCoco, which generates symbolic music from text descriptions with musical attributes as the bridge to break down the task into text-to-attribute understanding and attribute-to-music generation stages. MuseCoCo stands for Music Composition Copilot that empowers musicians to generate music directly from given text descriptions, offering a significant improvement in efficiency compared to creating music entirely from scratch. The system has two main advantages: Firstly, it is data efficient. In the attribute-to-music generation stage, the attributes can be directly extracted from music sequences, making the model training self-supervised. In the text-to-attribute understanding stage, the text is synthesized and refined by ChatGPT based on the defined attribute templates. Secondly, the system can achieve precise control with specific attributes in text descriptions and offers multiple control options through attribute-conditioned or text-conditioned approaches. MuseCoco outperforms baseline systems in terms of musicality, controllability, and overall score by at least 1.27, 1.08, and 1.32 respectively. Besides, there is a notable enhancement of about 20% in objective control accuracy. In addition, we have developed a robust large-scale model with 1.2 billion parameters, showcasing exceptional controllability and musicality.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

    cs.SD 2026-08 conditional novelty 7.0 of 10

    Controlled experiments show that a 10ms performance-timed token stream lowers Frechet Music Distance roughly twofold versus beat-grid tokens, across model sizes from 0.8B to 27B.

  2. BeatEdit: Symbolic Music Generation as Explicit Editing

    cs.SD 2026-07 conditional novelty 7.0 of 10

    Explicit edit operations on Beat encoding outperform AR and diffusion on music error correction, accompaniment editing, and segment completion while running under 100 ms.

  3. Text2Score: Generating Sheet Music From Textual Prompts

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    Text2Score turns text prompts into sheet music by having an LLM produce a bar-wise structural plan and a hierarchical decoder write ABC notation from that plan.

  4. MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing

    cs.SD 2025-07 conditional novelty 7.0 of 10

    MixAssist is the first audio-grounded, multi-turn conversational dataset for co-creative music mixing instruction, and fine-tuning Qwen-Audio on it yields human-comparable mixing advice.

  5. Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping

    cs.SD 2026-08 conditional novelty 6.0 of 10

    Using matched neutral and target-swap prompts, ACE-Step 1.5 and Stable Audio 3 show real key control and partial beat control, while LeVo2 does not, and much four-beat agreement is just the models' default output.

  6. MAJEPPA: Morphing and Assessing in a Unified Piano Performance Space

    eess.AS 2026-08 conditional novelty 6.0 of 10

    A score-conditioned piano model trained with a JEPA objective and contrastive losses learns embeddings that improve several performance-understanding benchmarks over existing MIDI foundation models.

  7. Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model

    cs.SD 2026-08 conditional novelty 6.0 of 10

    Diff-Symbo generates long, text-controlled symbolic music by autoregressively extending 8-bar latent diffusion segments conditioned on the previous segment's latent.

  8. Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A cross-modal model extracts performance style from audio and uses it to generate matching piano arrangements from lead sheets, with style transfer and audio-to-MIDI retrieval.

  9. DDSynth-RL: Audio Synthesizer Inversion via Discrete Diffusion with Reinforcement Learning

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A masked discrete diffusion model fine-tuned with GRPO on rendered-audio rewards improves out-of-domain synthesizer parameter estimation on the Dexed FM synthesizer.

  10. Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A neuro-symbolic generate-verify-repair harness raises audited twelve-tone delivery yield from 13.3% to 48.1% and improves expert preference without claiming whole-piece legality.

  11. MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI

    cs.HC 2025-09 conditional novelty 6.0 of 10

    MusicScaffold reports that scaffolding generative AI with symbolic explanations and reflective refinement improves 12-14 year olds' structured music expression, strategic adjustments, and self-efficacy compared with d...

  12. Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music

    cs.SD 2025-08 conditional novelty 6.0 of 10

    Amadeus generates symbolic music by autoregressively predicting note-level latents and decoding their attributes in parallel with a masked discrete diffusion model, yielding faster and more controllable generation tha...

  13. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.

  14. TOMI: Transforming and Organizing Music Ideas for Multi-Track Compositions with Full-Song Structure

    cs.SD 2025-06 conditional novelty 6.0 of 10

    TOMI uses a four-node structure plus LLM in-context learning to turn sample clips into whole multi-track electronic songs with planned section structure.

  15. Learning Musical Representations for Music Performance Question Answering

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Amuse fuses audio, video, and language early and adds rhythm and instrument-source supervision over time to reach state-of-the-art accuracy on Music AVQA and Music AVQA-v2.

  16. Text2midi: Generating Symbolic Music from Captions

    cs.SD 2024-12 conditional novelty 6.0 of 10

    An encoder-decoder transformer conditioned on frozen FLAN-T5 text embeddings generates MIDI from captions, with pretraining on pseudo-captioned SymphonyNet and fine-tuning on MidiCaps.

  17. Large Language Models' Internal Perception of Symbolic Music

    cs.CL 2025-07 conditional novelty 5.0 of 10

    LLM-generated MIDI data carries enough genre and style signal to train above-chance classifiers and melody predictors, but far less than real music data.

  18. ASTAR-NTU solution to AudioMOS Challenge 2025 Track1

    cs.SD 2025-07 conditional novelty 5.0 of 10

    DORA-MOS, a dual-branch MuQ/RoBERTa model with cross-attention and Gaussian label softening, achieved the top system-level SRCC of 0.991 for MI and 0.952 for TA on the AudioMOS 2025 Track 1 test set.

  19. From Generality to Mastery: Composer-Style Symbolic Music Generation via Large-Scale Pre-training

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A two-stage pre-train-then-fine-tune transformer with style adapters improves composer-style symbolic piano generation over from-scratch training and the NotaGen baseline.

  20. Workflow-Based Evaluation of Music Generation Systems

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.

  21. Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment

    cs.SD 2025-05 conditional novelty 5.0 of 10

    By mutating captions and re-ranking MIDI tokens with CLAP and key-consistency scores, the method improves caption agreement and key matching in text-to-MIDI generation at inference time.

  22. VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features

    cs.SD 2024-12 conditional novelty 5.0 of 10

    An adaptation of MusicGen that adds semantic conditioning from CLIP global features and rhythmic conditioning from CLIP local inter-frame similarity, trained in two stages, generates video-aligned background music.

  23. FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration

    cs.SD 2026-07 reject novelty 4.0 of 10

    FlowSonic combines deterministic rectified-flow inversion, cached cross-attention injection, and a 'seeded' third-order Adams-Bashforth solver to report better timbre and genre edits on small datasets.

  24. Genre Controlled Music Generation via Activation Steering

    cs.SD 2025-06 conditional novelty 4.0 of 10

    Activation steering with linear probe weights on MusicGen's residual stream shifts generated music between genres at inference time, outperforming text prompting in CLAP and listener preference but with incomplete reporting.

  25. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.

  26. Improving Controllability and Editability for Pretrained Text-to-Music Generation Models

    cs.SD 2024-11 conditional novelty 2.0 of 10

    A thesis compilation presenting three complementary approaches to improving editing and control of pretrained text-to-music models, with Instruct-MusicGen demonstrating the strongest stem-level editing results.

Pith tools