Pith. sign in

REVIEW 7 cited by

NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training Paradigms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.18008 v5 pith:QT4WDBKW submitted 2025-02-25 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords musicgenerationnotagensymbolicmodelsparadigmsadvancingclamp-dpo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce NotaGen, a symbolic music generation model aiming to explore the potential of producing high-quality classical sheet music. Inspired by the success of Large Language Models (LLMs), NotaGen adopts pre-training, fine-tuning, and reinforcement learning paradigms (henceforth referred to as the LLM training paradigms). It is pre-trained on 1.6M pieces of music in ABC notation, and then fine-tuned on approximately 9K high-quality classical compositions conditioned on "period-composer-instrumentation" prompts. For reinforcement learning, we propose the CLaMP-DPO method, which further enhances generation quality and controllability without requiring human annotations or predefined rewards. Our experiments demonstrate the efficacy of CLaMP-DPO in symbolic music generation models with different architectures and encoding schemes. Furthermore, subjective A/B tests show that NotaGen outperforms baseline models against human compositions, greatly advancing musical aesthetics in symbolic music generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Text2Score: Generating Sheet Music From Textual Prompts

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    Text2Score turns text prompts into sheet music by having an LLM produce a bar-wise structural plan and a hierarchical decoder write ABC notation from that plan.

  2. Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music

    cs.SD 2025-08 conditional novelty 6.0 of 10

    Amadeus generates symbolic music by autoregressively predicting note-level latents and decoding their attributes in parallel with a masked discrete diffusion model, yielding faster and more controllable generation tha...

  3. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.

  4. Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation

    cs.SD 2025-08 unverdicted novelty 5.0 of 10

    A bar-level symbolic-score song generator (BACH) is claimed to beat published systems and commercial Suno on human-rated quality, duration, and efficiency, but the supporting full text is corrupted and unverifiable.

  5. Large Language Models' Internal Perception of Symbolic Music

    cs.CL 2025-07 conditional novelty 5.0 of 10

    LLM-generated MIDI data carries enough genre and style signal to train above-chance classifiers and melody predictors, but far less than real music data.

  6. WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation

    cs.SD 2025-09 reject novelty 4.0 of 10

    An open multi-agent system that orchestrates specialized music models for understanding, composition, and synthesis, with local or hosted deployment.

  7. CoComposer: LLM Multi-agent Collaborative Music Composition

    cs.SD 2025-08 conditional novelty 4.0 of 10

    A five-agent LLM system for ABC-notation composition scores modestly higher than ComposerX and a single LLM on an automated aesthetic model, but no error bars or significance tests are reported.

Pith tools