Pith. sign in

REVIEW 3 cited by

Content-based Controls For Music Large Language Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.17162 v3 pith:QU6C5LX7 submitted 2023-10-26 cs.AI cs.SDeess.AS

classification cs.AIcs.SDeess.AS
keywords musiccontent-basedcontrolsgenerationmodelsaudiocontrollanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent years have witnessed a rapid growth of large-scale language models in the domain of music audio. Such models enable end-to-end generation of higher-quality music, and some allow conditioned generation using text descriptions. However, the control power of text controls on music is intrinsically limited, as they can only describe music indirectly through meta-data (such as singers and instruments) or high-level representations (such as genre and emotion). We aim to further equip the models with direct and content-based controls on innate music languages such as pitch, chords and drum track. To this end, we contribute Coco-Mulla, a content-based control method for music large language modeling. It uses a parameter-efficient fine-tuning (PEFT) method tailored for Transformer-based audio models. Experiments show that our approach achieved high-quality music generation with low-resource semi-supervised learning, tuning with less than 4% parameters compared to the original model and training on a small dataset with fewer than 300 songs. Moreover, our approach enables effective content-based controls, and we illustrate the control power via chords and rhythms, two of the most salient features of music audio. Furthermore, we show that by combining content-based controls and text descriptions, our system achieves flexible music variation generation and arrangement. Our source codes and demos are available online.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Expotion fine-tunes MusicGen with facial-expression and body-motion features plus text, and reports improved music quality and video-music alignment over text-only and video-only baselines.

  2. MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A lightweight adapter that adds rotary position embeddings to decoupled cross-attention enables efficient time-varying style control and audio inpainting/outpainting for text-to-music diffusion Transformers.

  3. FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration

    cs.SD 2026-07 reject novelty 4.0 of 10

    FlowSonic combines deterministic rectified-flow inversion, cached cross-attention injection, and a 'seeded' third-order Adams-Bashforth solver to report better timbre and genre edits on small datasets.

Pith tools