Pith. sign in

REVIEW 3 cited by

Music FaderNets: Controllable Music Generation Based On High-Level Features via Low-Level Feature Modelling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.15474 v1 pith:S46K2KCE submitted 2020-07-29 eess.AS cs.LGcs.SDstat.ML

classification eess.AScs.LGcs.SDstat.ML
keywords featurehigh-levellow-levelattributesmusicrepresentationsarousalframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

High-level musical qualities (such as emotion) are often abstract, subjective, and hard to quantify. Given these difficulties, it is not easy to learn good feature representations with supervised learning techniques, either because of the insufficiency of labels, or the subjectiveness (and hence large variance) in human-annotated labels. In this paper, we present a framework that can learn high-level feature representations with a limited amount of data, by first modelling their corresponding quantifiable low-level attributes. We refer to our proposed framework as Music FaderNets, which is inspired by the fact that low-level attributes can be continuously manipulated by separate "sliding faders" through feature disentanglement and latent regularization techniques. High-level features are then inferred from the low-level representations through semi-supervised clustering using Gaussian Mixture Variational Autoencoders (GM-VAEs). Using arousal as an example of a high-level feature, we show that the "faders" of our model are disentangled and change linearly w.r.t. the modelled low-level attributes of the generated output music. Furthermore, we demonstrate that the model successfully learns the intrinsic relationship between arousal and its corresponding low-level attributes (rhythm and note density), with only 1% of the training set being labelled. Finally, using the learnt high-level feature representations, we explore the application of our framework in style transfer tasks across different arousal states. The effectiveness of this approach is verified through a subjective listening test.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revisiting Your Memory: Reconstruction of Affect-Contextualized Memory via EEG-guided Audiovisual Generation

    cs.AI 2024-11 conditional novelty 7.0 of 10

    A nine-participant proof-of-concept shows that EEG during memory recall can be decoded into positive/neutral/negative trajectories, and those trajectories can steer text-to-music and text-to-image generation into pers...

  2. AffectMachine-Pop: A controllable expert system for real-time pop music generation

    cs.HC 2025-06 conditional novelty 5.0 of 10

    A rule-based system generates retro-pop music at target levels of arousal and valence, validated by a listening study with high correspondence between target and perceived ratings.

  3. Improving Controllability and Editability for Pretrained Text-to-Music Generation Models

    cs.SD 2024-11 conditional novelty 2.0 of 10

    A thesis compilation presenting three complementary approaches to improving editing and control of pretrained text-to-music models, with Instruct-MusicGen demonstrating the strongest stem-level editing results.

Pith tools