REVIEW 6 cited by
MusicRL: Aligning Music Generation to Human Preferences
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose MusicRL, the first music generation system finetuned from human feedback. Appreciation of text-to-music models is particularly subjective since the concept of musicality as well as the specific intention behind a caption are user-dependent (e.g. a caption such as "upbeat work-out music" can map to a retro guitar solo or a techno pop beat). Not only this makes supervised training of such models challenging, but it also calls for integrating continuous human feedback in their post-deployment finetuning. MusicRL is a pretrained autoregressive MusicLM (Agostinelli et al., 2023) model of discrete audio tokens finetuned with reinforcement learning to maximise sequence-level rewards. We design reward functions related specifically to text-adherence and audio quality with the help from selected raters, and use those to finetune MusicLM into MusicRL-R. We deploy MusicLM to users and collect a substantial dataset comprising 300,000 pairwise preferences. Using Reinforcement Learning from Human Feedback (RLHF), we train MusicRL-U, the first text-to-music model that incorporates human feedback at scale. Human evaluations show that both MusicRL-R and MusicRL-U are preferred to the baseline. Ultimately, MusicRL-RU combines the two approaches and results in the best model according to human raters. Ablation studies shed light on the musical attributes influencing human preferences, indicating that text adherence and quality only account for a part of it. This underscores the prevalence of subjectivity in musical appreciation and calls for further involvement of human listeners in the finetuning of music generation models.
Forward citations
Cited by 6 Pith papers
-
SpectroStream: A Versatile Neural Codec for General Audio
SpectroStream, a 2D time-frequency neural codec, reconstructs 48 kHz stereo music at 4-16 kbps with better ViSQOL and subjective quality than DAC.
-
SMART: Tuning a symbolic music generation system with an audio domain aesthetic reward
SMART uses an audio aesthetic reward to fine-tune a symbolic MIDI piano model, raising human enjoyment ratings by about 1.2 points in a small listening study while over-optimization collapses output diversity.
-
Exploring listeners' perceptions of AI-generated and human-composed music for functional emotional applications
Preference and perceived emotional efficacy dissociate for AI-generated versus human-composed music, with listeners preferring AI tracks but crediting human tracks with stronger functional emotion elicitation.
-
Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment
By mutating captions and re-ranking MIDI tokens with CLAP and key-consistency scores, the method improves caption agreement and key matching in text-to-MIDI generation at inference time.
-
Toward Rich Video Human-Motion2D Generation
A new 150K-video 2D skeleton dataset with text captions and a diffusion model for single- and double-character motion generation, though the claimed FID-rewarded RL training is misrepresented.
-
WavChat: A Survey of Spoken Dialogue Models
WavChat categorizes spoken dialogue models into cascaded and end-to-end paradigms and surveys speech representations, training strategies, streaming, duplex interaction, datasets, and evaluation benchmarks.
Discussion (0). Continue with ORCID to comment.