Pith. sign in

REVIEW 1 cited by

Multi-Instrumentalist Net: Unsupervised Generation of Music from Body Movements

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.03478 v1 pith:SKR6AXH5 submitted 2020-12-07 cs.SD cs.CVeess.AS

classification cs.SDcs.CVeess.AS
keywords musicbodyinstrumentslatentmovementspipelineinstrumentaudio
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a novel system that takes as an input body movements of a musician playing a musical instrument and generates music in an unsupervised setting. Learning to generate multi-instrumental music from videos without labeling the instruments is a challenging problem. To achieve the transformation, we built a pipeline named 'Multi-instrumentalistNet' (MI Net). At its base, the pipeline learns a discrete latent representation of various instruments music from log-spectrogram using a Vector Quantized Variational Autoencoder (VQ-VAE) with multi-band residual blocks. The pipeline is then trained along with an autoregressive prior conditioned on the musician's body keypoints movements encoded by a recurrent neural network. Joint training of the prior with the body movements encoder succeeds in the disentanglement of the music into latent features indicating the musical components and the instrumental features. The latent space results in distributions that are clustered into distinct instruments from which new music can be generated. Furthermore, the VQ-VAE architecture supports detailed music generation with additional conditioning. We show that a Midi can further condition the latent space such that the pipeline will generate the exact content of the music being played by the instrument in the video. We evaluate MI Net on two datasets containing videos of 13 instruments and obtain generated music of reasonable audio quality, easily associated with the corresponding instrument, and consistent with the music audio content.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Controllable Video-to-Music Generation with Multiple Time-Varying Conditions

    cs.MM 2025-07 reject novelty 6.0 of 10

    A two-stage video-to-music model with four time-varying controls (rhythm, melody, intensity, emotion) claims better controllability and alignment than prior V2M systems.

Pith tools