Pith. sign in

REVIEW 2 cited by

Jointist: Joint Learning for Multi-instrument Transcription and Its Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.10805 v2 pith:OPC5ADAE submitted 2022-06-22 cs.SD cs.AIcs.LGeess.AS

classification cs.SDcs.AIcs.LGeess.AS
keywords transcriptionmodelmodulemulti-instrumentinstrumentjointistconsistsexperiment
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we introduce Jointist, an instrument-aware multi-instrument framework that is capable of transcribing, recognizing, and separating multiple musical instruments from an audio clip. Jointist consists of the instrument recognition module that conditions the other modules: the transcription module that outputs instrument-specific piano rolls, and the source separation module that utilizes instrument information and transcription results. The instrument conditioning is designed for an explicit multi-instrument functionality while the connection between the transcription and source separation modules is for better transcription performance. Our challenging problem formulation makes the model highly useful in the real world given that modern popular music typically consists of multiple instruments. However, its novelty necessitates a new perspective on how to evaluate such a model. During the experiment, we assess the model from various aspects, providing a new evaluation perspective for multi-instrument transcription. We also argue that transcription models can be utilized as a preprocessing module for other music analysis tasks. In the experiment on several downstream tasks, the symbolic representation provided by our transcription model turned out to be helpful to spectrograms in solving downbeat detection, chord recognition, and key estimation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Classical Guitar Duet Separation using GuitarDuets -- a Dataset of Real and Synthesized Guitar Recordings

    eess.AS 2025-07 conditional novelty 6.0 of 10

    GuitarDuets provides about three hours of real and synthetic classical guitar duet audio, and experiments show that combining both data types improves Demucs-based separation of similar-timbre guitars.

  2. Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio

    eess.AS 2025-01 conditional novelty 5.0 of 10

    A MAML-trained CRNN outperforms supervised and transfer-learning baselines for low-resource tabla stroke transcription, and two simple scoring methods identify tala from transcribed strokes.

Pith tools