Pith. sign in

REVIEW 1 cited by

Optical Music Recognition with Convolutional Sequence-to-Sequence Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1707.04877 v1 pith:JNBN5XQ7 submitted 2017-07-16 cs.CV cs.IRcs.SD

classification cs.CVcs.IRcs.SD
keywords learningmodelsmusicaccuracyavailabledatadeepmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Optical Music Recognition (OMR) is an important technology within Music Information Retrieval. Deep learning models show promising results on OMR tasks, but symbol-level annotated data sets of sufficient size to train such models are not available and difficult to develop. We present a deep learning architecture called a Convolutional Sequence-to-Sequence model to both move towards an end-to-end trainable OMR pipeline, and apply a learning process that trains on full sentences of sheet music instead of individually labeled symbols. The model is trained and evaluated on a human generated data set, with various image augmentations based on real-world scenarios. This data set is the first publicly available set in OMR research with sufficient size to train and evaluate deep learning models. With the introduced augmentations a pitch recognition accuracy of 81% and a duration accuracy of 94% is achieved, resulting in a note level accuracy of 80%. Finally, the model is compared to commercially available methods, showing a large improvements over these applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synthesising Handwritten Music with GANs: A Comprehensive Evaluation of CycleWGAN, ProGAN, and DCGAN

    cs.CV 2024-11 conditional novelty 4.0 of 10

    CycleWGAN, a CycleGAN variant with Wasserstein loss, beats DCGAN and ProGAN at generating handwritten music images, with FID 41.87, IS 2.29, and KID 0.05.

Pith tools