Pith. sign in

REVIEW 4 cited by

End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.12105 v4 pith:26ZVLZMS submitted 2024-05-20 cs.CV

classification cs.CV
keywords musicdataend-to-endfull-pagesystemapproachapproachescommercial
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Optical Music Recognition (OMR) has made significant progress since its inception, with various approaches now capable of accurately transcribing music scores into digital formats. Despite these advancements, most so-called end-to-end OMR approaches still rely on multi-stage processing pipelines for transcribing full-page score images, which entails challenges such as the need for dedicated layout analysis and specific annotated data, thereby limiting the general applicability of such methods. In this paper, we present the first truly end-to-end approach for page-level OMR in complex layouts. Our system, which combines convolutional layers with autoregressive Transformers, processes an entire music score page and outputs a complete transcription in a music encoding format. This is made possible by both the architecture and the training procedure, which utilizes curriculum learning through incremental synthetic data generation. We evaluate the proposed system using pianoform corpora, which is one of the most complex sources in the OMR literature. This evaluation is conducted first in a controlled scenario with synthetic data, and subsequently against two real-world corpora of varying conditions. Our approach is compared with leading commercial OMR software. The results demonstrate that our system not only successfully transcribes full-page music scores but also outperforms the commercial tool in both zero-shot settings and after fine-tuning with the target domain, representing a significant contribution to the field of OMR.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Dataset and Benchmark for Optical Music Recognition of String Quartet Scores

    cs.CV 2026-08 conditional novelty 7.0 of 10

    OSSQ-OMR is the first multi-part OMR dataset, pairing 116 aligned string quartet scores with three symbolic encodings, and its benchmark shows encoding choice matters more than architecture.

  2. LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding

    cs.CV 2026-07 conditional novelty 6.5 of 10

    System-by-system autoregressive OMR with text-aware ABC transcription outperforms prior neural and rule-based systems and boosts VLM sheet-music QA.

  3. MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A synthetic music sheet QA dataset and a LoRA-fine-tuned Phi-3 model show large accuracy gains on OMR and chord tasks, but only within the synthetic distribution.

  4. Sheet Music Benchmark: Standardized Optical Music Recognition Evaluation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A public 685 page sheet music benchmark (SMB) and a category-level edit distance metric (OMR-NED) for Optical Music Recognition are introduced, with region-level baselines from a transformer model.

Pith tools