{"id":"4ee204e2-a62a-4537-93e3-1e176f7f67a1","arxiv_id":"1908.03608","paper_version":3,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper defines OMR as computationally reading music notation in documents and classifies applications into metadata extraction, search, replayability, and structured encoding.","lead":"This tutorial proposes a precise definition of Optical Music Recognition and a four-part taxonomy of its applications. It aims to give newcomers, librarians, and researchers a shared vocabulary for a field that has long lacked one.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The four-application taxonomy is not entailed by Definition 1 or the A/B distinction: metadata extraction and search are added ad hoc, so the central systematization claim is under-supported.","rationale":"The reader's ACCEPT verdict is reasonable for a tutorial that organizes a fragmented field, and the paper is honest that its novelty is systematization rather than new experimental results. However, the central claim is not just that a taxonomy is proposed; it is that the taxonomy follows from a robust definition and a fundamental A/B distinction. My reading of Sections IV and VI-B finds that only two of the four categories are consequences of the A/B distinction; the other two are added to accommodate applications that do not output notation or semantics. The paper even acknowledges the replayability/structured-encoding boundary is vague, which is the point where the A/B grounding is clearest; the boundaries around metadata extraction and search are much less constrained. This does not make the paper valueless, but it does mean the headline systematization claim is internally under-supported. A CONDITIONAL verdict is appropriate: the taxonomy should either be presented as a pragmatic proposal explicitly independent of the A/B derivation, or the authors should supply a principled rule for why these four categories, and only these, emerge from the definition. The proposed annotation test would settle whether the categories are natural in practice, and the derivability check would settle whether the definition actually entails them. Agreement with the reader is partial because the reader identified the taxonomy's lack of empirical validation and the A/B boundary as the weak spot, which is related, but the reader did not pinpoint the internal non-derivation of the two lower-comprehension categories.","tokens_in":28884,"tokens_out":7585,"duration_ms":90497,"concrete_test":"Test coverage and derivability separately. First, take the application-oriented papers cited in Section VI-B (at least [9], [118], [10], [66], [1], [100], [35], [102], [16], [91], [116], [33], [5], [83], [39], [41], [136]) and have two independent annotators, blind to Fig. 13's ordering, assign each paper to one or more of the four categories; if Cohen's kappa is below 0.6 or more than 20% of papers straddle or fit none, the taxonomy is not a natural grouping. Second, write out for each category which of the two A/B outputs (notation recovery or semantic recovery) is strictly required; if Document Metadata Extraction and Search require neither, the taxonomy is not derivable from Definition 1 and Section IV.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Definition 1 defines OMR as computationally reading music notation, and Section IV frames reading as two prongs: (A) recovering music notation and (B) recovering musical semantics. Yet Section VI-B introduces four application categories. Only Replayability (Def. 4) and Structured Encoding (Def. 5) map cleanly onto B and A, respectively. Document Metadata Extraction (Def. 2) outputs classification or regression answers about the document (e.g., writer identity, whether an image contains music), and Search (Def. 3) outputs relevance scores; neither is a recovered notation layer or a recovered semantic layer. The paper itself says it 'needed to broaden the scope' beyond the two prongs before proposing the four categories. Thus the central claim, as the reader states it, that the A/B distinction grounds exactly four application categories with shared evaluation protocols, does not follow from the definition. The monotonic 'level of comprehension' ordering in Fig. 13 is also asserted rather than derived: search can be implemented by image hashing or query-by-example (e.g., [100], [35]) at comprehension levels that need not fall strictly between metadata extraction and replayability. This is not an external empirical gap; it is an internal structural gap in the argument for the taxonomy.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a tutorial and survey of Optical Music Recognition (OMR). It proposes a formal definition: OMR is a field of research that investigates how to computationally read music notation in documents (Definition 1). The paper then analyzes OMR as the process of inverting music encoding, distinguishing between recovering the music notation itself (stage A) and recovering musical semantics (stage B). On this basis, it proposes a taxonomy of OMR inputs, system architectures, and, most centrally, four application categories with increasing levels of comprehension: document metadata extraction, search, replayability, and structured encoding. The paper also reviews traditional pipeline-based approaches and deep learning alternatives, and it concludes with a list of open issues. The contributions are definitional and organizational rather than experimental, with a substantial curated bibliography provided as supplementary material.","tokens_in":29136,"tokens_out":7327,"duration_ms":76682,"significance":"If the definition and taxonomy are adopted by the community, this paper would materially improve the clarity, comparability, and accessibility of OMR research. The definition is carefully argued to capture the field's breadth while excluding non-OMR tasks, and the taxonomy provides a practical vocabulary for researchers and stakeholders. The paper's strength lies in its systematic organization, its concrete examples, and the extensible curated bibliography. The taxonomy is proposed pragmatically rather than proven from first principles, which is appropriate for a tutorial; this does not undermine the core definitional contribution but does leave room for clarification in the presentation.","major_comments":[{"comment":"The four application categories are not derived from Definition 1 or from the A/B distinction presented in Section IV. The paper itself states that it \"needed to broaden the scope\" beyond the two prongs of Replayability and Structured Encoding to also include Search and Document Metadata Extraction. Moreover, the \"level of comprehension\" ordering in Fig. 13 is asserted rather than operationalized. Because the taxonomy is one of the paper's three headline contributions, the authors should make explicit that this is a practical systematization rather than a logical consequence of the definition, and they should discuss boundary cases such as query-by-example search, which may not require comprehension strictly between metadata extraction and replayability. A short paragraph acknowledging the heuristic nature of the ordering and pointing to possible alternative groupings would address this concern.","section":"Section VI-B / Fig. 13"}],"minor_comments":[{"comment":"The abstract refers to the paper as a \"tutorial,\" while the opening of the full text says \"In this work\"; the wording should be aligned.","section":"Abstract / Section I"},{"comment":"The labels in Fig. 13 are difficult to parse in the current rendering; please ensure the figure is typeset clearly so that the four category names and the \"Level of Comprehension\" axis are legible.","section":"Fig. 13"},{"comment":"Definition 3 says a musical query \"must convey musical semantics,\" but the same section mentions image queries (query-by-example); please clarify whether image queries are interpreted semantically or are treated as raw visual patterns.","section":"Section VI-B2"},{"comment":"The paper frequently cites the authors' own prior work as illustrations; this is acceptable, but adding an explicit sentence that these works are used as examples of existing research rather than as evidence for the taxonomy would help the reader.","section":"Section VI-B4"}],"recommendation":"minor_revision","confidential_remarks":"This is a well-written and useful tutorial. The central definition is sound and the taxonomy, while heuristic, is likely to be influential. The only substantive point I want addressed is the framing of the taxonomy's status; this requires a modest textual revision, not new experiments. The paper is within the scope of the journal and would be a valuable contribution once the authors clarify the nature of the systematization."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is the closest thing OMR has to a state-of-the-field paper, and its four-way application taxonomy will be cited for years. But the taxonomy is more of a useful organizing proposal than a consequence of the paper's definition, and the paper is honest enough to show its own seams.\n\nWhat's actually new: Definition 1 (OMR as a field that investigates how to computationally read music notation in documents) is reasonable but not the main event. The genuinely new piece is Section VI-B: Document Metadata Extraction, Search, Replayability, Structured Encoding, ordered by level of comprehension. That framing is concrete enough to give researchers a shared vocabulary and, importantly, a way to talk about evaluation. The discussion of why OMR evaluation has been stuck is the best short account I know. The appendix bibliography is a real service to the community. The deep learning update over the traditional pipeline is competent and balanced.\n\nCredit where it is earned: the authors cite prior reviews and primary sources extensively, including work that challenges the field's self-image. They also flag their own limitations: they openly say the replayability/structured encoding boundary is vague, and they do not pretend the taxonomy is empirically validated. Self-citations appear often, but they point to concrete systems the reader can check, not to evidence for the central claim.\n\nSoft spots: the stress-test note is right. The four categories are not entailed by Definition 1 or by the A/B distinction (recovering notation vs. recovering semantics). Metadata extraction and Search are added because real use-cases exist; the paper says so. The monotonic \"level of comprehension\" ordering is asserted rather than derived, and search via image hashing or query-by-example can require arbitrarily little comprehension. That is a structural gap, but it is the kind of gap that is normal in a taxonomy paper: the categories earn their keep through usefulness, not proof. I do not think it sinks the contribution.\n\nThe citation pattern looks solid, and the paper ships a reproducible bibliography alongside the text. This is a paper for newcomers to OMR, for reviewers who need a shared vocabulary, and for MIR/Digital Humanities people who want to know what OMR can and cannot do. It deserves a serious referee and publication, with only light pressure to acknowledge the ad hoc nature of the category scheme in Section VI-B.","headline":"A very good state-of-the-field paper whose four-part application taxonomy is genuinely useful, even though it is not entailed by the paper's own definition.","tokens_in":29587,"tokens_out":2280,"would_cite":true,"duration_ms":25122,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper defines Optical Music Recognition as a research field that computationally reads music notation in documents, and organizes its applications into four levels of comprehension.","keywords":["optical music recognition","music notation","document analysis","taxonomy of applications","music information retrieval","deep learning","evaluation"],"falsifier":"Survey OMR papers from the past five years in the field's main venues and try to assign each paper's stated goal to exactly one of the four application categories; if a substantial share cannot be assigned uniquely, or if papers inside a single category report no common evaluation metric, then the taxonomy's promised shared evaluation protocols fail in practice.","tokens_in":28722,"feed_emoji":"🎼","tokens_out":6040,"duration_ms":60686,"temperature":0.7,"pith_summary":"After five decades of scattered contributions, Optical Music Recognition (OMR) still lacked a shared definition and a way to compare results. This tutorial supplies both: it defines OMR as a research field that investigates how to computationally read music notation in documents, and it derives a taxonomy of applications from the process of inverting music encoding. The four application classes—document metadata extraction, search, replayability, and structured encoding—demand increasing levels of comprehension and each comes with its own natural evaluation strategy. If the taxonomy is adopted, the vague question 'Does OMR work?' becomes answerable only in the form 'Does OMR work for application X?', which is exactly what the authors argue the field needs. The payoff is that otherwise incomparable systems can finally be measured against one another within each class.","feed_headline":"Four rungs define what music-reading software can do","feed_subtitle":"A new definition and taxonomy split 'does OMR work?' into metadata, search, playback, and full-encoding tasks.","key_machinery":"The load-bearing mechanism is the inverse-encoding model: music is conceptualized as notes in time, engraved with a notation system, and embodied in a document, and OMR runs this process backwards. The inversion splits into two prongs—recovering the notation itself and recovering the musical semantics—and the paper's four application categories are generated by how far back along this chain a system must go and how much comprehension is required. The level-of-comprehension ladder, from partial understanding for metadata extraction to complete understanding for structured encoding, is the organizing device that ties output representations to evaluation strategies.","core_discovery":"The central claim is that OMR is not a single task or process but a research field, defined as investigating how to computationally read music notation in documents. Reading itself has two distinct targets: recovering the music notation as laid out on the page, and recovering the musical semantics—pitches, velocities, onsets, and durations—encoded by that notation. Because these targets require different outputs and tolerate different errors, the paper argues OMR has at least four natural application classes: metadata extraction, search, replayability, and structured encoding, arranged by increasing comprehension. The taxonomy's practical point is that evaluation can be shared within each class: classification metrics for metadata, information-retrieval metrics for search, sequence comparison for replayability, and, once a formal model of notation exists, an intrinsic metric for structured encoding.","pith_inferences":["The taxonomy implies that a single OMR system cannot be judged by one number; a system can be excellent for replayability yet useless for structured encoding, so the field should report results per application class.","If search is defined by semantic queries rather than full transcription, then OMR systems for search can be trained to ignore notation details that do not affect retrieval, which may make them more error-tolerant than transcription-first pipelines.","A natural next step would be a benchmark suite with one shared dataset and metric per application class; its absence would itself be evidence of the fragmentation the paper describes.","Adopting the notation-versus-semantics distinction would let digital-musicology studies that only need musical semantics proceed without waiting for full structured encoding, since MIDI-level output suffices for many corpus-wide questions."],"forward_implications":["The question 'Does OMR work?' becomes well-posed only when followed by a target application: metadata extraction, search, replayability, or structured encoding.","Search applications can adopt standard information-retrieval metrics such as precision, recall, and mean average precision, since they output ranked matches to a musical query.","Replayability systems can be evaluated by comparing pitch-onset-duration sequences, reusing symbolic melodic-similarity metrics from music information retrieval.","Metadata extraction reduces to classification or regression and can be scored with accuracy or mean squared error.","Structured encoding remains the hard case: without a formal model of music notation and an intrinsic edit distance between scores, no meaningful evaluation metric exists."],"supporting_citations":[{"why":"State-of-the-art review of OMR that the paper uses as the baseline technical pipeline and as evidence that prior reviews avoided defining application taxonomies.","marker":"[133]"},{"why":"Byrd and Simonsen's analysis of OMR difficulty and evaluation, including the notational complexity categories the paper adapts for its input taxonomy.","marker":"[29]"},{"why":"Miyao and Haralick's ground-truth-format work, cited as the source for the two broad application areas (replayability and structured encoding) the paper folds into its four-category taxonomy.","marker":"[105]"},{"why":"Bainbridge and Bell's formulation of the OMR challenge and the traditional pipeline stages that the paper updates for deep learning.","marker":"[7]"},{"why":"Blostein and Baird's critical survey, used to anchor the relationship between OMR and other graphics recognition and to document the limits of grammar-based methods.","marker":"[22]"},{"why":"Hajič's argument for intrinsic evaluation of OMR, which the paper draws on for the structured-encoding evaluation discussion.","marker":"[81]"}],"fun_headline_variants":["Four application classes define the OMR field","OMR taxonomy: metadata, search, playback, encoding","Two targets: notation recovery vs. semantics recovery","Optical Music Recognition: a field with four tasks","OMR's core: reading notation and meaning from sheet"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The taxonomy rests on the assumption that reading a music document splits cleanly into recovering the notation and recovering the semantics, and that every real OMR application falls into exactly one of the four categories with a shared evaluation protocol.","fun_headline_variants_meta":{"raw":{"variants":["Four application classes define the OMR field","OMR taxonomy: metadata, search, playback, encoding","Two targets: notation recovery vs. semantics recovery","Optical Music Recognition: a field with four tasks","OMR's core: reading notation and meaning from sheet"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000222,"raw_usage":{"total_tokens":1417,"prompt_tokens":871,"completion_tokens":546,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":487,"completion_tokens_details":{"reasoning_tokens":471}},"tokens_in":487,"tokens_out":546,"duration_ms":6739,"temperature":1.0,"reasoning_tokens":471,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:41:25.388221+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Survey OMR papers from the past five years in the field's main venues and try to assign each paper's stated goal to exactly one of the four application categories; if a substantial share cannot be assigned uniquely, or if papers inside a single category report no common evaluation metric, then the taxonomy's promised shared evaluation protocols fail in practice.","supporting_citations":[{"cited_title":"Marcal, Carlos Guedes, and Jamie dos Santos Cardoso","cited_arxiv_id":null,"evidence_quote":"State-of-the-art review of OMR that the paper uses as the baseline technical pipeline and as evidence that prior reviews avoided defining application taxonomies."},{"cited_title":"Format of ground truth data used in the evaluation of the results of an optical music recognition system","cited_arxiv_id":null,"evidence_quote":"Miyao and Haralick's ground-truth-format work, cited as the source for the two broad application areas (replayability and structured encoding) the paper folds into its four-category taxonomy."}],"review_version":1}