{"id":"41e01c03-c3a4-4aef-a1bf-26a0d23bdd4b","arxiv_id":"2411.15741","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A workshop proceedings volume compiling 15 short OMR papers; the most substantive are a multitask transformer for sheet music information retrieval, a faster YOLO-based layout detector, and a new tree-based evaluation format.","lead":"These are the proceedings of the 6th International Workshop on Reading Music Systems (WoRMS 2024), a collection of 15 short papers on optical music recognition, sheet music layout analysis, and evaluation formats. The standout contributions are a multitask sheet-music transformer (SMIReT), a YOLOv8m layout detector with released models, and a proposed tree-based evaluation standard (MTN).","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SMIReT's 'successful multitask' claim lacks single-task baselines and an ablation of the PRIMENS curriculum; without these, Table I cannot establish that task unification works.","rationale":"This submission is a workshop proceedings volume, not a single-research-paper claim, and the reader's UNVERDICTED verdict is the appropriate disposition: no individual paper rises to ACCEPT-level support for a transformative claim, and none contains a load-bearing error warranting REJECT. My stress-test focuses on the volume's strongest visible claim, the SMIReT multitask transformer. The reader's weakest_assumption correctly flags the synthetic-to-real transfer from PRIMENS as untested, but I see an even more load-bearing gap: the paper never compares SMIReT against single-task baselines on the same data. Without such baselines, 'all SMIR tasks successfully' is uninterpretable; the reported degradation in full-parsing text CER and the poor selective OMR SER could stem from multitasking itself, from the unified vocabulary, or from the small corpus, and the paper cannot discriminate. My proposed test—a no-pretraining ablation plus a single-task SMT baseline—would settle both whether the curriculum matters and whether multitasking hurts. This concern reinforces the reader's decision to leave the volume UNVERDICTED; it does not move the verdict, so I set verdict_should_be to UNCHANGED. I credit the paper for its honest discussion of limitations, the YOLO paper for released models and reproducible evaluation, and the MTN paper for a shipped evaluation toolkit; these are real independent supports even though they do not validate SMIReT's central claim.","tokens_in":50965,"tokens_out":4211,"duration_ms":38897,"concrete_test":"Retrain SMIReT on the MOTTECTA training split with the PRIMENS curriculum pretraining step removed (same architecture, hyperparameters, and task prompts) and compare all seven Table I metrics against the published values. Also train a single-task SMT model on the same OMR-only data and compare OMR SER. If the no-pretraining run matches the published OMR SER within roughly 1 point and the single-task SMT is substantially better (e.g., SER below 4), then the curriculum is not load-bearing and the multitask success claim lacks a baseline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that SMIReT can perform all SMIR tasks successfully (Section IV-B) rests entirely on Table I, yet the paper provides no baseline for any of the seven reported metrics. The authors themselves note that region detection IoU 70.23 is 'notably below state-of-the-art Layout Analysis techniques' and that selective OMR SER 41.55 shows the model 'struggles' with pixel-wise context, so 'successful' is already qualified. But the missing comparison is more fundamental: without a single-task SMT baseline trained on the same MOTTECTA split, we cannot attribute the reported OMR SER 5.92 to the multitask prompting and unified vocabulary rather than to the underlying SMT architecture, nor can we judge whether multitasking causes the 51.78% relative CER degradation between isolated OCR (10.08) and full parsing (15.30). The load-bearing premise identified by the reader—that curriculum pretraining on synthetic PRIMENS incipits transfers to real 17th-century MOTTECTA pages—is untested: Section III-C describes the curriculum only verbally, with no ablation isolating PRIMENS pretraining from target fine-tuning. If the synthetic pretraining were removed and performance collapsed, transfer would be load-bearing; if it did not change, the claim would reduce to ordinary fine-tuning on 297 pages, making 'successful multitask' an overstatement in the absence of any comparative baseline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This submission is the proceedings of the 6th International Workshop on Reading Music Systems (WoRMS 2024), consisting of a set of short contributed papers on optical music recognition and music-reading systems. The topics range from an exploratory study of multimodal LLMs on score images, to a multitask transformer for sheet-music information retrieval (SMIReT), a graph-neural-network approach to semantic reconstruction, a YOLO-based staff layout detector, a tree-based evaluation representation (MTN), and several system and tool reports. The most defensible archival contribution is the YOLO paper, which harmonizes a 7,013-image corpus, evaluates on one in-domain and three out-of-domain settings, and releases models and scripts; it demonstrates comparable or better accuracy than Faster R-CNN at roughly 26x lower latency and 4.4x lower memory. The most ambitious claim, that SMIReT can perform all proposed SMIR tasks successfully, is made on the basis of a single table of results without single-task baselines or an ablation of the synthetic pretraining stage.","tokens_in":51195,"tokens_out":6968,"duration_ms":67422,"significance":"If the stronger claims hold, the volume contributes practically useful resources to the OMR community. The YOLO paper stands out because it ships harmonized datasets, converted annotation scripts, trained models, and a reproducible evaluation protocol, which is exactly the kind of infrastructure the field needs; its finding that modern one-stage detection is substantially cheaper and often more accurate for staff layout is directly actionable. The MTN paper is also valuable as a community-level proposal for a common evaluation representation, and its public dataset and toolkit are a concrete step away from per-methodology metrics. The SMIReT paper addresses a timely question, namely whether an end-to-end multitask transformer can unify parsing, layout, and query tasks, but the current evidence is preliminary: the absence of baselines and ablations prevents the reader from attributing the reported numbers to the proposed multitask mechanism. Several other contributions are explicitly preliminary or experience reports; they are useful as workshop documentation but not as completed research claims.","major_comments":[{"comment":"The central claim that SMIReT can perform all SMIR tasks successfully rests on seven metrics reported for a single trained model, with no single-task baseline. In particular, the OMR SER of 5.92 cannot be attributed to the multitask prompting and unified vocabulary unless it is compared with a Sheet Music Transformer trained on the same MOTTECTA split with the same curriculum but without the added tasks and prompts. Without such a baseline, the 51.78% relative CER increase in full parsing (10.08 to 15.30) is also uninterpretable as an effect of task interference. Please add at least an SMT-only OMR baseline and, ideally, a sequential fine-tuning baseline.","section":"SMIReT paper, Section IV-B, Table I"},{"comment":"The load-bearing premise that curriculum pretraining on synthetically rendered PRIMENS incipits transfers to the 297-page MOTTECTA corpus is untested: the training procedure is described only verbally, and no ablation removes the PRIMENS stage. If the synthetic stage is essential, the paper should demonstrate this experimentally; if it is not, the curriculum-learning narrative should be removed or substantially weakened. As written, the claim that the approach is made viable by synthetic pretraining goes beyond the evidence presented.","section":"SMIReT paper, Section III-C and Section IV-A"},{"comment":"The empirical basis for the paper's title question is three score crops selected for simplicity and cultural familiarity, with outputs scored by the authors on a subjective three-level scale and prompts generated with the assistance of the models themselves. This design cannot support general conclusions about MLLM capability for music score reading; it is an anecdotal pilot. The authors should either enlarge the sample with a stratified selection and report quantitative agreement, or explicitly frame all conclusions as observations on three specific images.","section":"MLLM paper, Section II-B and Section III, Table I"},{"comment":"The paper's own data show that the k=20 candidate graph contains only 80% of ground-truth edges for the MUSCIMA++ measure-cut set and 91% for the DoReMi measure-cut set. Since the GNN only prunes edges, these numbers set an upper bound on achievable recall and should be stated whenever the reported MER values are interpreted. The conclusion that GNNs can effectively recover relations between musical primitives should be qualified by this ceiling, which also suggests that candidate-graph construction is itself a load-bearing component of the pipeline.","section":"GNN paper, Section VI"}],"minor_comments":[{"comment":"The displayed expression is not a coherent probability statement: as written, the left-hand side is equated to a sum over the vocabulary, which appears to be a typesetting omission of an argmax or of a distribution over tokens.","section":"SMIReT paper, Equation (1)"},{"comment":"There are several typos and spacing inconsistencies, including 'effort hat has been put', 'adress', and the inconsistent rendering of 'MOTTECTA'; these should be corrected in a revision.","section":"SMIReT paper, Section IV-A"},{"comment":"The sentence 'no previuos work has evaluated this scenario' is too strong unless it is restricted to the specific combination of models and tasks tested here; at minimum, the claim should be scoped and the typo 'previuos' fixed.","section":"MLLM paper, Section I"},{"comment":"The abstract contains the typo 'MeausreDetector', and the affiliation line contains 'Linquistics' for 'Linguistics'; these should be corrected.","section":"YOLO paper, Abstract and author affiliation"},{"comment":"The phrase 'nearest neighbors retrieved using K-means' conflates clustering with nearest-neighbor search; the intended procedure appears to be a k-nearest-neighbor lookup in the UMAP embedding space.","section":"Suzipu paper, Section III-B"},{"comment":"The reference to the tree edit distance algorithm should be 'Zhang and Shasha', not 'Zhang and Sasha'.","section":"MTN paper, Section III-C"}],"recommendation":"major_revision","confidential_remarks":"The submission is a workshop proceedings rather than a single research article, and its papers are of uneven depth. If the editorial model is to accept the collection as workshop documentation, I would not object; however, at the standard of a serious journal, the SMIReT paper's central multitask claim needs a single-task baseline and an ablation of the synthetic pretraining stage, and the MLLM paper needs to be repositioned as a pilot study. The YOLO paper is already at archival quality and could stand alone as a journal contribution. The GNN and MTN papers are useful and transparent about their limitations, but the GNN paper should explicitly report the recall ceiling imposed by its candidate graphs."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is the WoRMS 2024 proceedings—a preface plus 15 short papers—so there is no single central claim to referee, and the reader's UNVERDICTED call is the right one. The two pieces worth your time are the YOLO staff-layout paper and the SMIReT multitask paper.\n\nThe YOLO paper (Dvořák, Hajič, Mayer) is the soundest thing here. They harmonize three existing datasets into a 7,013-image corpus, add a new negative dataset (MZKBlank), evaluate in-domain and on three out-of-domain splits with standard COCO metrics, and ship models and scripts. The headline result—YOLOv8m matches or beats Pacha's Faster R-CNN MeasureDetector while running roughly 26x faster and using about 4x less memory—is real and directly useful for library-scale digitization. They also document YOLO's remaining weakness on out-of-domain handwritten MUSCIMA++ and disclose that parts of OSLiC/AudioLabs annotations were machine-generated or filtered, which is honest reporting.\n\nSMIReT is a genuine if incremental extension of the authors' own Sheet Music Transformer: task prompting plus a unified vocabulary (music tokens, text, Pix2Seq-style box coordinates) so one autoregressive decoder covers parsing, layout, and query tasks. Framing SMIR as one sequence-generation problem is worth airing. But the stress-test concern holds: Table I reports seven numbers with no single-task baseline and no ablation of the synthetic PRIMENS curriculum, so you cannot attribute the 5.92 SER to multitasking rather than to the underlying SMT. The authors themselves qualify \"successful\"—region IoU sits below dedicated layout-analysis systems and selective OMR is 41.55 SER. For a four-page workshop paper that is forgivable; the abstract just overstates it.\n\nAlso worth noting: the MTN evaluation-format paper ships a converter, a 435k-image dataset, and a tiered metric suite—reproducible infrastructure. The GNN paper is candid about training instability and reports that its k=20 candidate graphs contain only 80–91% of ground-truth edges, which caps achievable recall; that kind of honest reporting is useful. The weakest entry is the MLLM evaluation: three images, five proprietary models, subjective checkmarks, and prompts co-written by the models themselves. Fine as an informal probe, thin as evidence.\n\nWho this is for: OMR researchers and digital-library practitioners, not a general ML audience. It deserves a serious referee—not the volume as a whole, but the strongest constituent papers, which should go through full review with baselines and ablations requested.","headline":"A workshop proceedings volume, not a single paper: the YOLO layout-analysis study is solid and reusable, the SMIReT multitask work is promising but under-evidenced, and the rest is a useful state-of-the-field snapshot.","tokens_in":51831,"tokens_out":3445,"would_cite":true,"duration_ms":31765,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single transformer with task prompts, SMIReT, can handle every sheet-music retrieval task end-to-end on 17th-century prints, with music error near 6 percent and region-classification F1 at 97.","keywords":["optical music recognition","sheet music information retrieval","multitask transformers","task prompting","staff layout analysis","semantic reconstruction","mensural notation","music score datasets"],"falsifier":"Train SMIReT on the same MOTTECTA fine-tuning schedule from random initialization, omitting the synthetic-pretraining curriculum, and compare music symbol error rate, region intersection-over-union, and pattern-match accuracy on the held-out test pages: if those numbers stay near 6, 70, and 74, the transfer premise is not load-bearing and the multitask claim must be re-attributed; if they collapse, the premise is confirmed. A second check sits ready inside the graph-neural-network paper: its candidate graphs contain only 80-91 percent of ground-truth edges for two of its datasets, so no recall figure from that pipeline can exceed those fractions no matter how the edge classifier is tuned.","tokens_in":50635,"feed_emoji":"🎼","tokens_out":18998,"duration_ms":157729,"temperature":0.7,"pith_summary":"These proceedings of the 6th International Workshop on Reading Music Systems present a field shifting from single-purpose recognition pipelines toward unified end-to-end models, cheaper layout analysis, and human-in-the-loop tools for historical collections. The strongest claim is that one autoregressive transformer, SMIReT, can perform every task in a newly proposed Sheet Music Information Retrieval challenge on the MOTTECTA corpus of 297 printed 17th-century mensural pages: full-page parsing, music-only and text-only recognition, region detection and classification, partial transcription of user-specified regions, and pattern-search queries. The reported numbers are music symbol error rates of 6.05 for full parsing and 5.92 for music-only recognition, text error of 10.08, region-classification F1 of 97.00, and pattern-match accuracy of 73.80, with the weak spots being spatial localization (intersection-over-union 70.23) and selective transcription (error rate 41.55). A second load-bearing result is that a YOLOv8m detector matches or slightly beats the previous Faster R-CNN measure detector on staff layout analysis while running more than 20 times faster and using more than 4 times less memory, which would make layout indexing cheap enough for library-scale digitization. A sympathetic reader would care because these results, if they hold, mean one model could replace several task-specific services and that efficient multitask reading of historical music is no longer out of reach.","feed_headline":"One transformer parses, transcribes and searches sheet music","feed_subtitle":"One model performs parsing, layout and musical queries on 17th-century prints; staff detection is 20x faster.","key_machinery":"The carrying mechanism is task prompting over a unified multimodal vocabulary inside an autoregressive transformer: a short sequence of prompt tokens is prepended to the decoder input, redirecting the same encoder features and the same vocabulary toward parsing, layout analysis, OCR, selective transcription, or pattern queries. The vocabulary is deliberately multimodal, binding music symbols in an agnostic shape-plus-position encoding, text characters, absolute-position bounding-box tokens in the style of the Pix2Seq detection procedure cited in the paper, and special tokens for region categories. Curriculum learning carries the model from the synthetic PRIMENS incipit collection to the real MOTTECTA corpus, feeding pages with progressively more staves and text before interleaving synthetic and target samples; this transfer step is the load-bearing premise of the multitask claim. For the volume's layout-analysis result, the equivalent mechanism is a single-stage YOLOv8m detector that replaces the two-stage region-proposal pipeline and avoids the confidence collapse the older model shows when a grand staff coincides with a full system.","core_discovery":"In the authors' own framing, the central discovery is that task unification for music documents is feasible: SMIReT adapts the Sheet Music Transformer, an autoregressive encoder-decoder that generates one output token at a time conditioned on a convolutional feature map of the page and on all tokens generated so far, by prepending a prompt-token sequence to the decoder input. With task prompts, a single unified vocabulary covering music symbols (encoded agnostically as shape plus staff position), text characters, bounding boxes written as absolute position tokens, and region-category tokens, plus a curriculum that moves from synthetic images of mensural incipits to real 17th-century pages, the same weights produce full parsing, OMR, OCR, layout recognition, selective OMR, and pattern-matching queries; the authors state that the model learns all of them successfully with acceptable performance. They read the uneven error profile as a feasibility result with clear next targets: text recognition degrades by 51.78 percent when mixed with music, localization lags classification, and the hardest task is the one where the user points at pixels and asks for a partial transcription. Around this result the volume documents adjacent findings: a graph-neural-network pipeline reconstructs music-notation graphs but is seed-sensitive and its candidate graphs already cap recall, a tree-based evaluation format with tiered metrics is proposed so that systems can be compared fairly, general multimodal language models can identify tonality and texture from score images but cannot yet transcribe them, and the YOLO layout result makes cheap page analysis practical.","pith_inferences":["The paper describes the curriculum verbally and reports no ablation that isolates the synthetic pretraining; a natural experiment is to train SMIReT on the target corpus alone. If the synthetic pretraining is truly load-bearing, that version should be clearly worse; if it is not, the multitask credit belongs to the unified vocabulary and fine-tuning schedule rather than the curriculum.","The contrast between pattern-match queries (73.80 accuracy) and selective OMR (41.55 error rate) is revealing because both consume bounding-box information; the difference suggests the bottleneck is instruction-following under dense, per-symbol guidance rather than localization itself, which could be tested by varying how many regions a query specifies.","The out-of-domain asymmetry in the layout paper, where YOLO handles grand staffs well but staffs and measures poorly on handwritten pages, points to a cheap extension: adding synthetic handwritten staff lines to the training mix would show whether the gap is data or architecture.","The GNN paper already computes the ceiling for its own approach, with candidate graphs containing only 80-91 percent of ground-truth edges for two datasets; folding music-grammar rules into candidate-graph construction, as the authors suggest, is the directly testable route past that cap."],"forward_implications":["A single SMIReT-style model could replace separate OMR, OCR, and layout services in a digitization workflow, removing the compute and maintenance overhead the paper identifies as a practical obstacle.","Adding a new sheet-music reading task would need no new architecture: a new prompt token and matching training data extend the same model, exactly the mechanism used for the six task families evaluated here.","With YOLO-class layout detection running at 0.83 seconds per page on a CPU, staff and system indexing of million-page collections becomes feasible as a cheap gate before expensive full-page recognition is invoked.","Comparisons between OMR systems become meaningful if the field adopts a shared representation like the proposed Music Tree Notation with its tree-edit-distance metric, replacing per-methodology evaluation.","The failure modes the papers identify, especially selective transcription at 41.55 error rate, put a concrete target on attention mechanisms that can ground instructions in specific regions of the score image."],"supporting_citations":[{"why":"The Sheet Music Transformer that SMIReT extends; supplies the autoregressive encoder-decoder and full-page transcription machinery on which the multitask model is built.","marker":"[22]"},{"why":"Supplies both training worlds: the PRIMENS synthetic mensural incipits for curriculum pretraining and the fully labelled MOTTECTA corpus of 297 printed 17th-century pages used for fine-tuning and evaluation.","marker":"[27]"},{"why":"Provides the procedure for writing object detection as absolute-position tokens inside a language model; the bounding-box tokens of SMIReT's unified vocabulary follow it.","marker":"[26]"},{"why":"Foundational attention mechanism for the transformer decoder that generates the output symbol sequence.","marker":"[21]"},{"why":"Precedent for task prompting in document understanding, the mechanism SMIReT adopts to turn one decoder into six task specialists.","marker":"[19]"},{"why":"The state-of-the-art region-based layout analysis system whose performance is the explicit comparison point for SMIReT's region IoU of 70.23.","marker":"[15]"},{"why":"The Faster R-CNN measure detector (cited in the staff-layout paper) that serves as the baseline for the YOLOv8m accuracy, speed, and memory comparisons.","marker":"[16]"}],"fun_headline_variants":["One transformer unifies parsing, transcription, and music queries","Single model reads 17th-century sheet music across tasks","SMIReT: one model for OMR, OCR, and pattern queries","Unified transformer tackles all music reading tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire multitask result rests on one unisolated premise: that curriculum pretraining on synthetic images of early mensural notation transfers enough to real 17th-century pages for a single vocabulary and a few prompt tokens to cover all six task families; the paper asserts this curriculum in words only, reports no ablation isolating it, and its most visible weak point is the 41.55 error rate of the selective-transcription task, where the premise must do the most work.","fun_headline_variants_meta":{"raw":{"variants":["One transformer unifies parsing, transcription, and music queries","Single model reads 17th-century sheet music across tasks","SMIReT: one model for OMR, OCR, and pattern queries","Unified transformer tackles all music reading tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00056,"raw_usage":{"total_tokens":2686,"prompt_tokens":996,"completion_tokens":1690,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":1633}},"tokens_in":612,"tokens_out":1690,"duration_ms":12250,"temperature":1.0,"reasoning_tokens":1633,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:57:30.504078+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train SMIReT on the same MOTTECTA fine-tuning schedule from random initialization, omitting the synthetic-pretraining curriculum, and compare music symbol error rate, region intersection-over-union, and pattern-match accuracy on the held-out test pages: if those numbers stay near 6, 70, and 74, the transfer premise is not load-bearing and the multitask claim must be re-attributed; if they collapse, the premise is confirmed. A second check sits ready inside the graph-neural-network paper: its candidate graphs contain only 80-91 percent of ground-truth edges for two of its datasets, so no recall figure from that pipeline can exceed those fractions no matter how the edge classifier is tuned.","supporting_citations":[{"cited_title":"Le-Tien, T","cited_arxiv_id":null,"evidence_quote":"The Sheet Music Transformer that SMIReT extends; supplies the autoregressive encoder-decoder and full-page transcription machinery on which the multitask model is built."},{"cited_title":"Calvo-Zaragoza, A","cited_arxiv_id":null,"evidence_quote":"Provides the procedure for writing object detection as absolute-position tokens inside a language model; the bounding-box tokens of SMIReT's unified vocabulary follow it."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Precedent for task prompting in document understanding, the mechanism SMIReT adopts to turn one decoder into six task specialists."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The state-of-the-art region-based layout analysis system whose performance is the explicit comparison point for SMIReT's region IoU of 70.23."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Faster R-CNN measure detector (cited in the staff-layout paper) that serves as the baseline for the YOLOv8m accuracy, speed, and memory comparisons."}],"review_version":1}