{"id":"0ebcaa8c-f0f9-47b4-aadc-7b7a86c56917","arxiv_id":"2607.06015","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"A meta-benchmark that auto-generates multimodal music-perception multiple-choice tests from user symbolic music, demonstrated on ChoraleBricks with text-only and white-noise controls.","lead":"MusICA-MetaBench turns user-supplied music into automatic multimodal quizzes that test whether AI models can perceive music in audio, sheet images, and symbolic files. It aims to let people evaluate music AI on the repertoire they care about, with controls meant to check real perception rather than language shortcuts.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Text-only and white-noise baselines do not establish content-specific music perception for notation or symbolic modalities, so the claim that generated questions “do measure music perception” is under-supported.","rationale":"The Reader correctly identified the sufficiency of templates plus text-only/white-noise controls as the weakest assumption behind the claim that questions measure music perception, and correctly left the paper UNVERDICTED given incomplete methods access. My concern is the same load-bearing point, made more concrete: white-noise is modality-asymmetric, and neither baseline tests content-specificity (correct piece vs wrong piece) or symbolic parsing vs perception. The meta-benchmark idea (on-demand, pedagogy-aligned, cross-modal, size-calibrated) is coherent and addresses real evaluation gaps; the soft spot is specifically the perception-validity argument, not the overall framing. Without full template inventory, control design, statistics, and artifacts, the claim cannot be confirmed or refuted, so the verdict should stay UNVERDICTED. A full-text review that includes mismatched-input or equivalent ablations, clear distractor construction, and released code/data could move the verdict to CONDITIONAL or ACCEPT; absence of content-specific controls would justify rejecting the strong “do measure music perception” wording while still allowing the framework contribution under a weaker claim. No formal verification or parameter-free derivation is at issue. Agreement with the Reader is full on both the weak assumption and the withheld verdict.","tokens_in":2136,"tokens_out":688,"duration_ms":44728,"concrete_test":"On ChoraleBricks, run a mismatched-content control: keep each item’s question text and options, but replace the multimodal input with that of a different chorale (same modality: audio / notation image / symbolic). Compare accuracy to the matched condition and to the paper’s text-only baseline. If mismatched accuracy stays significantly above text-only (or near matched), the reported baselines do not establish perception of the intended musical content and the central perception claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim’s load-bearing step is that fixed pedagogy-aligned templates applied to symbolic encodings, plus text-only and white-noise baselines, suffice to show answers require genuine perception of the intended multimodal input. White-noise is audio-specific and does not control notation images or symbolic files. For MusicXML/symbolic input, template questions about pitch, key, harmony, or rhythm can be answered by structured parsing of tags without any auditory or visual music perception—yet that path still beats text-only and is untouched by white-noise. For notation images, OCR-plus-language reasoning can similarly succeed without musical perception. Even for audio, white-noise only shows that non-noise acoustic input helps; it does not show the model uses the musically relevant content of the correct piece rather than generic audio statistics or option–template correlations that appear only when real audio is present. Without content-mismatch, scrambled-score, or modality-appropriate ablations, residual shortcuts (metadata, template artifacts, OCR/spectrogram cues, language priors) can survive the two reported controls. That is the least secure condition for the abstract’s assertion that the questions measure music perception.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript introduces MusICA-MetaBench, a meta-benchmark framework that automatically derives on-demand multimodal multiple-choice benchmarks from user-provided symbolic music (e.g., MusicXML) via fixed pedagogy-aligned question templates. Generated items are intended to probe music perception competencies across three modalities—audio, notation images, and symbolic files—and to support systematic cross-modal comparison. The authors demonstrate the pipeline on the ChoraleBricks dataset, experimentally choose benchmark sizes argued to yield statistically reliable model comparisons, and report text-only and white-noise baseline controls as evidence that the questions measure music perception rather than non-perceptual shortcuts. The central claim is that this dataset-specific paradigm overcomes three limitations of static music benchmarks: evaluation cost, unclear transfer beyond the fixed corpus, and lack of cross-modal assessment.","tokens_in":2366,"tokens_out":1174,"duration_ms":32618,"significance":"If the framework and controls hold, the contribution is practically significant for multimodal music evaluation: on-demand, user-data-driven benchmarking reduces the need for large static suites and enables modality-matched comparisons on music the community actually cares about. Pedagogy-aligned templates and explicit attention to sample-size reliability are strengths relative to ad-hoc music LLM leaderboards. The work is constructive and falsifiable in principle (baseline gaps, size experiments). Significance of the headline claim that the questions “do measure music perception,” however, rests on whether the reported controls close residual shortcut paths in every modality; that is the load-bearing empirical hinge for the paper’s positioning against prior “music understanding” benchmarks.","major_comments":[{"comment":"The abstract’s assertion that text-only and white-noise baselines show the questions “do measure music perception” is load-bearing for the paper’s second claimed advance (benchmarks that require perception). White-noise is an audio-domain control and does not address notation-image or symbolic-file pathways. For MusicXML/symbolic input, template questions on pitch, key, harmony, or rhythm can be solved by structured parsing of tags without auditory or visual music perception, yet still beat text-only and remain untouched by white-noise. For notation images, OCR-plus-language reasoning can succeed without musical perception. Even for audio, white-noise only shows that non-noise acoustic input helps; it does not show use of piece-specific musical content versus generic audio statistics or option–template correlations that appear only when real audio is present. Content-mismatch, scrambled-","section":null},{"comment":"The claim that experimentally determined benchmark sizes “ensure statistically reliable model comparisons” on ChoraleBricks is central to the efficiency argument. The manuscript must report the design of that size experiment (power analysis or resampling procedure, effect-size assumptions, multiple-comparison handling across models/modalities/templates, and the decision rule for N). If reliability is only shown for a subset of templates or one modality, the cross-modal comparison claim should be narrowed accordingly. Tables or figures of pairwise separation vs. N, with uncertainty, are needed so readers can judge whether the chosen sizes actually support the stated comparisons.","section":null},{"comment":"“Pedagogy-aligned” competencies are part of the framing that distinguishes MusICA from generic music QA. The paper should specify which pedagogical sources or skill taxonomies the templates implement, how coverage was validated (expert review, curriculum mapping), and whether template difficulty or construct validity was checked beyond face validity. If alignment is only author-asserted, the claim should be softened to “pedagogy-inspired templates” unless supporting evidence is added.","section":null}],"minor_comments":[{"comment":"Clarify early (abstract/intro) the exact symbolic formats supported at generation time versus at evaluation time, and whether MIDI is treated as first-class input or only as a render path from MusicXML.","section":null},{"comment":"Define “music perception” operationally versus “music understanding” when contrasting prior benchmarks, so the baseline design can be judged against a stated construct.","section":null},{"comment":"When reporting baseline gaps, give absolute accuracies and chance levels per modality and template family, not only relative improvements, so shortcut residual performance is visible.","section":null},{"comment":"State licensing and redistribution constraints for user-provided music and for ChoraleBricks-derived items, since on-demand generation from third-party scores raises reuse questions for shared leaderboards.","section":null},{"comment":"If code and template catalogs are released, document the exact generation seed, option-shuffling policy, and any exclusion rules for degenerate items (e.g., trivial keys, empty voices) so results are reproducible.","section":null}],"recommendation":"major_revision","confidential_remarks":"I only had the abstract plus the provided reader/skeptic notes in this packet; section/equation/table numbers for the size experiment and baseline tables were not available to cite. My major comments target the load-bearing perception-control and reliability claims as stated in the abstract. If the full manuscript already contains content-mismatch/scrambled-score ablations and a transparent power/size analysis, those two majors can be downgraded after verification. Fit for cs.SD / multimodal evaluation venues looks appropriate if the control gap is closed. No integrity red flags from the materials seen."},"author_rebuttal":{"model":"grok-4.5","summary":"We thank the referee for a careful and constructive report. The three major comments correctly identify load-bearing points in our positioning: (i) what the text-only and white-noise controls actually establish about “music perception,” (ii) the transparency of the sample-size experiment that underwrites reliable comparisons on ChoraleBricks, and (iii) the strength of the “pedagogy-aligned” framing. We agree that the abstract and related claims over-reach relative to the controls and evidence currently reported, and we will revise the manuscript accordingly: soften and modality-qualify the perception claim, fully document the size experiment (design, decision rule, and separation-vs-N evidence), and either ground or soften the pedagogy language. Below we respond point by point. We believe these revisions address the recommendation for major revision without changing the core technical contribution of the meta-benchmark pipeline.","responses":[{"response":"We agree that the abstract’s unqualified claim is stronger than the controls support, and that white-noise is an audio-domain control only. Text-only is a modality-agnostic “no music input” baseline; white-noise only probes whether non-noise audio helps. Neither rules out (a) MusicXML tag parsing for symbolic input, (b) OCR-plus-linguistic reasoning for notation images, nor (c) residual audio shortcuts (generic acoustics, option–template correlations that appear only with real audio). We will revise the abstract and main text to state only what the controls establish: that performance drops without music-bearing input (text-only) and, for audio, without structured acoustic content (white-noise), and that residual non-perceptual pathways remain possible in each modality. We will add an explicit limitations subsection on shortcut risks (symbolic parsing, OCR, audio statistics) and, where feasible within revision, report content-mismatch / scrambled-content or modality-scrambled controls on a subset of templates to tighten the claim. We will not claim that the present baselines “show the questions do measure music perception” across all modalities.","revision_made":"yes","referee_comment":"The abstract’s assertion that text-only and white-noise baselines show the questions “do measure music perception” is load-bearing. White-noise is audio-only and does not address notation-image or symbolic pathways. Symbolic MusicXML items can be solved by structured tag parsing without perception; notation images by OCR-plus-language reasoning. Even for audio, white-noise only shows non-noise input helps, not piece-specific musical content vs. generic audio statistics or option–template correlations. Content-mismatch / scrambled controls are needed; residual shortcut paths remain."},{"response":"We agree that the size claim is central to the efficiency argument and that the current manuscript does not report the experiment design at the level required for readers to judge reliability. In revision we will fully document: (1) the procedure used (resampling / bootstrap of item subsets and pairwise model separation as a function of N), (2) the effect-size and separation criteria that drove the decision rule for chosen N, (3) how we handled multiplicity across models, modalities, and templates (or explicitly note if we did not correct and why), and (4) whether the chosen N was validated for all templates and modalities or only a subset. We will add tables and/or figures of pairwise separation versus N with uncertainty bands so that the chosen sizes can be audited. If reliability is only demonstrated for a subset of templates or modalities, we will narrow the cross-modal comparison claim accordingly rather than assert global statistical reliability.","revision_made":"yes","referee_comment":"The claim that experimentally determined benchmark sizes “ensure statistically reliable model comparisons” on ChoraleBricks is central. The manuscript must report the design of that size experiment (power analysis or resampling procedure, effect-size assumptions, multiple-comparison handling across models/modalities/templates, and the decision rule for N). If reliability is only shown for a subset of templates or one modality, the cross-modal comparison claim should be narrowed. Tables or figures of pairwise separation vs. N, with uncertainty, are needed."},{"response":"The referee is correct that “pedagogy-aligned” currently rests on author design of templates around standard ear-training / theory competencies (pitch, key, harmony, rhythm, etc.) rather than a documented mapping to a named curriculum or taxonomy with external validation. We did not conduct formal expert review, curriculum mapping, or psychometric checks of difficulty/construct validity beyond face validity of the templates. We will therefore either (a) add a short subsection naming the pedagogical sources that informed the template set and describing any informal expert feedback we can document, or—if that evidence remains thin—(b) soften the language throughout (abstract, introduction, and claims) to “pedagogy-inspired” templates that target common music-perception competencies, and state clearly that construct validity and curriculum coverage are not formally validated. We will not retain the stronger “aligned with music pedagogy” phrasing without supporting evidence.","revision_made":"yes","referee_comment":"“Pedagogy-aligned” competencies distinguish MusICA from generic music QA. The paper should specify which pedagogical sources or skill taxonomies the templates implement, how coverage was validated (expert review, curriculum mapping), and whether template difficulty or construct validity was checked beyond face validity. If alignment is only author-asserted, soften to “pedagogy-inspired templates” unless supporting evidence is added."}],"tokens_in":1991,"tokens_out":1273,"duration_ms":23145,"standing_objections":["We cannot fully close residual non-perceptual shortcut paths in every modality (MusicXML structured parsing, notation OCR-plus-language reasoning, and audio generic-statistics / option–template correlations) with the present control suite alone; even with added content-mismatch/scrambled controls on a subset, a complete multi-modality shortcut audit is beyond what a single revision can guarantee, so the perception claim will remain carefully qualified."]},"desk_editor":{"model":"grok-4.5","letter":"The one thing to know: this is a practical meta-benchmark that turns user symbolic music (MusicXML and kin) into pedagogy-templated multimodal MCQs for MLLMs, with a ChoraleBricks demo and sizing work for reliable comparisons. The second: the abstract’s assertion that text-only and white-noise baselines show the questions “do measure music perception” is thinner than it sounds once you look past audio.\n\nWhat is actually new is the packaging. Static music benchmarks are costly and transfer poorly; MusICA-MetaBench generates on-demand tests from repertoire the user cares about, across audio, notation images, and symbolic files, with fixed templates aimed at perception competencies. That combination is a real methods contribution for music-AI eval. Naming the three problems—cost/transfer, weak perception requirements, missing cross-modal comparison—and building something aimed at all three is honest engineering. The size experiments for ChoraleBricks are the kind of detail people skip and shouldn’t. Credit for that.\n\nThe soft spot is load-bearing and matches the stress-test note. White-noise only controls audio. For symbolic input, pitch/key/harmony/rhythm questions can be answered by structured parsing of tags with no musical perception; that path still beats text-only and is untouched by white-noise. For notation images, OCR plus language priors can do similar work. Even for audio, white-noise only shows non-noise acoustics help, not that the model uses the musically relevant content of the correct piece rather than generic stats or option–template correlations. Without content-mismatch, scrambled-score, or modality-matched ablations, residual shortcuts can survive. The central constructive idea still holds; the perception claim needs stronger controls than the two reported baselines.\n\nWho it’s for: music-AI evaluation people and product teams scoring multimodal music models on their own repertoire. Not a theory paper. It deserves a serious referee—methods contribution is real enough that peer review should pressure the control design rather than desk-reject. I’d bring it to a reading group if we have multimodal-eval people in the room; otherwise maybe. I would not cite it in my own work in the next year unless I were building a music benchmark.\n\nSend to peer review. Ask reviewers to push hard on modality-appropriate controls before accepting the perception claim at face value.","headline":"Useful on-demand multimodal music-perception MCQ framework from user scores; the claim that text-only + white-noise prove genuine perception is under-supported for notation and symbolic paths.","tokens_in":3005,"tokens_out":580,"would_cite":false,"duration_ms":19455,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"On-demand benchmarks from your own scores test whether MLLMs actually hear and read music","keywords":["music perception","multimodal large language models","benchmarking","symbolic music","MusicXML","multiple-choice questions","cross-modal evaluation","on-demand benchmarks"],"falsifier":"On a held-out set of user scores, a strong multimodal model that scores near chance on both the text-only and white-noise controls still scores far above chance on the rendered audio/notation questions without using any musical content—e.g., by exploiting OCR artifacts, spectrogram shortcuts, or template regularities that survive the baselines.","tokens_in":3030,"feed_emoji":"🎵","tokens_out":860,"duration_ms":63740,"temperature":0.7,"pith_summary":"Large static music benchmarks for multimodal language models are expensive to run, hard to trust on music outside their fixed set, and often can be solved without ever perceiving the music itself. This paper introduces MusICA-MetaBench, a framework that turns user-provided symbolic scores (such as MusicXML) into multiple-choice questions automatically, using fixed templates drawn from music pedagogy. The same questions can be asked in audio, notation-image, or symbolic-file form, so models can be compared across modalities on the same musical material. Text-only and white-noise controls are used to check that correct answers require genuine perception of the musical input rather than language priors or residual metadata. On a demonstration set of Bach-style chorales (ChoraleBricks), the authors also find how many questions are needed for statistically reliable model rankings. If the method holds, anyone with a score library can generate a focused, perception-centered benchmark for the music they actually care about.","feed_headline":"Your own scores become MLLM music-perception tests on demand","feed_subtitle":"Templates plus text-only and white-noise controls check that models actually hear or read the music","key_machinery":"Automatic template-based generation of multiple-choice perception questions from symbolic encodings (e.g., MusicXML), rendered into audio, notation images, and symbolic files, with text-only and white-noise baselines that isolate genuine cross-modal music perception from language priors and residual cues.","core_discovery":"MusICA-MetaBench automatically builds on-demand multimodal multiple-choice benchmarks from user-supplied symbolic music. Pedagogy-aligned question templates applied to structured encodings produce questions that, when paired with text-only and white-noise baselines, measure music perception rather than non-perceptual shortcuts, and experimentally chosen sizes support statistically reliable model comparisons (shown on ChoraleBricks).","pith_inferences":["If symbolic-to-multimodal rendering is faithful, the same pipeline could stress-test whether models that claim score following actually track voice leading, cadence, or key rather than surface pattern matching.","The method may transfer to other structured cultural media (e.g., dance notation or chess scores) where templates plus modality controls can separate perception from language priors.","Failure modes that survive white-noise and text-only baselines—such as OCR of printed accidentals or spectrogram texture cues—would become the next natural control targets.","Community score libraries could become shared “what we care about” testbeds, shifting evaluation from fixed corpora toward domain-specific perception stress tests."],"forward_implications":["Anyone with a MusicXML (or similar) library can generate a custom perception benchmark without building a new static dataset.","Model rankings can be compared across audio, notation images, and symbolic files on identical musical material.","Benchmark size can be chosen so that pairwise model differences are statistically reliable for a given score collection.","Claims of “music understanding” can be stress-tested against text-only and white-noise controls before they are accepted as perception.","Evaluation cost scales with the user’s own data rather than with ever-larger fixed public suites."],"fun_headline_variants":["On-demand MLLM music tests auto-built from your symbolic scores","Your music becomes multimodal perception quizzes for MLLMs","Templates turn scores into MLLM tests that measure real perception","Baselines prove the questions test music perception not shortcuts","Statistically reliable MLLM music benchmarks from any user scores"],"cache_read_input_tokens":2176,"weakest_assumption_plain":"That fixed pedagogy-style question templates on structured symbolic scores, plus text-only and white-noise checks, are enough to guarantee that correct answers need real perception of the music rather than leftover metadata, template artifacts, or language shortcuts that those two controls miss.","fun_headline_variants_meta":{"raw":{"variants":["On-demand MLLM music tests auto-built from your symbolic scores","Your music becomes multimodal perception quizzes for MLLMs","Templates turn scores into MLLM tests that measure real perception","Baselines prove the questions test music perception not shortcuts","Statistically reliable MLLM music benchmarks from any user scores"]},"model":"grok-4.5","cost_usd":0.021364,"raw_usage":{"total_tokens":4180,"prompt_tokens":833,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":213640000,"prompt_tokens_details":{"text_tokens":833,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3263,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":833,"tokens_out":84,"duration_ms":45911,"temperature":1.0,"reasoning_tokens":3263,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T19:24:49.487942+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a held-out set of user scores, a strong multimodal model that scores near chance on both the text-only and white-noise controls still scores far above chance on the rendered audio/notation questions without using any musical content—e.g., by exploiting OCR artifacts, spectrogram shortcuts, or template regularities that survive the baselines.","supporting_citations":[],"review_version":1}