{"id":"9f7517db-f123-4724-ab6a-8806bd2476e3","arxiv_id":"2507.15590","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 51-respondent survey of computational musicology practitioners identifies limited functionality, integration difficulties, and poor documentation as the main gaps in current tools.","lead":"This paper reports a survey of 51 computational musicology practitioners, mapping which tools they use, what tasks they perform, and where tools fall short. It finds limited functionality, integration difficulties, and poor documentation are the top complaints across symbolic music, audio, images, and text domains.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'limited functionality' gap is inferred solely from a closed-ended checkbox; the survey's own open-ended gap question is never reported, and Figure 6 lacks per-modality denominators.","rationale":"The reader identifies sampling representativeness as the weakest assumption. I disagree in emphasis. Representativeness is a limitation the authors themselves acknowledge (Sec. 5) and, for a descriptive community snapshot, 51 responses from a niche field can still be informative if framed as a pilot. The more damaging issue is internal validity: the survey collected the exact evidence that would substantiate 'significant gaps'—participants' free-text descriptions of missing tasks and features—but the results section never reports it. The only quantitative support for the central claim is a checkbox labeled 'Limited functionality' in a 'what do you not like' question. That item conflates many possible meanings: missing features, insufficient depth, poor integration with the user's workflow, or simply a generic dissatisfaction. Without the open-ended data, we cannot know what 'limited functionality' meant, so the paper's primary finding is uninterpretable. The missing per-modality denominators further weaken the cross-modality consistency claim. None of this makes the paper wrong, but it makes the central claim conditional on a re-analysis that the authors can easily perform since they already collected the data. Hence the verdict remains CONDITIONAL, with the condition being the reporting of the open-ended analysis and per-modality N, rather than (only) softening the language. This is an honest non-fatal concern: the study could be sound, but the evidence as presented doesn't yet support the headline.","tokens_in":9033,"tokens_out":4622,"duration_ms":47070,"concrete_test":"Obtain the anonymized response data (which the paper should release). Then: (a) for each modality, report N and compute 95% Wilson intervals for the proportion reporting 'Limited functionality'; (b) code the free-text answers to 'List tasks for which you lack adequate tools' and check whether every respondent who selected 'Limited functionality' named at least one concrete missing feature; (c) recompute Figure 6 percentages under alternative denominators (participants vs. tool mentions). If image/text strata have N<10, or if fewer than half of 'Limited functionality' selections are corroborated by a specific free-text gap, the cross-modal 'limited functionality' claim loses support. This single re-analysis would settle whether the reported gaps are real or an artifact of the closed-ended format.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline finding—'limited functionality' as the most consistent gap across all four modalities (Fig. 6)—rests on a closed-ended checkbox in the 'What do you not like about them?' question (Sec. 3.2). Three problems make this load-bearing. (1) The survey also asked an open-ended question, 'List tasks for which you lack adequate tools. What specific features are missing...' (Sec. 3.2), which is the direct evidence for the claimed capability/need gap, yet Section 4 never reports or analyzes those responses. The central claim therefore relies on an indirect checkbox rather than participants' own descriptions of missing functionality. (2) The denominator for the percentages in Figure 6 is unspecified: it could be participants, tool mentions, or issue-tool pairs. If a participant who uses five tools can flag 'limited functionality' only once per participant but 'difficult installation' per tool, the relative rankings are an artifact of the response format. (3) Per-modality respondent counts are not given for Figs. 6/7; with only 51 total responses across four modalities, the image and text strata may have very small N, making the 'consistent across modalities' claim statistically fragile. The acknowledged sample bias (Sec. 5) is secondary to these internal-validity gaps because even a representative sample would not rescue the claim if the gap measure itself is unvalidated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a survey-based study of computational musicology (CM) practitioners, reporting on tool usage patterns, common analytical tasks, satisfaction levels, data characteristics, and prioritized features across four data modalities: symbolic music, audio, images, and text. Based on 51 valid responses, it identifies limited functionality, integration difficulties, and poor documentation as the most common complaints, and derives recommendations for tool development, interoperability, and GUI accessibility. The authors acknowledge the convenience sample and its potential biases, but frame the findings as revealing significant gaps between current tooling capabilities and user needs.","tokens_in":9409,"tokens_out":3654,"duration_ms":38362,"significance":"If the findings are robust, the paper offers a valuable, community-facing evidence base for tool developers in a field where user needs are often inferred anecdotally. Its strengths include a clear survey instrument covering multiple modalities, explicit attention to tasks, tools, and satisfaction, and a transparent discussion of sample limitations. The study is one of few attempts to systematically collect practitioner-level data on CM tooling, and the descriptive summaries could inform funding and development priorities. However, the central claim of 'significant gaps' currently rests on incompletely reported evidence, most notably the omission of the survey's open-ended gap question and ambiguous denominator definitions in the headline figures.","major_comments":[{"comment":"The open-ended question 'List tasks for which you lack adequate tools. What specific features are missing or need improvement?' is the most direct evidence for the paper's central claim of gaps between capabilities and user needs, yet Section 4 never reports or analyzes these responses. The headline finding that 'limited functionality' is the consistent primary concern across modalities (Fig. 6) rests instead on a closed-ended checkbox from 'What do you not like about them?' This is an indirect measure that does not capture the specifics of missing features or tasks. The authors should either report a thematic analysis of the open-ended answers or substantially qualify the gap claim as based on complaint categories rather than user-described unmet needs.","section":"Section 3.2 / Section 4 / Figure 6"},{"comment":"The denominators used to compute the percentages in Figures 6 and 7 are not defined in the text. For Figure 6, it is unclear whether percentages are over participants, tool mentions, or issue-tool pairs; since respondents could name up to five tools per modality and could select multiple issues, the relative rankings (e.g., limited functionality vs. difficult installation) may be an artifact of the response format. In addition, per-modality respondent counts are not reported, so the claim that the pattern is 'consistent across all modalities' is fragile given that only 51 total responses are spread over four modalities, with image and text likely having much smaller N. Please provide the denominator definitions and per-modality sample sizes.","section":"Figure 6 / Figure 7"},{"comment":"The statement 'over 85% of surveyed users engage with multiple modalities in their research workflows' appears without any supporting statistic in Section 4. No figure or table in the results reports this percentage, so the claim cannot be verified from the presented data. This number is used to motivate the 'Towards Integrated Music Analysis' discussion, so it should be either added to the results with a clear computation or removed.","section":"Section 5"}],"minor_comments":[{"comment":"The list of programming proficiency options has a formatting issue: 'None / Very Basic (e.g., only using GUIs; Basic: Using existing scripts/tools...' appears to be missing a closing parenthesis or semicolon after 'GUIs'. Please correct the typography.","section":"Section 3.1"},{"comment":"The caption 'Corpora size' does not explain what the bars represent; please add axis labels (e.g., percentage of respondents) and specify which modality each group corresponds to.","section":"Figure 4"},{"comment":"The paper describes itself as 'the first systematic investigation of the practical needs of CM practitioners,' but reference [12] (Inskip and Wiering, 2015) already surveyed musicologists' attitudes toward technology. Please soften the novelty claim to acknowledge prior related surveys and specify what is new (e.g., focus on tool functionality across modalities).","section":"Section 5"},{"comment":"The phrase 'joint to the advent of personal computing technology' should be 'coupled with the advent of personal computing technology'.","section":"Section 1"},{"comment":"The 'top 80%' and 'top 75%' filtering is mentioned only in the captions; please explain in the text how the filtering was performed and whether the excluded values affect any conclusions.","section":"Captions of Figures 2 and 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for DLfM and addresses a genuine gap in the literature. The core problem is that the main claim—'significant gaps'—is not yet supported by the reported data analysis: the open-ended gap question is unused, and the headline figure lacks defined denominators and per-modality sample sizes. These are fixable with additional analysis and re-presentation, so major revision rather than rejection seems appropriate. The authors' self-citations are relevant to the topic and do not appear to distort the conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a clearly written descriptive survey of 51 computational musicology practitioners, asking about tool usage, satisfaction, and unmet needs across four modalities. What's genuinely new is the original data and the modality-specific breakdown: e.g., the co-occurrence analysis showing GUI tools complement rather than replace libraries, and the observation that medium-sized corpora dominate. The methodology is transparent, the related work is fair, and the discussion offers pragmatic, sensible takeaways for tool developers. I buy the general picture: integration is hard, documentation matters, and many musicologists prefer GUIs.\n\nThe soft spots are real but not disqualifying. As the stress-test note says, the central claim that \"limited functionality\" is the consistent primary concern rests on a closed-ended checkbox in the \"What do you not like about them?\" question. The survey also asked an open-ended question about tasks lacking adequate tools and specific missing features, but Section 4 never reports those responses. That is a genuine missed opportunity, because those free-text answers are the direct evidence for capability gaps. I also agree that the denominators in Figures 6 and 7 are unclear: are these percentages of participants, tool mentions, or issue-tool pairs? Without per-modality Ns, the \"consistent across all modalities\" claim is fragile given only 51 total responses. And the abstract's \"significant gaps\" is a stretch for a descriptive convenience sample with no inferential tests.\n\nThat said, these are problems of degree, not kind. The authors explicitly acknowledge the sample limitations in Section 5, and the paper reads as an honest first pass at mapping the field rather than an overclaiming study. The open-ended data could be added in a revision, and softening the language would make the paper a solid community snapshot. I would not call it a breakthrough, and I would not generalize from it to the whole field, but it is a useful, citable source of practitioner-reported pain points for anyone building or evaluating computational musicology tools.\n\nFor peer review: yes, send it. A serious referee should ask for the open-ended analysis, the per-modality denominators, and a more cautious abstract, but the paper deserves the time. I'd recommend conditional accept with those revisions.","headline":"A useful but modest descriptive survey of computational musicology tool users; the headline 'significant gaps' overstates what a 51-person convenience sample can support, but the data and transparency make it worth engaging.","tokens_in":9774,"tokens_out":1759,"would_cite":true,"duration_ms":21522,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of 51 practitioners across four music data modalities finds that limited functionality, integration difficulties, and poor documentation are the most consistent unmet needs in computational musicology tools.","keywords":["Computational musicology","Musicological tools","Digital humanities","User survey","Software usability","Tool interoperability","Symbolic music","Audio analysis"],"falsifier":"A replication survey with a larger, more geographically diverse sample that finds different leading concerns (for example, cost or speed instead of limited functionality) or satisfaction averages above 4 on the same 5-point scale would refute the paper's generalization about the field.","tokens_in":8832,"feed_emoji":"🎼","tokens_out":6311,"duration_ms":66967,"temperature":0.7,"pith_summary":"Computational musicology has grown from mainframe experiments in the 1960s into a heterogeneous ecosystem of libraries, command-line tools, and graphical applications, but the paper argues that no one has systematically measured whether these tools actually serve the researchers who use them. This paper reports on a survey of 51 computational musicology practitioners across four data modalities—symbolic music, audio, music-related images, and text—and claims that consistent gaps separate tool capabilities from user needs. Limited functionality is the most frequently reported problem in every modality, followed by difficulties integrating tools and poor documentation, and average satisfaction sits near neutral in all domains. If the finding holds, tool developers have a concrete, evidence-based set of priorities: close functionality gaps, improve interoperability, and write better documentation, with attention to the GUI preference of musicologists.","feed_headline":"Survey: musicology tools share the same gaps in all four data domains","feed_subtitle":"Limited functionality, integration trouble, and poor documentation top the list for symbolic, audio, image, and text tools.","key_machinery":"The carrying object is the survey instrument itself: a structured online questionnaire that combines demographic questions (role, field, experience, programming proficiency, usage frequency), modality-specific questions (common tasks, tools used, satisfaction, disliked aspects, corpus size, missing features), and cross-cutting questions (support channels, tool-selection criteria, integration difficulties). A modality here is a data representation of music: symbolic notation, audio signal, image, or text. The analysis turns these self-reports into per-modality frequency patterns, co-occurrence counts between tasks and tools, and concern rankings, which are the evidence for the claimed gaps.","core_discovery":"The paper's central claim is that the current landscape of computational musicology tools is misaligned with practitioners' actual workflows. Based on 51 valid responses, it reports that limited functionality, integration difficulties, and poor documentation are the dominant unmet needs across symbolic, audio, image, and text modalities; that satisfaction averages between 3.31 and 3.45 on a 5-point scale; that medium-sized collections (100–1,000 items) dominate research practice; and that nearly half of respondents report cross-modal integration difficulties. The paper presents this as the first systematic investigation of practical needs in computational musicology and as quantitative support for a long-observed communication gap between musicologists and software developers.","pith_inferences":["Beyond the paper: pairing the self-reported gaps with behavioral traces from issue trackers of widely used tools could test whether 'limited functionality' means genuinely missing features or mostly discoverability and training problems.","Beyond the paper: a standing, recurring version of this survey would turn the one-shot snapshot into a longitudinal signal that tool maintainers could use to track whether the reported gaps close over time.","Beyond the paper: qualitative interviews with the minority of respondents who report satisfaction at 4 or 5 could reveal which workflows are already well served and which tool-design patterns deserve replication."],"forward_implications":["If the reported gaps are real, the highest-leverage development targets are adding missing functionality, smoothing tool-to-tool integration, and improving documentation, in that order of reported importance.","Because over 85% of respondents work with more than one modality and nearly half report integration difficulties, tools that treat modalities in isolation will continue to fall short of researchers' needs.","The prevalence of small and medium collections (100–1,000 items) means deep-learning-heavy tools will not automatically transfer to musicological research without strategies for scarce data.","The clear GUI preference among musicologists, set against the library and command-line focus of developers, implies that visual, interactive tools are a necessary complement to programmatic ones rather than a nice-to-have."],"supporting_citations":[{"why":"Prior survey of musicologists' attitudes toward technology; supplies the framing of a communication gap that this survey quantifies.","marker":"[12]"},{"why":"Establishes music's multimodal nature, motivating the survey's four-modality structure.","marker":"[31]"},{"why":"Documents the interdisciplinary challenges and lack of comprehensive approaches in computational musicology, the problem this survey addresses.","marker":"[33]"}],"fun_headline_variants":["Survey: musicology tools share same gaps across all data types","Functionality, integration, docs: top unmet needs in music tools","51 experts rate musicology tools: just 3.4/5 satisfaction","Cross-modal integration fails for half of musicology tool users"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 51 people who answered the survey are treated as standing in for the whole field of computational musicology, even though they were recruited from the authors' own contacts and academic mailing lists and the paper admits this may skew toward established Western researchers.","fun_headline_variants_meta":{"raw":{"variants":["Survey: musicology tools share same gaps across all data types","Functionality, integration, docs: top unmet needs in music tools","51 experts rate musicology tools: just 3.4/5 satisfaction","Cross-modal integration fails for half of musicology tool users"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000489,"raw_usage":{"total_tokens":2338,"prompt_tokens":804,"completion_tokens":1534,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":420,"completion_tokens_details":{"reasoning_tokens":1460}},"tokens_in":420,"tokens_out":1534,"duration_ms":13261,"temperature":1.0,"reasoning_tokens":1460,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:26:48.875063+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication survey with a larger, more geographically diverse sample that finds different leading concerns (for example, cost or speed instead of limited functionality) or satisfaction averages above 4 on the same 5-point scale would refute the paper's generalization about the field.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior survey of musicologists' attitudes toward technology; supplies the framing of a communication gap that this survey quantifies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the interdisciplinary challenges and lack of comprehensive approaches in computational musicology, the problem this survey addresses."}],"review_version":1}