{"id":"8621f911-916e-4d54-b79a-a187abd6b0cc","arxiv_id":"2608.06903","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"AI music research concentrates technical and frontier-method investment in generation and content tasks, while education, health, and governance receive less support and adopt new methods years later.","lead":"This paper measures where AI music research attention goes, using 6,839 papers from 2015 to April 2026. It finds that content-generation tasks get more technical support and adopt new AI methods faster than education, health, and governance applications.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The quantitative lag/TIR headlines rest on first-adoption years and residuals in sparse single-label cells; with only aggregate validation and no sensitivity/uncertainty analysis, the reported 4.33/5.00-year gaps are not yet established.","rationale":"The paper's central contribution is a measurement claim, so the load-bearing condition is that the measured indicators reflect the field rather than noise or labeling artifacts. The reader's weakest assumption was corpus representativeness and LLM-label accuracy. My concern sharpens one specific mechanism by which that assumption can fail: FMAL is computed as a difference of first-adoption years, which are extreme order statistics in sparse task-method cells. For health, education, and governance the cells are small, so a single missing or mislabeled early paper can shift the lag by years. The validation in Table 1 is aggregate and does not report per-task error rates, and Section 7 explicitly notes first-adoption sensitivity without a sensitivity analysis. The same sparsity affects TIR and TMIR through small expected-count denominators, and no uncertainty quantification is provided. This is not an objection to the framework or to the qualitative conclusion, which is plausible and honestly qualified; it is an objection to the precision of the headline numbers. Minor internal inconsistencies in taxonomy counts (12 vs 13 application categories, 11 vs 12 technical families) should also be fixed but are not the primary concern. A CONDITIONAL verdict with a request for robustness checks remains appropriate, so I do not change the reader's verdict.","tokens_in":12048,"tokens_out":6970,"duration_ms":74677,"concrete_test":"Release the corpus and labels (or a blinded replica) and recompute FMAL and TIR under two perturbations. (1) Order-statistic robustness: for each frontier method m, remove the single earliest technical paper in the whole corpus and the single earliest paper in each task cell, then recompute lags; also require at least two papers in a task-method cell before counting adoption, otherwise treat the method as not adopted. (2) Label-noise simulation: bootstrap the validation sample to estimate per-task label error rates, corrupt the full corpus labels at those rates, and rerun the complete indicator pipeline. If generation remains within about one year of zero lag and education/health remain at least three years behind under both perturbations, the headline stands; if lags move by multiple years, the reported numbers are fragile artifacts of sparse cells.","verdict_should_be":"UNCHANGED","load_bearing_attack":"FMAL (Eqs. 7-8) is defined as t_{a,m} - t0_m, where t_{a,m} is the first publication year in the cell (task a, method m) and t0_m is the first year the method appears anywhere in the technical corpus. For sparse applications - health has 114 papers in the full corpus and a much smaller method-mapped subset, and several cells have very few papers - the earliest paper in a cell is an extreme order statistic. One missing early paper from an unindexed practice-oriented venue, one LLM mislabel, or one multi-task paper assigned to the wrong primary label changes t_{a,m} by years. Section 3.3 reports only aggregate Kappa (0.93-0.98), accuracy, and macro-F1 (93.5-94.6%) across a 403-500-paper validation sample; no per-task or per-cell error rates are given, so it cannot be ruled out that label noise is concentrated in exactly the small education/health/governance cells that drive the headline. The same sparsity affects TIR and TMIR: their denominators are sums of small expected counts, and no confidence intervals or bootstrap errors are reported. Section 7 acknowledges first-adoption sensitivity but offers no sensitivity analysis. Thus the qualitative direction (content tasks get more technical attention) is plausible, but the quantitative strength of the central claim - '0.33 years vs 4.33/5.00 years' - is not yet supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 6,839 AI music publications from 2015 to April 2026 and introduces a Research Attention Profile with four indicators: Technical Investment Residual (TIR), Task-Method Investment Residual (TMIR), Normalized Methodological Diversity (NMD), and Frontier Method Adoption Lag (FMAL). Using a joint taxonomy of application tasks and technical method families, with LLM-assisted labeling validated by human annotation, the authors report that technical and frontier-method support is concentrated in content-oriented tasks such as generation, MIR, and audio processing, while education, health, and governance remain under-supported. The headline quantitative finding is that generation adopts frontier methods with an average lag of 0.33 years, versus 4.33 years for education and 5.00 years for health. The paper concludes with a call for a more socially responsive AI music research agenda.","tokens_in":12298,"tokens_out":6269,"duration_ms":65794,"significance":"If the measurement framework is robust, the Research Attention Profile is a useful diagnostic tool for science-of-science studies of applied AI fields, and the paper makes a constructive contribution by tying bibliometric attention to methodological support, diversity, and adoption timing. The strengths are the explicit proportional-allocation baselines, the multi-source corpus, the validation protocol with high reported inter-annotator agreement (Cohen's kappa 0.93–0.98, accuracy above 93%), and the clear statement of the four indicators. However, the quantitative edge of the paper, especially the adoption-lag contrast, is not yet established because the headline statistics are computed from sparse cells using extreme-order statistics and because no uncertainty or sensitivity analysis is reported. The qualitative direction—content-heavy tasks receive more technical attention—is plausible, but the precise residual magnitudes and lag differences require additional support before the central claims can be fully accepted.","major_comments":[{"comment":"The headline quantitative result—adoption lags of 0.33 years for generation versus 4.33 and 5.00 years for education and health—rests on first-adoption years in cells with very few papers (health has 114 papers in the full corpus and an even smaller method-mapped subset; governance has 39 method-mapped papers). The first publication year in a cell is an extreme order statistic: a single missing early paper from an unindexed practice-oriented venue, or one LLM mislabel, shifts t_{a,m} by years. Table 1 reports only aggregate Cohen's kappa, accuracy, and macro-F1 over validation samples of 403–500 papers; there are no per-task or per-method error rates, and no confidence intervals or bootstrap estimates are reported for TIR, TMIR, NMD, or FMAL. Section 7 acknowledges that first-adoption years are sensitive to a few early papers, but no sensitivity analysis is provided. This gap is load-bearing because the abstract and discussion state the lag contrast as a main finding; without per-cell error rates and a leave-one-out or bootstrap analysis, the quantitative strength of that finding is not yet established.","section":"Section 5.3 / Eqs. (7)–(8)"},{"comment":"The frontier-method set F is introduced in Eq. (7) only as 'F⊆M', and the three frontier families used in Section 5.3 are never named in the methods; the reader must infer Transformer, diffusion, and foundation models from context. Because FMAL_a is averaged only over the methods in F_a that were adopted, the set's membership directly changes the reported average lags, and unadopted methods are excluded rather than penalized. The manuscript also gives inconsistent taxonomy counts: the abstract says 12 application categories and 11 method families, Section 3.2 says 12 primary categories for each dimension, Figure 2 says 13 application categories and 12 method categories, Table S1 lists A0–A12, and Table S2 lists M0–M11. For NMD's denominator log|M| to be interpretable and for FMAL to be reproducible, the paper must define F and M explicitly and reconcile these counts.","section":"Section 4 / Figure 2"},{"comment":"The claim that education, health, and governance are 'under-supported' depends on the corpus covering those literatures and on the LLM-assigned primary labels being accurate in the small cells. The retrieval is keyword-based from Semantic Scholar, ISMIR, and arXiv, which may under-represent practice-oriented venues where music education, music therapy, and policy/governance research is published; the limitations paragraph concedes that the corpus 'does not cover all relevant publications' but does not quantify this risk. In addition, the stratified validation in Algorithm 1 validates labels over all tasks pooled, so it cannot rule out that label errors are concentrated in exactly the small categories that drive TIR and TMIR. I would like to see per-category validation statistics and a sensitivity analysis that re-computes the main indicators under a multi-label assignment or under alternative corpus-retrieval rules; the current evidence supports a qualitative direction but not the precise residual magnitudes.","section":"Section 3.1 / Algorithm 1"}],"minor_comments":[{"comment":"The petal lengths are described qualitatively; please add a legend or a quantitative mapping so the reader can relate petal length to TIR, TMIR, NMD, and FMAL values.","section":"Figure 1"},{"comment":"The text sets the acceptance threshold tau_Acc = 0.90, but the pseudocode only checks tau_kappa and tau_F1; either include tau_Acc in the algorithm or remove it from the prose.","section":"Algorithm 1"},{"comment":"The LLM used for taxonomy mapping is not identified; Gemini-3.1-Flash-Lite is named only for taxonomy generation. Please state the model, prompt version, and inference settings used for the final label assignment to support reproducibility.","section":"Section 3.3"},{"comment":"The caption says 'negative values indicate low expected investment'; this should read 'below-expected investment' to match the definition of TIR in Eq. (2).","section":"Figure 5 caption"}],"recommendation":"major_revision","confidential_remarks":"This is an honest measurement paper whose framework is coherent and whose qualitative conclusions are plausible. The main risk is overclaiming precision: the adoption-lag and residual magnitudes are presented as crisp numbers despite sparse cells, pooled validation, and no uncertainty analysis. The required fixes are feasible within the manuscript's scope: add per-cell validation, bootstrap/leave-one-out sensitivity for FMAL and TIR/TMIR, define F and the taxonomy counts explicitly, and address the venue-coverage concern. The taxonomy count inconsistencies should also be reconciled before resubmission. No concerns about novelty or scope; the paper fits a bibliometric/science-of-science venue. The reader's and stress-test concerns align with my reading of the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this paper is a bibliometric measurement with a real new contribution—a joint application/method taxonomy for AI music and a four-indicator profile (TIR, TMIR, NMD, FMAL) that lets you see where technical attention goes. The qualitative finding is credible: content tasks like generation and audio processing get more technical investment and adopt frontier methods earlier; education, health, and governance lag. That direction will not surprise anyone who works in the field, but the paper is the first to quantify it in one framework.\n\nWhat it does well: the methods are transparent. The residual definitions are simple and clearly specified, the proportional baselines are sensible, and the LLM labeling is validated on stratified samples with strong inter-annotator agreement (kappa 0.93–0.98). The limitations section is honest—it flags incomplete coverage, single-label simplification, and sensitivity of first-adoption years.\n\nThe soft spots are real but proportionate. The headline numbers—0.33 years vs 4.33 and 5.00 lag—depend on first-adoption years in very sparse cells. Health has 114 papers total and a smaller technical subset; one missing early paper or one mislabel changes the lag by years. There is no uncertainty propagation, no bootstrap, no per-task error rates. The paper mentions first-adoption sensitivity in the limitations but does no sensitivity analysis, which is exactly where it is needed. Also, the data and code are not public (available from author upon request), so the central estimates are not independently reproducible. There is also an internal inconsistency: the text says 12 primary application categories and 11 method families in the abstract, then 13 and 12 in Section 3.2. Minor, but sloppy.\n\nI disagree with the stress-test note in one way: it overstates fragility. The qualitative direction is robust across multiple indicators—TIR, TMIR, and NMD all point the same way, and the adoption-lag gaps are consistent with the residuals. A single missing paper would move a lag, but it would not erase the pattern. So the framework holds; only the precise year gaps should be read as illustrative until sensitivity analysis is done.\n\nWho this is for: people working on bibliometrics of creative AI, or research-policy folks in music technology. It deserves a serious referee, but the referee should ask for a sensitivity analysis on FMAL and TIR, and ideally a release of dataset/code. Recommend: accept with major revision, conditional on uncertainty quantification.","headline":"A transparent, useful bibliometric framework for AI music, but the headline lag numbers rest on sparse cells and need sensitivity analysis before they carry weight.","tokens_in":12869,"tokens_out":1696,"would_cite":false,"duration_ms":16704,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI music research attention is systematically skewed toward scalable content tasks, leaving education, health, and governance under-supported.","keywords":["AI music","research attention imbalance","scientometrics","frontier-method adoption lag","technical investment residual","music education","music health","methodological diversity"],"falsifier":"Re-run the four indicators on an expanded corpus that adds education, health, and governance venues beyond the three sources used in the paper—for example, music-education journals, clinical music-therapy journals, and arts-policy venues—and check whether education, health, and governance still show negative TIR and multi-year FMAL values; if their technical support becomes comparable once those venues are counted, the imbalance would not reproduce.","tokens_in":11798,"feed_emoji":"🎵","tokens_out":10007,"duration_ms":87960,"temperature":0.7,"pith_summary":"AI music research has grown rapidly, but this paper argues its attention is lopsided: technical investment and frontier-method adoption concentrate in tasks with scalable data and standardized benchmarks, particularly music generation, information retrieval, and audio processing, while education, health, and governance receive less technical support and adopt new methods years later. To make this imbalance measurable, the paper introduces the Research Attention Profile, four indicators computed over a corpus of 6,839 publications from 2015 to April 2026, labeled with a joint taxonomy of application tasks and technical method families. The headline numbers are clear: generation adopted the Transformer, diffusion, and foundation-model families with an average lag of 0.33 years, versus 4.33 years for education and 5.00 years for health. A sympathetic reader would take the paper as establishing that the field's methodological progress is channeled not by societal relevance but by task-level capacity to absorb emerging methods.","feed_headline":"Frontier music AI reaches health five years later than generation","feed_subtitle":"A 6,839-paper analysis shows effort and adoption cluster in generation, retrieval, and audio processing.","key_machinery":"The carrying mechanism is the Research Attention Profile, a set of four indicators computed from a jointly labeled corpus of AI music publications. Technical Investment Residual (TIR) compares a task's observed number of technical papers to the number expected from its publication volume and the overall annual share of technical papers. Task-Method Investment Residual (TMIR) performs the same comparison for a specific method family within a task. Normalized Methodological Diversity (NMD) is a normalized Shannon entropy of method proportions within a task, ranging from 0 (one method dominates) to 1 (methods evenly spread). Frontier Method Adoption Lag (FMAL) measures the years between a frontier method's first appearance in the AI music corpus and its first appearance within a given task. Together, the four indicators separate the amount of technical support, its allocation across methods, the structural diversity of that support, and the timing of adoption.","core_discovery":"The paper's central discovery is that research attention in AI music is systematically uneven, and the unevenness has a structure: content- and model-intensive tasks receive stronger and earlier technical support than socially embedded tasks. Music generation and creation (1,522 papers) and music information retrieval (1,496) dominate the corpus, while governance (163) and health (114) remain limited. Technical Investment Residuals are positive for audio processing, generation, and singing-voice technologies, and negative for education, governance, health, and datasets and benchmarks. Frontier methods are allocated most heavily to generation (diffusion TMIR of 1.317) and MIR (foundation-model TMIR of 0.973), while education, health, and governance show below-expected residuals across all three frontier families. The paper concludes that the diffusion of frontier methods is shaped more strongly by compatibility with scalable datasets and benchmark pipelines than by the relative societal importance of the applications.","pith_inferences":["The adoption-lag metric, as defined, is sensitive to a method's very first year of appearance, so a single early paper can dominate the number; measuring the year a method reaches a meaningful share of a task's technical papers would test whether generation's 0.33-year lag is a robust early-adoption pattern or an artifact of one paper.","The same four-indicator profile could be applied to other AI domains, such as educational technology, clinical AI, or computational law, to test whether a systematic under-support of socially embedded, non-benchmark-driven tasks is a general feature of AI research rather than specific to music.","The corpus cutoff in April 2026 may undercount very recent diffusion-based studies in health and governance; re-running the analysis after additional years would clarify whether those areas have a lasting lag or simply a longer publication cycle."],"forward_implications":["The gap between content-oriented and socially embedded applications is structural, not just a matter of volume: even after controlling for publication count, generation and audio processing receive more technical investment than expected while education, governance, and health receive less.","Frontier methods diffuse unevenly across tasks, with AI music generation adopting the Transformer, diffusion, and foundation-model families within months on average, while education and health lag by over four years and no diffusion-based health study appears in the study period.","Methodological diversity does not compensate for limited investment, as governance shows high diversity (NMD of 0.864) but on only 39 method-mapped papers, reflecting dispersion within a small literature rather than broad support.","Closing the gap requires not only model development but task-appropriate datasets, context-sensitive evaluation, longitudinal validation, and interdisciplinary collaboration in education, health, and governance."],"supporting_citations":[{"why":"Grounds the application-task definition of music generation that anchors the content-oriented side of the imbalance.","marker":"[3]"},{"why":"Establishes the model-driven perspective the paper's technical-method taxonomy extends.","marker":"[9]"},{"why":"Provides the prior bibliometric inequality result the paper shifts from authorship to technical investment.","marker":"[11]"},{"why":"Anchors the Transformer family in the frontier-method set used for adoption-lag measurement.","marker":"[18]"},{"why":"Anchors the diffusion family in the frontier-method set used for adoption-lag measurement.","marker":"[19]"},{"why":"Defines the foundation-model family, the third frontier-method family whose allocation is measured.","marker":"[21]"},{"why":"Supplies a field-level bibliometric baseline across creation, performance, and education that the imbalance analysis refines.","marker":"[35]"}],"fun_headline_variants":["AI music research: 5-year frontier gap for health vs generation","Music AI: generation adopts frontier in 0.3 yrs, health in 5","AI music neglects health: a 6,839-paper study of attention","Where does AI music innovate? Not in health, says new analysis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the retrieved corpus of 6,839 publications, the single-primary-label mapping validated on samples of 403–500 papers, and the indicators derived from those labels faithfully represent the field's true research attention; if relevant practice-oriented venues were missed, or if labeling errors correlate with task type, the reported under-support of education, health, and governance could be exaggerated.","fun_headline_variants_meta":{"raw":{"variants":["AI music research: 5-year frontier gap for health vs generation","Music AI: generation adopts frontier in 0.3 yrs, health in 5","AI music neglects health: a 6,839-paper study of attention","Where does AI music innovate? Not in health, says new analysis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000676,"raw_usage":{"total_tokens":3067,"prompt_tokens":928,"completion_tokens":2139,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":2057}},"tokens_in":544,"tokens_out":2139,"duration_ms":15744,"temperature":1.0,"reasoning_tokens":2057,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:39:08.098566+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the four indicators on an expanded corpus that adds education, health, and governance venues beyond the three sources used in the paper—for example, music-education journals, clinical music-therapy journals, and arts-policy venues—and check whether education, health, and governance still show negative TIR and multi-year FMAL values; if their technical support becomes comparable once those venues are counted, the imbalance would not reproduce.","supporting_citations":[{"cited_title":"Springer, 2020","cited_arxiv_id":null,"evidence_quote":"Establishes the model-driven perspective the paper's technical-method taxonomy extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prior bibliometric inequality result the paper shifts from authorship to technical investment."},{"cited_title":"Non-technical / Not model-focused","cited_arxiv_id":null,"evidence_quote":"Supplies a field-level bibliometric baseline across creation, performance, and education that the imbalance analysis refines."}],"review_version":1}